HITL method and system for unloading large model reasoning task in heterogeneous GPU computing power cloud

By introducing human expert feedback and multi-agent collaboration mechanisms into heterogeneous GPU computing cloud, combined with optimized prompt word technology, the problem of user intent recognition bias in large models in heterogeneous environments is solved, achieving efficient and accurate task scheduling and resource utilization.

CN121979667APending Publication Date: 2026-05-05NARI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NARI TECH CO LTD
Filing Date
2025-12-24
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing large models lack a deep understanding of the physical resource environment in heterogeneous GPU computing cloud environments, leading to biases in user intent recognition and difficulty in accurately scheduling tasks in complex scenarios.

Method used

By introducing human expert feedback and multi-agent collaboration mechanisms, combined with optimized chained prompting technology, and using the HITL method to accurately map user intent into executable utility functions, and by employing modular agent design and multi-round interaction, the system's adaptability and interpretability are enhanced.

Benefits of technology

It improves the accuracy and personalization of large model inference services, reduces end-to-end task response latency, increases GPU cluster computing power utilization, optimizes resource collaborative utilization, and reduces overall computing power costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979667A_ABST
    Figure CN121979667A_ABST
Patent Text Reader

Abstract

The invention discloses an HITL method and system for unloading a large model reasoning task in a heterogeneous GPU computing power cloud. The method comprises the following steps: constructing a system model of the heterogeneous GPU computing power cloud; an HITL system architecture based on multiple agents is designed; human expert feedback is introduced through an HITL method, and recognition of a large model on a user intention is corrected; human expert feedback is introduced through an HITL method, and recognition of a large model on a user intention is corrected; the multi-agent cooperation module maps the intention of a user to a utility function through interaction among a plurality of agents; an optimized cue word method is used, the thinking process of a large model is decomposed, and the intention recognition accuracy is improved in a chain reasoning mode. According to the method, dynamic scheduling of computing power resource management and large model reasoning tasks can be optimized, and the service quality and the user experience in the heterogeneous GPU computing power cloud environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of cloud computing, large model inference, and heterogeneous GPU computing resource scheduling and optimization, and in particular to a HITL method and system for offloading large model inference tasks in a heterogeneous GPU computing cloud. Background Technology

[0002] With the development of artificial intelligence technology, especially the widespread application of large pretrained models, the demand for computing power for AI inference tasks is growing exponentially. The inference process of large language models often involves the calculation of billions or even hundreds of billions of parameters, resulting in a significant consumption of computing resources, especially high-performance GPU resources. Against this backdrop, heterogeneous GPU computing cloud has gradually become the core infrastructure for supporting large model inference tasks. By dynamically scheduling inference tasks among GPU resources of different models and performance levels, heterogeneous GPU computing cloud achieves optimal resource utilization, becoming an important means to improve AI inference efficiency and service quality.

[0003] In large-scale model inference tasks, task offloading technology becomes crucial for improving system performance and resource utilization due to the massive scale and high resource consumption. However, in actual deployment, task offloading decisions are influenced not only by computing resource availability, task characteristics, and network conditions, but also by the user's actual needs and goals, i.e., user intent. For large-scale model inference services, user intent typically affects task priority, resource scheduling strategies, and the way inference results are generated. Especially in heterogeneous GPU environments, user intent directly impacts the execution strategies of inference tasks (such as model selection, scheduling priority, and whether to use parallel or distributed inference). Therefore, accurately identifying user intent is key to ensuring the intelligence and personalization of large-scale model inference services.

[0004] However, current methods for identifying user intent primarily rely on traditional approaches, small models, or mathematical methods. These methods typically depend on rules, statistical models, or small datasets for training, enabling inference over trained data ranges. However, due to their low model complexity and poor generalization ability, they struggle to adapt to the diversity and complexity of user needs. Large models, on the other hand, can learn more complex patterns from massive amounts of data and exhibit strong reasoning and predictive capabilities. However, current large model methods cannot address the understanding of the physical world and lack a deep understanding of the dynamic changes in devices and environments, leading to inaccurate identification of user intent and limiting their application in complex scenarios. Furthermore, due to the inherent errors in understanding, limitations in reasoning capabilities, and context management issues of large models, a gap exists between the actual execution of task scheduling and the user's intent during the unloading of large model inference tasks. Therefore, accurately identifying user intent and optimizing intent targeting have become one of the core challenges in unloading large model inference tasks. Summary of the Invention

[0005] Purpose of the Invention: The purpose of this invention is to provide a HITL method and system for offloading large model inference tasks in heterogeneous GPU computing cloud environments. By introducing human expert feedback and multi-agent collaboration mechanisms, combined with optimized chained prompting word technology, the user intent is accurately mapped into an executable utility function. This solves the problem of intent recognition bias caused by the lack of understanding of the physical resource environment in existing large models, thereby narrowing the gap between task scheduling execution results and the user's true intent, and improving the accuracy, personalization, and scheduling efficiency of large model inference services in heterogeneous computing environments.

[0006] Technical solution: The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud, as described in this invention, includes the following steps:

[0007] (1) Construct a system model for a heterogeneous GPU computing power cloud, the system model including a user access layer, an edge cloud computing power layer and a central cloud computing power layer;

[0008] (2) Design a multi-agent HITL system architecture, the HITL system architecture including:

[0009] Intelligent agent model: The basic unit of the architecture, which adopts a modular design, including the agent, the memory module that interacts and integrates with the agent, the action module, the perception module, and the brain module;

[0010] HITL module: Interacts with the brain module and perception module of the intelligent agent model, providing a channel for human expert intervention in the architecture;

[0011] Multi-agent collaboration module: Defines and manages multiple role-based agents, including operations engineer agents, dispatcher agents, analyst agents, and customer agents; each agent encapsulates specific toolsets and knowledge bases based on its role and function, and conducts multi-round, structured dialogues and information exchanges;

[0012] The prompt word engineering module, based on an improved thought chain method, structures prompt words into "understanding-analysis-output". It dynamically embeds the current task context, multi-agent interaction rounds, specified interaction objects, forced interaction modes and structured output formats for each stage. This module receives interaction status from the multi-agent collaboration module and expert feedback from the HITL module, and dynamically generates and optimizes prompt words for each agent's brain module.

[0013] (3) Human expert feedback is introduced through the HITL method to correct the large model’s recognition of user intent; the human experts are personnel with specific domain expertise and system operation permissions to supervise, evaluate and correct the artificial intelligence decision-making process.

[0014] (4) The multi-agent collaboration module maps subjective user intentions into objective utility functions through multi-round collaboration among multiple agents;

[0015] (5) Use optimized prompt words to decompose the thinking process of the large model and improve the accuracy of intent recognition through chain reasoning.

[0016] Furthermore, in step (1), the user access layer handles lightweight inference tasks, the edge cloud computing layer undertakes complex tasks, and the central cloud computing layer handles computing power intensive loads.

[0017] Furthermore, in step (2), the memory module is divided into long-term memory, medium-term memory and short-term memory. The medium-term memory extracts key information through word slots. The action module integrates tool invocation and action execution, and invokes tools in real time through a large model. The perception module includes the system model and prompt word design, and dynamically generates input text. The brain module integrates multiple large language models to evaluate and adjust QoSE goals.

[0018] Furthermore, in step (2), the multi-agent collaboration involves an operations engineer agent, a scheduler agent, an analyst agent, and a customer agent, and achieves dynamic optimization of task scheduling through a multi-round interaction process.

[0019] Furthermore, the operations engineer agent acquires the system network status and summarizes the information; the scheduler agent performs task scheduling based on the summarized information and analysis results; the analyst agent analyzes the scheduling results and provides assurance status; and the customer agent proposes adjustment suggestions.

[0020] Furthermore, the prompting word method described in step (5) is based on an improvement on thought chain, including:

[0021] The reasoning process is structured and divided into three stages: understanding, analysis, and output.

[0022] Add interactive information for the character, including the interactive object, interactive mode, and interactive format;

[0023] Add information about the current interaction round number.

[0024] A HITL system for offloading large model inference tasks in a heterogeneous GPU computing cloud includes the following modules:

[0025] Heterogeneous GPU computing power cloud resource management module: performs unified abstraction, discovery, monitoring and pooling management of heterogeneous GPU resources in the user access layer, edge cloud computing power layer and central cloud computing power layer, and provides resource status query interface and task unloading execution interface for the upper layer;

[0026] Intelligent Agent Engine Module: This module includes: (a) Intelligent Agent Instantiation Unit: Dynamically creates or calls intelligent agent instances of different roles such as operation and maintenance engineers, schedulers, analysts, and customers according to task requirements; (b) Intelligent Agent Kernel: Encapsulates memory, action, perception, and brain modules for each intelligent agent instance;

[0027] Human-machine collaborative interaction module: Provides an expert operation interface. When the agent's decision confidence is low or a preset rule is triggered, it actively suspends the automated process, visualizes the decision context, inference chain, and candidate solutions to the human expert, and receives the expert's correction instructions or direct decision.

[0028] Multi-agent collaboration module: Defines the communication protocol, interaction process and conflict resolution mechanism between agents, drives and manages agents with different roles to collaborate according to the established process, and jointly completes the complete closed loop from understanding user intent to generating an optimized scheduling scheme.

[0029] Dynamic Prompt Engineering and Optimization Module: Maintains a structured prompt template library and can dynamically assemble and optimize prompts sent to the "brains" of each agent based on the current task type, collaboration stage, interaction rounds, and expert feedback from the HITL module.

[0030] Task scheduling and unloading execution module: Receives the scheduling strategy finally generated by the multi-agent collaboration module and confirmed by the HITL module, and converts it into executable instructions;

[0031] System support modules: The processor is responsible for executing the logical calculations and scheduling of each module; the storage medium is used for persistent storage of system configuration, agent's long-term / medium-term memory, historical interaction logs, prompt word templates, expert knowledge base and task data, etc.

[0032] Furthermore, the brain module integrates and manages calls to multiple large language model APIs to perform core reasoning.

[0033] A computer device, comprising:

[0034] Memory: Used to store executable instructions and system configuration data, agent model parameters, prompt word template library and interaction logs generated by the multi-agent HITL system architecture;

[0035] Processor: Coupled to the memory, configured to execute the executable instructions; call and run the prompt word engineering module to perform dynamic optimization and assembly of prompt words; coordinate the interaction and communication between multiple agents in the multi-agent collaboration module; and manage the interaction process with human experts through the interface provided by the HITL module.

[0036] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a HITL method for offloading the large model inference task in a heterogeneous GPU computing cloud.

[0037] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. It integrates multi-agent collaboration and human-in-the-loop feedback, providing an innovative solution for efficient scheduling of large-scale model inference tasks; by constructing a multi-agent HITL system architecture, it achieves accurate mapping from user intent to utility function; 2. By adopting a hierarchical edge-central cloud system model and a multi-agent collaboration mechanism, and through optimized prompt word engineering and chained reasoning methods, it can improve the accuracy of large-scale model intent recognition; in actual deployment, this method can reduce end-to-end response latency of tasks, while improving the overall GPU cluster computing power utilization rate, effectively improving system processing efficiency; 3. Through an improved thought chain (CoT) prompt word structure and modular agent design, it enhances the interpretability of the system decision-making process; combined with a human expert feedback correction mechanism, it enables the system to maintain automated scheduling efficiency while possessing the ability for human intervention in key decisions, significantly improving the system's adaptability and reliability in complex scenarios; 4. Through intelligent load balancing and resource scheduling strategies, it optimizes the collaborative utilization of edge and central cloud resources, reducing overall computing power costs while ensuring service quality. Attached Figure Description

[0038] Figure 1 This is a system model diagram of the heterogeneous GPU computing power cloud of the present invention;

[0039] Figure 2 This is a diagram of the HITL system architecture based on multiple agents according to the present invention;

[0040] Figure 3 This is a diagram of the intelligent agent model of the present invention;

[0041] Figure 4 This is a basic flowchart of the HITL process of the present invention;

[0042] Figure 5 This is a flowchart illustrating the interaction process of the multi-agent system in this invention.

[0043] Figure 6 This is a flowchart illustrating the overall workflow of the present invention. Detailed Implementation

[0044] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0045] The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud, as described in this invention, includes the following steps:

[0046] (1) Constructing a system model for a heterogeneous GPU computing cloud, the system model of which includes three layers: user access layer, edge cloud computing layer, and central cloud computing layer, such as Figure 1 As shown.

[0047] User Access Layer: As the system front-end, this layer consists of lightweight GPU nodes (such as embedded GPUs or T4 / A10-level GPUs on edge servers) deployed near the user. Its functions include:

[0048] It can receive large model inference requests submitted by users in real time, such as large language model question answering and image generation tasks.

[0049] Preprocessing and offloading: Lightweight processing such as compression and format conversion is performed on the input data, and decisions are made on whether to offload the task to the upper layer based on the task complexity and local GPU resource status (memory margin / computing load);

[0050] Low latency response: For latency-sensitive lightweight tasks (such as small-scale inference after model pruning), it directly calls local heterogeneous GPU resources to execute, reducing cloud transmission overhead.

[0051] Edge cloud computing layer: Composed of mid-range heterogeneous GPU clusters (such as A100 / L40 clusters) in regional data centers, possessing medium-scale computing power and high-speed network interconnection. This layer's functions include:

[0052] Undertake offloading tasks: Receive complex inference tasks (such as inference for a model with tens of billions of parameters) forwarded by the edge access layer.

[0053] Heterogeneous resource collaboration: Integrate heterogeneous GPUs across nodes (such as hybrid deployment of A100 / H100 / MI300X) through virtualization technology and dynamically allocate computing resources;

[0054] Near-end optimization: Utilize geographical proximity to reduce network latency and provide tiered computing power guarantees for latency-sensitive tasks (such as prioritizing scheduling to GPUs with high video memory).

[0055] Central Cloud Computing Layer: Composed of top-tier heterogeneous GPU clusters in the core data center (such as H100 SuperPOD and DGX systems), providing ultra-large-scale computing power and global scheduling capabilities. This layer's functions include:

[0056] Heavy-duty task execution: handling computationally intensive workloads such as inference of large models with hundreds of billions of parameters and multimodal fusion tasks;

[0057] Global resource pooling: Abstracts heterogeneous GPUs across regions (cloud, edge, and device resources) into a unified computing power pool, supporting fine-grained resource slicing.

[0058] (2) Design a HITL system architecture based on multiple agents. For example... Figure 2 As shown, the system architecture consists of an agent model, a HITL module, a multi-agent collaboration module, and a prompt word engineering module.

[0059] The system model re-divides the intelligent agent into four modules: agent, memory, action, perception, and brain. This allows the intelligent agent to better adapt to the characteristics of heterogeneous GPU computing cloud environments, thereby solving problems such as resource constraints, dispersed computing power, high latency, and dynamic changes in these environments, and improving the system's real-time response and adaptability. The intelligent agent model is described as follows:

[0060] (a) Agent: The intelligent agent acts as a proxy for the large language model (LLM), responsible for interacting with the user and various auxiliary systems (including knowledge, memory, tools, and external systems) to achieve efficient task execution optimization.

[0061] (b) Memory Module: The traditional agent model divides memory into long-term memory and short-term memory. The memory module adds intermediate memory to this, that is, divides memory into long-term memory, intermediate memory and short-term memory.

[0062] Long-term memory contains information that an intelligent agent needs over a long period of time, and it usually contains a lot of information, such as project information and background. Intermediate memory contains some key information, which contains less information, but it is also essential information for an intelligent agent over a long period of time, such as time and place.

[0063] Short-term memory contains information that an agent needs in the short term, and it contains relatively little information, such as the history of interactions in previous rounds.

[0064] Long-term memory is accessed through RAG. RAG is a model architecture that combines retrieval and generation, allowing the agent to enhance its generation capabilities by retrieving relevant information from external databases or knowledge bases, thereby improving the access efficiency of long-term memory. Intermediate memory is extracted through slot extraction. Slots refer to placeholders for specific information or data, representing the storage location of specific semantic information, thus enabling the agent to quickly extract specific contextual information.

[0065] Short-term memory is retrieved directly from historical records. Because the contents of short-term memory are overwritten or forgotten over time, and long-term memory may be too vast to store frequently used critical information, medium-term memory stores this crucial information, avoiding information loss and forgetting, and ensuring that the agent can accurately retrieve relevant information when needed.

[0066] Furthermore, intermediate memory reduces reliance on external databases or knowledge bases by storing critical information locally. This design not only improves the system's robustness (allowing it to continue functioning even when the network is unavailable or external services are unreachable) but also reduces latency in accessing external resources.

[0067] (c) Action Module: This includes various tools that agents can use, such as simulation tools, query tools, and algorithm simulation tools. For tool usage, the large model's function calling feature can be used, allowing the large language model (LLM) to autonomously select the necessary tool functions to call based on the user's problem description, the associated context, and background information.

[0068] The Action module is responsible for how the agent selects appropriate action steps and executes tasks through decision-making and planning. It directly integrates the Tool module and the Action module, synchronizing tool invocation and execution. By combining the tool set of the Tool module with the Function Calling function of the Large Language Model (LLM), the Action module can invoke tools in real-time via Function Calling when executing tasks. In this way, the Action module enables the agent to adjust its strategy based on real-time feedback and select appropriate execution tools (such as simulation tools, query tools, etc.). Simultaneously, the agent can achieve timely and accurate actions even under network latency and bandwidth limitations.

[0069] (d) Perception Module: This module comprises two parts: the system model and the design of prompt words. Prompt words are used to guide the model in generating specific output input text. The design of prompt words involves creating appropriate prompts based on different scenarios. By dynamically generating suitable input text and combining environmental feedback, interaction history, and simulation results, the accuracy and response speed of the perception system are improved. This allows perception data from edge nodes to be rapidly transmitted to the model, making the agent more flexible and adaptable, and reducing reliance on the central cloud.

[0070] (e) Brain Module: The agent's brain, integrating multiple large language models (e.g., ChatGpt, llama, Qwen, etc.) and responsible for evaluating and adjusting QoSE (Quality of Service and Experience) goals. Based on real-time evaluation results and user feedback, the Brain Module supports seamless collaboration between edge nodes and the cloud through a distributed computing framework. This enables the agent to perform efficient decision-making and reasoning in a heterogeneous GPU computing cloud environment and achieve dynamic task optimization.

[0071] HITL module: Interacts with the brain module and perception module of the intelligent agent model, providing a channel for human expert intervention in the architecture;

[0072] Multi-agent collaboration module: Defines and manages multiple role-based agents, including operations engineer agents, dispatcher agents, analyst agents, and customer agents; each agent encapsulates specific toolsets and knowledge bases based on its role and function, and conducts multi-round, structured dialogues and information exchanges;

[0073] The prompt word engineering module, based on the improved Chain of Thought (CoT) method, structures prompt words into an "understand-analyze-output" process. It dynamically embeds the current task context, multi-agent interaction rounds, specified interaction objects, forced interaction modes, and structured output formats for each stage. This module receives interaction status from the multi-agent collaboration module and expert feedback from the HITL module, and dynamically generates and optimizes prompt words for each agent's brain module.

[0074] The basic workflow of HITL is the interaction process between the user and the agent, and the steps of the interaction process are as follows:

[0075] Step 1: The user inputs a description of the task into the Agent, including the type / arrival rate / priority / performance metrics (power, latency), as well as a description of the target formula, template / performance metric scoring rules, and other parameter information.

[0076] Step 2: Based on the user-input task description and target formula description, the Agent uses RAG (Retrieval Augmented Generation) to retrieve relevant knowledge bases using augmented generation techniques, returning knowledge documents related to the user's description. Then, it uses the memory module to query the user's and Agent's interaction history to understand the dialogue context. The knowledge retrieved by RAG is used as the Agent's background information, and the user's and Agent's interaction history is used as the Agent's historical dialogue information. Combining the background information and historical dialogue information, HITL cue words are constructed and used for inference queries in the Large Language Model (LLM).

[0077] Step 3: The Large Language Model (LLM) considers the user's direct requirements, incorporates relevant information retrieved from the knowledge base, and their reasoning ability to generate query results.

[0078] Step 4: The agent comprehensively considers the results returned by the Large Language Model (LLM) and generates an optimized objective function.

[0079] Step 5: The agent analyzes and calls the most suitable optimization simulation tool from the tool library to perform simulation, extracts the simulation results, uses the simulation results as background information to form new prompt words, and requests the LLM to analyze the results.

[0080] Step 6: Results of Large Language Model (LLM) analysis.

[0081] Step 7: The agent summarizes and organizes the interaction process of this round, and combines it with historical information to provide complete feedback to the user.

[0082] Step 8: Users comment on and adjust the results generated by the Agent, and input the content that needs to be adjusted into the Agent.

[0083] Step 9: The system loops through steps 2 through 8 until the user is satisfied with the result. After proposing the basic HITL workflow, a multi-agent collaboration method is introduced. For task offloading in edge cloud, four agents are defined—an operations engineer, a scheduler, an analyst, and a customer—who collaborate to complete task offloading. Finally, combining the HITL method with multi-agent collaboration, the multi-agent workflow is used as an agent. HITL allows users to monitor and provide feedback on the entire multi-agent workflow process. Then, the four modules of the agent model are integrated to finally construct a multi-agent HITL system architecture.

[0084] (3) Human expert feedback is introduced through the HITL method to correct the large model's recognition of user intent. The human experts are personnel with specific domain expertise and system operation permissions who supervise, evaluate, and correct the AI ​​decision-making process. HITL emphasizes human participation and supervision in the AI ​​decision-making process, including the interaction flow between the user and the agent, such as... Figure 4 As shown, the Agent is replaced with a multi-agent workflow, and the HITL method allows users to supervise and provide feedback on the entire multi-agent workflow process.

[0085] (4) The multi-agent collaboration module maps subjective user intentions into objective utility functions through multiple rounds of collaboration between multiple agents.

[0086] The interaction process between multiple agents is as follows: Figure 5 As shown, after multiple rounds of collaboration among multiple agents, subjective user intentions are mapped to objective utility functions, enabling more accurate identification of user intentions. The overall workflow diagram is as follows: Figure 6 As shown.

[0087] exist Figure 5 The multi-agent interaction flow shown involves four agents: an operations engineer, a scheduler, an analyst, and a customer. The job description of each agent is as follows:

[0088] Operations Engineer: The operations engineer agent is responsible for interacting with the dispatcher agent. The operations engineer agent needs to use tools to obtain system network status, such as the overall network usage, base station wireless bandwidth usage, MEC computing resource usage, and the computing resource usage of each terminal, and then summarize these data.

[0089] Scheduler: The scheduler agent is responsible for interacting with the operations engineer agent. In the first round, the scheduler agent needs to perform the first scheduling based on the information summarized by the operations engineer agent (such as the overall network usage, base station wireless bandwidth usage, MEC computing resource usage, and computing resource usage of each terminal). In subsequent rounds, in addition to the information summarized by the operations engineer agent, the scheduler also needs to gather the analysis results from the analyst agent, the adjustment requests from the customer agent, and the scheduling results from the previous round, and then call the scheduling tool to perform the next round of scheduling.

[0090] Analyst: The analyst agent is responsible for interacting with the scheduler agent. The analyst agent needs to analyze the scheduling results from the scheduler agent and, in conjunction with the technical solution, provide guarantees for scheduling across various users, services, and metrics.

[0091] Client: The client agent is responsible for interacting with the analyst agent. The client agent needs to comment on the results analyzed by the analyst agent and, based on the current background information, propose corresponding adjustment suggestions.

[0092] (5) Use optimized prompt words to decompose the thinking process of the large model and improve the accuracy of intent recognition through chain reasoning.

[0093] By using optimized prompts to break down the thought process, large models can more accurately decompose problems and perform chained reasoning during task execution, thereby better understanding user intent and compensating for the lack of flexibility and context adaptability in the reasoning process of the Chain of Thought (CoT) method. Furthermore, the new prompt method enhances the system's adaptability in heterogeneous GPU computing cloud environments, enabling it to better cope with complex network conditions and dynamically changing hardware environments.

[0094] The improvements made based on the CoT (CoT) framework are as follows:

[0095] Structured Reasoning Process: Based on the ideas of Cognitive Load Theory (CLT) and Systems Thinking, the requirement of a structured reasoning process is added to the prompts, that is, the reasoning process is required to be executed step by step in three clear steps: understanding, analysis and output.

[0096] Cognitive load theory posits that working memory, with its limited capacity, is prone to overload when processing complex tasks, thus impacting learning effectiveness and problem-solving abilities. However, structured reasoning processes can break down complex tasks into three or more steps, reducing the burden on working memory and improving information processing efficiency.

[0097] Systems thinking emphasizes viewing problems from a holistic, interconnected, and dynamic perspective, believing that any problem is part of a complex system in which the various elements are interconnected and interact with each other.

[0098] By using a structured reasoning process (understanding, analyzing, and outputting), we can better reveal the essence of the problem, identify the key factors in the system and their interrelationships, and thus propose more effective solutions.

[0099] In the "understanding" phase, the Large Language Model (LLM) first needs to carefully read and understand the background and key information of the problem. The purpose of this phase is to ensure that all input information is fully interpreted and to identify the core requirements of the problem.

[0100] In the “analysis” phase, the Large Language Model (LLM) performs detailed reasoning and step decomposition, breaking down complex problems into smaller, more manageable parts.

[0101] The "output" stage is the summary and presentation of the reasoning process. The Large Language Model (LLM) generates the final conclusion based on the previous analysis stage and expresses it clearly and concisely.

[0102] By dividing the reasoning process of the Thinking Chain (CoT) method into three stages—"understanding," "analysis," and "output"—the reasoning process of large models becomes more systematic and structured.

[0103] Furthermore, each stage is optimized to improve understanding, reasoning, and output quality, enhancing the reasoning ability and accuracy of the large model. This improvement enables the large model to better understand context, avoid inconsistencies in the reasoning process, and ultimately generate higher-quality answers that better meet user needs.

[0104] Add interactive information for the characters: The prompts now include information about the interactive objects, such as the interactive object (who the current character is interacting with), the interaction mode (whether it is input or output during the interaction), and the interaction format (what the current character's input and output are specifically). This improvement enhances the large model's ability to understand the context.

[0105] Adding information about the current interaction round: In multi-round interaction experiments, the input and output of the large model may contain repeated content. Therefore, adding information about the current interaction round at the end of the prompt word avoids bias in the large model's understanding due to content repetition in a heterogeneous GPU computing cloud environment. This improvement not only enhances the accuracy of artificial intelligence results but also strengthens the system's adaptability in heterogeneous GPU computing cloud deployments.

[0106] The designed prompt word method involves four agents (operations engineer, scheduler, analyst, and customer) designing corresponding prompt words. Then, basic scheduling is performed based on these prompt words; specifically, the scheduler performs the first round of scheduling based on the summarized information provided by the operations engineer. The input is the summarized information from the operations engineer, and the output is the scheduling result from the scheduler in this round.

[0107] Optimizing intent based on prompts involves four steps. Step 1: Analyze the initial scheduling results from the perspective of an analyst. The input is the initial scheduling results from the scheduler, and the output is the initial analysis results from the analyst.

[0108] Step 2: Empathize with the client, providing feedback on the analyst's findings based on the current context, and requesting adjustments. The input is the analyst's initial analysis, and the output is the client's initial adjustment requests.

[0109] Step 3: Entering the second round, the dispatcher takes on the role of the dispatcher, gathers the previous dispatch results, the summary information from the operations engineer, the analysis results, and the adjustment requests, draws a conclusion on the convergence of dispatch indicators, and calls the dispatching tool to conduct the next round of dispatching. The inputs are the dispatch results from the first round, the analyst's analysis results, the client's adjustment requests, and the summary content from the operations engineer; the output is the dispatch results for this round (the second round).

[0110] Step 4: Repeat the above three processes until the client is satisfied with the analysis results. Finally, after multiple rounds of interactive experiments, obtain a summary of parameter adjustments, observe the changes in parameters, and analyze whether the user's needs have been met in the system.

Claims

1. A HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud, characterized in that, Includes the following steps: (1) Construct a system model for a heterogeneous GPU computing power cloud, the system model including a user access layer, an edge cloud computing power layer and a central cloud computing power layer; (2) Design a multi-agent HITL system architecture, the HITL system architecture including: Intelligent agent model: The basic unit of the architecture, which adopts a modular design, including the agent, the memory module that interacts and integrates with the agent, the action module, the perception module, and the brain module; HITL module: Interacts with the brain module and perception module of the intelligent agent model, providing a channel for human expert intervention in the architecture; Multi-agent collaboration module: Defines and manages multiple role-based agents, including operations engineer agents, dispatcher agents, analyst agents, and customer agents; each agent encapsulates specific toolsets and knowledge bases based on its role and function, and conducts multi-round, structured dialogues and information exchanges; The prompt word engineering module, based on an improved thought chain method, structures prompt words into "understanding-analysis-output". It dynamically embeds the current task context, multi-agent interaction rounds, specified interaction objects, forced interaction modes and structured output formats for each stage. This module receives interaction status from the multi-agent collaboration module and expert feedback from the HITL module, and dynamically generates and optimizes prompt words for each agent's brain module. (3) Human expert feedback is introduced through the HITL method to correct the large model’s recognition of user intent; the human experts are personnel with specific domain expertise and system operation permissions to supervise, evaluate and correct the artificial intelligence decision-making process. (4) The multi-agent collaboration module maps subjective user intentions into objective utility functions through multi-round collaboration among multiple agents; (5) Use optimized prompt words to decompose the thinking process of the large model and improve the accuracy of intent recognition through chain reasoning.

2. The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 1, characterized in that, In step (1), the user access layer handles lightweight inference tasks, the edge cloud computing layer undertakes complex tasks, and the central cloud computing layer handles computing-intensive loads.

3. The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 1, characterized in that, In step (2), the memory module is divided into long-term memory, medium-term memory and short-term memory. The medium-term memory extracts key information through word slots. The action module integrates tool invocation and action execution, and invokes tools in real time through a large model. The perception module includes the system model and prompt word design, and dynamically generates input text. The brain module integrates multiple large language models to evaluate and adjust QoSE goals.

4. The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 1, characterized in that, In step (2), the multi-agent collaboration involves an operations engineer agent, a scheduler agent, an analyst agent, and a customer agent, and achieves dynamic optimization of task scheduling through a multi-round interaction process.

5. The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 4, characterized in that, The operations engineer agent acquires the system network status and summarizes the information; the scheduler agent schedules tasks based on the summarized information and analysis results; the analyst agent analyzes the scheduling results and provides assurance status; and the customer agent proposes adjustment suggestions.

6. The HITL method for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 1, characterized in that, The prompting word method described in step (5) is based on an improvement on the thought chain, including: The reasoning process is structured and divided into three stages: understanding, analysis, and output. Add interactive information for the character, including the interactive object, interactive mode, and interactive format; Add information about the current interaction round number.

7. A HITL system for offloading large model inference tasks in a heterogeneous GPU computing cloud, characterized in that, Includes the following modules: Heterogeneous GPU computing power cloud resource management module: performs unified abstraction, discovery, monitoring and pooling management of heterogeneous GPU resources in the user access layer, edge cloud computing power layer and central cloud computing power layer, and provides resource status query interface and task unloading execution interface for the upper layer; Intelligent Agent Engine Module: This module includes: (a) Intelligent Agent Instantiation Unit: Dynamically creates or calls intelligent agent instances of different roles such as operation and maintenance engineers, schedulers, analysts, and customers according to task requirements; (b) Intelligent Agent Kernel: Encapsulates memory, action, perception, and brain modules for each intelligent agent instance; Human-machine collaborative interaction module: Provides an expert operation interface. When the agent's decision confidence is low or a preset rule is triggered, it actively suspends the automated process, visualizes the decision context, inference chain, and candidate solutions to the human expert, and receives the expert's correction instructions or direct decision. Multi-agent collaboration module: Defines the communication protocol, interaction process and conflict resolution mechanism between agents, drives and manages agents with different roles to collaborate according to the established process, and jointly completes the complete closed loop from understanding user intent to generating an optimized scheduling scheme. Dynamic Prompt Engineering and Optimization Module: Maintains a structured prompt template library and can dynamically assemble and optimize prompts sent to the "brains" of each agent based on the current task type, collaboration stage, interaction rounds, and expert feedback from the HITL module; Task scheduling and unloading execution module: Receives the scheduling strategy finally generated by the multi-agent collaboration module and confirmed by the HITL module, and converts it into executable instructions; System support modules: The processor is responsible for executing the logical calculations and scheduling of each module; the storage medium is used for persistent storage of system configuration, agent's long-term / medium-term memory, historical interaction logs, prompt word templates, expert knowledge base and task data, etc.

8. The HITL system for offloading large model inference tasks in a heterogeneous GPU computing cloud according to claim 7, characterized in that, The brain module integrates and manages calls to multiple large language model APIs, performing core reasoning.

9. A computer device, characterized in that, include: Memory: Used to store executable instructions and system configuration data, agent model parameters, prompt word template library and interaction logs generated by the multi-agent HITL system architecture; Processor: Coupled to the memory, configured to execute the executable instructions; call and run the prompt word engineering module to perform dynamic optimization and assembly of prompt words; coordinate the interaction and communication between multiple agents in the multi-agent collaboration module; and manage the interaction process with human experts through the interface provided by the HITL module.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1-6.