A system for energy-conscious LLM-based workflow keying with dynamic resource allocation
The integration of LLM inference and reinforcement learning in the workflow planning system addresses the challenges of HPC systems by achieving adaptive, energy-efficient, and transparent scheduling with real-time human intervention, balancing energy and performance in heterogeneous computing environments.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-03-12
AI Technical Summary
Existing high-performance computing (HPC) systems face challenges in dynamically balancing energy consumption with performance requirements, adapting to changing workload patterns and resource availability, and lack transparency in scheduling decisions, necessitating an intelligent, energy-efficient, and adaptive workflow scheduling system that integrates Large Language Model (LLM) capabilities.
An energy-conscious workflow planning system that combines LLM inference with reinforcement learning for dynamic resource allocation, continuous monitoring, multi-criteria optimization, and human-machine interaction, enabling intelligent, adaptive, and traceable scheduling in heterogeneous computing clusters.
The system achieves dynamic balance between energy consumption, throughput time, and system reliability through Pareto-optimal analysis, provides transparent scheduling decisions, and allows real-time human intervention, resulting in improved throughput and energy efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA OF INVENTION
[0001] The present disclosure relates to a system for energy-conscious workflow planning based on LLM with dynamic resource allocation. In particular, the present invention relates to a system for hybrid, LLM-supported workflow planning with dynamic resource allocation and performance optimization. BACKGROUND OF THE INVENTION
[0002] High-performance computing (HPC) systems execute complex scientific workflows, typically represented as directed acyclic graphs (DAGs), on heterogeneous computing clusters consisting of CPUs, GPUs, and specialized processors. These workflows are inherently energy-intensive and require careful balancing of performance goals such as minimal processing time, energy efficiency, and system reliability. Traditional workflow planning approaches rely on static heuristics, simple optimization algorithms, or basic reinforcement learning methods, which are inadequate for adapting to dynamic runtime conditions and offer limited transparency to the decision-making processes.
[0003] Existing scheduling systems are reaching their limits when it comes to handling the complex demands of modern HPC workloads. Conventional methods struggle to dynamically balance energy consumption with performance requirements, fail to adapt effectively to changing workload patterns and resource availability, and offer insufficient support for human intervention and decision transparency. The lack of traceability in the scheduling rationale makes it difficult for HPC operators to understand, validate, or, if necessary, override automated scheduling decisions.
[0004] Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in logical reasoning, natural language understanding, and solving complex problems across various application domains. However, no workflow scheduling framework has yet successfully integrated LLM-based logical capabilities into a comprehensive, energy-aware optimization system for high-performance computing (HPC) environments. There remains a pressing need for an intelligent scheduling system that combines LLM logical capabilities with adaptive learning algorithms to enable transparent, energy-efficient, and dynamically optimized workflow scheduling while ensuring meaningful human-machine collaboration in HPC operations. SUMMARY OF THE INVENTION
[0005] The present invention relates to an energy-conscious workflow planning system that integrates the inferential capabilities of a Large Language Model (LLM) with reinforcement learning-based optimization for high-performance computing environments. The system combines natural language processing, dynamic resource allocation, multi-criteria optimization, and human-machine interaction to enable intelligent, adaptive, and traceable workflow planning in heterogeneous computing clusters.
[0006] This disclosure relates to an energy-aware workflow planning system based on large language models (LLM) with dynamic resource allocation. The system comprises: a workflow input interface that receives workflow-directed acyclic graphs (DAGs), energy budget constraints, system performance constraints, and natural language queries from human operators; an LLM-based inference module connected to the workflow input interface that provides the following functions: analyzing the workflow specifications and system constraints in natural language, generating energy-aware planning recommendations based on the analyzed workflow specifications, and providing comprehensible planning rationales in natural language; and a reinforcement learning-based planning unit connected to the LLM inference module that provides the following functions: to receive the planning recommendations of the LLM inference agent, to fine-tune task-resource assignments by dynamically adapting to runtime fluctuations, and to adjust resource allocation online while considering runtime variability; a power monitoring unit that provides the following functions: to continuously monitor CPU and GPU utilization in heterogeneous clusters, to track power consumption and thermal limits per node, and to generate power profiles for system components; a multi-criteria optimization engine configured to: perform a Pareto-optimal scheduling analysis that balances energy consumption, throughput time, and reliability; apply statistical and AI-powered trade-off analyses and ensure optimal resource allocation based on Pareto frontier analysis;a dynamic resource allocation mechanism configured to: use predictive models that incorporate LLM inferences and reinforcement learning feedback; distribute tasks across nodes and clusters while minimizing power consumption; and improve system throughput based on the predictive models; a performance optimization module configured to optimize scheduling decisions using multi-criteria optimization analysis; and a user interface that allows human operators to override and refine scheduling strategies in real time based on traceable scheduling rationales.
[0007] The aim of this disclosure is to provide a system for energy-conscious workflow planning based on LLM with dynamic resource allocation.
[0008] Another objective of the present disclosure is the dynamic balance between energy consumption, throughput time and system reliability through multi-criteria Pareto-optimal analysis, taking into account runtime fluctuations.
[0009] Another objective of the present disclosure is to enable transparent scheduling decisions through explanations in natural language and real-time intervention options by human operators, thereby allowing for a refinement of guidelines and objective prioritization.
[0010] Another objective of the present disclosure is the implementation of a closed-loop rule system that continuously learns and optimizes task-resource assignments by combining feedback from reinforcement learning with predictive modeling to achieve improved throughput and higher energy efficiency.
[0011] To further clarify the advantages and features of the present disclosure, the invention is described in more detail with reference to specific embodiments illustrated in the accompanying drawings. It is understood that these drawings merely show typical embodiments of the invention and are therefore not to be understood as limiting its scope of protection. The invention is described and explained in more detail and with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE IMAGES
[0012] These and other features, aspects and advantages of the present disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings, in which identical symbols represent identical parts, wherein: Fig. Figure 1 shows a block diagram of a system for energy-conscious, LLM-based workflow planning with dynamic resource allocation according to an embodiment of the present disclosure; Fig. Figure 2 shows a block diagram illustrating the energy-conscious, LLM-supported workflow planning process according to one embodiment of the present disclosure; and Fig. Figure 3 shows a block diagram of an energy-conscious, LLM-based workflow planning system that uses a decision engine based on a Large Language Model (LLM) in combination with a Model Context Protocol (MCP) layer and zero-shot reasoning according to an embodiment of the present disclosure.
[0013] Furthermore, those skilled in the art will recognize that the elements in the drawings are simplified and not necessarily drawn to scale. For example, the flowcharts illustrate the process by highlighting the main steps to facilitate understanding of this disclosure. With regard to the construction of the device, one or more components may be represented in the drawings by conventional symbols. The drawings may show only those specific details relevant to understanding the embodiments of this disclosure, so as not to clutter the drawings with details that are already apparent to those skilled in the art from the description contained herein. DETAILED DESCRIPTION:
[0014] To facilitate understanding of the principles of the invention, reference is made below to the embodiment illustrated in the drawings, which is described using specific terms. It is understood, however, that this does not limit the scope of protection of the invention. Rather, modifications and further developments of the illustrated system, as well as further applications of the inventive principles depicted therein, are conceivable, insofar as they would typically occur to a person skilled in the art in the field of the invention.
[0015] It will be clear to those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not to be understood as a limitation of it.
[0016] References to “an aspect”, “another aspect”, or similar phrases in this description mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present disclosure. Therefore, phrases such as “in one embodiment”, “in another embodiment”, and similar expressions in this description may, but do not necessarily, all refer to the same embodiment.
[0017] The terms "includes," "comprehensive," or similar expressions denote non-exclusive inclusion. Thus, a procedure or method containing a list of steps does not only include those steps but may also include further steps not explicitly listed or inherent in the procedure or method. Likewise, the statement "includes..." for one or more devices, subsystems, elements, structures, or components, without further limitations, does not preclude the existence of other devices, subsystems, elements, structures, or components.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meanings generally known to those skilled in the art in the field to which this invention belongs. The systems, methods, and examples described herein serve only for illustration and are not to be understood as limiting.
[0019] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0020] The functional units described in this specification are referred to as devices. A device may be implemented in programmable hardware such as processors, digital signal processors, central processing units, FPGAs, PALs, PLDs, cloud processing systems, or similar. Devices may also be implemented in software for execution by various processor types. An identified device may contain executable code and, for example, comprise one or more physical or logical blocks of computer instructions, which may be organized as an object, procedure, function, or other construct. However, the executable files of an identified device need not be physically related; they may consist of different instructions stored in different locations that, when logically combined, constitute the device and fulfill its purpose.
[0021] The executable code of a device or module can consist of a single instruction or multiple instructions and can even extend across different code sections, applications, and storage media. Similarly, operational data within the device can be identified and represented, and can exist in any suitable form and be organized in any data structure. The operational data can be captured as a single data record or distributed across various storage media and may exist, at least partially, as electronic signals within a system or network.
[0022] References to “a selected embodiment”, “an embodiment”, or “an embodiment” in this description mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the disclosed subject matter. Therefore, the phrases “a selected embodiment”, “in an embodiment”, or “in an embodiment” appearing at different points in this description do not necessarily refer to the same embodiment.
[0023] Furthermore, the described features, structures, or properties can be combined in one or more embodiments in any suitable manner. The following description contains numerous specific details to enable a comprehensive understanding of the embodiments of the disclosed subject matter. However, a person skilled in the art will recognize that the disclosed subject matter can also be realized without one or more of the specific details or with other methods, components, materials, etc. In other cases, known structures, materials, or processes are not presented or described in detail so as not to obscure aspects of the disclosed subject matter.
[0024] According to the exemplary embodiments, the disclosed computer programs or modules can be executed in a variety of ways, for example, as an application running in the memory of a device or as a hosted application running on a server and communicating with the device application or browser via various standard protocols such as TCP / IP, HTTP, XML, SOAP, REST, JSON, and other suitable protocols. The disclosed computer programs can be written in programming languages that run either in the device's memory or on a hosted server, such as BASIC, COBOL, C, C++, Java, Pascal, or scripting languages such as JavaScript, Python, Ruby, PHP, Perl, or other suitable programming languages.
[0025] Some of the described embodiments involve data transmission over a network, such as the transmission of various inputs or files. The network may include, for example, the internet, wide area networks (WANs), local area networks (LANs), analog or digital wired and wireless telephone networks (e.g., PSTN, ISDN, cellular networks, and xDSL), radio, television, cable, satellite, and / or other transmission or tunneling mechanisms for data. It may include multiple networks or subnetworks, each of which may, for example, have a wired or wireless data path. The network may include a circuit-switched voice network, a packet-switched data network, or another network for transmitting electronic data. For example, it may be based on the Internet Protocol (IP) or Asynchronous Transfer Mode (ATM) and support voice communication using VoIP, Voice over ATM, or similar protocols.In one embodiment, the network comprises a mobile network configured for the exchange of text or SMS messages.
[0026] Examples of networks include Personal Area Networks (PAN), Storage Area Networks (SAN), Home Area Networks (HAN), Campus Area Networks (CAN), Local Area Networks (LAN), Wide Area Networks (WAN), Metropolitan Area Networks (MAN), Virtual Private Networks (VPN), Enterprise Private Networks (EPN), the Internet, Global Area Networks (GAN), and so on. Fig. Figure 1 shows a block diagram of a system (100) for energy-conscious workflow planning based on LLM with dynamic resource allocation according to an embodiment of the present disclosure.
[0027] According to Fig. 1 The system comprises: a workflow input interface (102) configured to receive workflow-directed acyclic graphs (DAGs), energy budget constraints, system performance constraints, and natural language queries from human operators; a large language model (LLM) reasoning module (104) connected to the workflow input interface (102) that provides the following functions: to analyze the workflow specifications and system constraints in natural language, to generate energy-aware scheduling recommendations based on the analyzed workflow specifications, and to provide explainable scheduling justifications in natural language; a reinforcement learning-based scheduling unit (106) connected to the LLM reasoning module (104) that provides the following functions: to receive the scheduling recommendations from the LLM reasoning agent,The fine-tuning of task-resource allocations through dynamic adaptation to runtime fluctuations and online resource redistribution under runtime variability; an energy monitoring unit (108) is configured for continuous monitoring of CPU and GPU utilization of heterogeneous clusters, tracking of power consumption and thermal limits per node, and creation of energy profiles for system components. A multi-objective optimization engine (110) is configured for performing a Pareto-optimal scheduling analysis considering energy consumption, throughput time, and reliability, applying statistical and AI-supported trade-off analyses, and ensuring optimal resource allocation based on Pareto frontier analysis. A dynamic resource allocation unit (112) is configured for using predictive models with LLM logic and reinforcement learning feedback.The reassignment of tasks between nodes and clusters while minimizing energy consumption and improving system throughput based on predictive models. A performance optimization module (114) is configured to optimize scheduling decisions using multi-objective optimization analysis. A user interface (116) allows human operators to override and refine scheduling strategies in real time based on verifiable scheduling reasons.
[0028] In one embodiment, the LLM inference module (104) comprises: a natural language processing module (104a) configured to interpret workflow goals and system constraints; a planning strategy generation module (104b) configured to translate the interpreted goals into specific planning strategies; and an explanation generation module (104c) configured to provide natural language justifications for recommended planning decisions.
[0029] In one embodiment, the reinforcement learning-based scheduling unit (106) comprises: a task resource mapping module (106a) configured to dynamically adapt task assignments to available resources; a runtime adaptation module (106b) configured to respond to system fluctuations and workload variations; and a learning module (106c) configured to continuously improve scheduling policies based on historical performance data and energy consumption patterns.
[0030] In one embodiment, the energy monitoring unit (108) comprises: a power measurement module (108a) for real-time acquisition of the power consumption of individual nodes; a thermal monitoring module (108b) for monitoring temperature thresholds and thermal limits; and a utilization tracking module (108c) for measuring the CPU and GPU utilization rates in the heterogeneous clusters.
[0031] In one embodiment, the multi-objective optimization engine (110) is configured to simultaneously optimize throughput time, energy consumption and system reliability, generate Pareto-optimal solutions that represent trade-offs between the optimization objectives, and select optimal scheduling configurations based on weighted priority factors defined by the human operators.
[0032] In one embodiment, the dynamic resource allocation unit (112) further comprises: a predictive modeling module (112a) that combines LLM inference outputs with feedback from reinforcement learning; a task migration module (112b) configured to move tasks between nodes to optimize energy efficiency; and a load balancing module (112c) configured to distribute workloads across available resources while maintaining performance targets.
[0033] In one embodiment, the system (100) further comprises: a fault tolerance monitoring module (118) configured to detect and respond to system errors and malware threats; and a security management module (120) configured to integrate fault correction and malware resilience mechanisms into the planning policy.
[0034] In one embodiment, the user interface (118) is configured to: interpret operator requests and policy changes; allow operators to modify scheduling decisions during execution; and prioritize certain objectives, such as the use of green energy or execution deadlines, based on operator instructions.
[0035] In one embodiment, the system (100) further comprises an execution module (122) configured to distribute tasks to assigned resources based on the optimized planning decisions, monitor runtime performance in the heterogeneous clusters, and provide feedback to the reinforcement learning-based scheduler for continuous improvement.
[0036] In one embodiment, the system (100) is further configured to: process workflow DAGs with associated energy budgets, performance deadlines, and priority levels; generate initial scheduling policies using the LLM argumentation agent; continuously adjust the scheduling policies using the reinforcement learning-based scheduler based on real-time system feedback; and maintain optimal energy-performance trade-offs through the dynamic resource allocation mechanism and the multi-objective optimization engine.
[0037] In one embodiment, the workflow input interface (102), the LLM inference module (104), the reinforcement learning-based planning unit (106), the energy monitoring unit (108), the multi-objective optimization engine (110), the dynamic resource allocation unit (112), the power optimization module (114), and the user interface (116) can be implemented in programmable hardware devices such as processors, digital signal processors, central processing units, field programmable gate arrays, programmable array logic, programmable logic devices, cloud processing systems, or the like.
[0038] The present invention provides a workflow planning system that uses LLM inference to analyze workflow specifications and generate energy-efficient planning recommendations. Simultaneously, a reinforcement learning-based scheduler optimizes the allocation of tasks and resources based on runtime conditions. The system continuously monitors energy consumption and system performance using dedicated monitoring units, applies multi-criteria optimization to make Pareto-optimal planning decisions, and provides operators with natural language explanations that allow them to adjust or refine planning strategies in real time. This results in an intelligent, adaptive, and transparent planning framework for heterogeneous computing environments.
[0039] Fig. Figure 2 shows a block diagram illustrating the energy-conscious, LLM-supported workflow planning process according to one embodiment of the present disclosure.
[0040] As in Fig. As shown in Figure 2, the hybrid, LLM-based workflow scheduling system enables energy-conscious and performance-optimized task execution in heterogeneous clusters. The system integrates Large Language Model Reasoning, reinforcement learning-based scheduling, energy monitoring, dynamic resource allocation, multi-criteria optimization, and human-machine interaction into a unified architecture. This integration allows the system to intelligently adapt scheduling decisions to varying runtime conditions and provide operators with comprehensible justifications.
[0041] Fig. As shown in Figure 2, the system architecture consists of an input layer that receives workflow DAGs, system constraints such as energy budgets, deadlines, and priority levels, as well as natural language input from human operators. The LLM logic layer processes this input, interpreting objectives and generating scheduling strategies with explanations. These strategies are refined by the dynamic scheduler, which is based on reinforcement learning and integrates malware and fault tolerance mechanisms. The system also includes a resource and energy monitor that tracks utilization, power consumption, and thermal conditions across clusters. Scheduling decisions are optimized by the optimization engine, which performs a Pareto-based multi-objective analysis.An execution layer distributes tasks according to optimized scheduling guidelines, monitors runtime performance, and provides continuous feedback for reinforcement learning-based improvements. Through this integration, the system represents a novel framework that uniquely combines LLM-based logic, reinforcement learning-based adaptation, and energy-conscious multi-objective optimization, while enabling transparent and traceable interaction with human operators.
[0042] In one embodiment, the system includes an LLM inference agent configured to analyze workflow specifications, energy budgets, and HPC system constraints in natural language. Based on this interpretation, the LLM inference agent generates optimized scheduling strategies and simultaneously provides natural language justifications to explain the scheduling recommendations. These explanations enable human operators to understand and evaluate system decisions. The system further includes a reinforcement learning-based scheduler that optimizes task-to-resource allocation by dynamically adapting to runtime variations. This scheduler continuously adjusts task assignments to available resources, responds to load fluctuations, and improves scheduling policies over time using feedback from historical performance data and real-time execution results.The system also includes an energy-aware resource manager that continuously monitors CPU and GPU utilization, thermal thresholds, and power consumption per node in heterogeneous clusters. This monitoring capability allows the system to maintain updated energy profiles that guide scheduling and resource allocation decisions. To improve energy efficiency and throughput, the system integrates a dynamic resource allocation mechanism. This mechanism utilizes predictive models that combine the outputs of the LLM logic system and reinforcement learning feedback. By leveraging these models, the system can distribute tasks across nodes and clusters, reducing power consumption and ensuring balanced utilization while optimizing performance targets. The system also includes a performance optimization layer that performs statistical and AI-powered trade-off analyses.This layer optimizes multiple objectives based on Pareto front analysis, ensuring that planning decisions strike a balanced compromise between throughput time, energy efficiency, and reliability. By selecting Pareto-optimal configurations, the system can implement energy-efficient yet performance-oriented planning strategies. A human-machine interaction module is also integrated. This module provides understandable planning rationales in natural language, generated by the LLM (Logistics Management Module). This allows operators to adjust or override the system's planning strategies in real time and enforce specific priorities, such as the use of green energy or adherence to strict deadlines. This ensures that operator-defined guidelines can be directly integrated into the automated planning system.
[0043] Fig. Figure 3 shows a block diagram of an energy-conscious, LLM-based workflow planning system that uses a decision engine based on a Large Language Model (LLM) in combination with a Model Context Protocol (MCP) layer and zero-shot reasoning according to an embodiment of the present disclosure.
[0044] As in Fig. As shown in Figure 3, the system is configured for energy-conscious workflow planning in heterogeneous computing environments. It uses a decision engine based on a Large Language Model (LLM) in combination with a Model Context Protocol (MCP) layer and zero-shot reasoning. The system dynamically distributes computing resources across clusters, nodes, and accelerators to minimize energy consumption while ensuring performance and reliability.
[0045] Fig.As shown in Figure 3, the system comprises a computing architecture with multiple heterogeneous nodes, including CPUs, GPUs, and domain-specific accelerators equipped with energy and power telemetry. The system also includes an MCP layer, which acts as a middleware protocol, aggregating real-time telemetry from the computing infrastructure, normalizing control interfaces, and providing this data as structured content for the scheduling engine. Furthermore, the system includes an LLM-based zero-shot scheduler. This is a large-language model that, using the context provided by the MCP layer, derives task placement, resource throttling, or scaling decisions for unknown workflows without explicit retraining. The system also includes a feedback channel mechanism for measuring actual energy consumption and performance, which feeds this data back to the MCP layer to optimize future scheduling actions.
[0046] In the implementation, the MCP layer acts as a context bridge between the physical infrastructure and the LLM-based scheduler. It collects telemetry data on power consumption, thermal state, and node utilization, and dynamically detects and registers new compute resources and control APIs. The MCP layer translates heterogeneous signals into standardized schemes that can be processed by the LLM and implements the scheduling decisions returned by the LLM to the underlying resource managers. The zero-shot component of the LLM scheduler enables generalizable, task-independent reasoning. It can interpret new workflow graphs, constraints, or SLAs without requiring pre-training for each workflow type. The zero-shot algorithm derives energy-optimal scheduling decisions directly from structured prompts generated by the MCP layer. It generates structured action tokens, such as...Task-to-node assignments and frequency scaling instructions can be executed immediately. This reduces training overhead and enables adaptive behavior in dynamic multi-tenant environments. By combining the MCP's real-time telemetry with zero-shot LLM inference, the system is capable of energy-aware optimization. It identifies underutilized nodes and reassigns tasks to minimize idle power consumption, selects energy-saving configurations for non-critical tasks, and predicts the energy impact of scheduling decisions before they are executed. The scheduler also performs continuous, closed-loop control. When the workload or node state changes, the MCP updates the context, and the LLM re-evaluates the placement decisions.This enables rapid resource scaling based on workload intensity, the migration of tasks to more energy-efficient nodes, and the integration of external energy policies or cost-conscious constraints. Since the LLM issues control actions, the MCP includes validation layers to verify the safety and compliance of the generated actions before execution. This ensures that changes to resource allocation do not violate SLAs or destabilize the system, thus increasing safety and reliability.
[0047] In one implementation of the energy-aware, LLM-based workflow scheduling system with dynamic resource allocation, two core components are the Model Context Protocol (MCP) interface and the zero-shot reasoning capability of the large language model. MCP provides a standardized, secure channel for the scheduler's LLM engine to access structured, real-time information about the computing environment. Through MCP clients and servers, the LLM retrieves real-time telemetry data on CPU / GPU utilization, node power consumption, and energy budgets; schemas and control endpoints for resource managers, autoscalers, and power controllers; as well as workflow descriptors and SLA / policy data.This protocol eliminates the need for custom connectors, minimizes prompt length by transmitting only relevant, summarized contextual information, and allows the LLM to send validated control actions back to the infrastructure. MCP also enables proactive tool discovery, allowing the LLM to identify new controllers or metrics without retraining.
[0048] In this implementation, the LLM decision engine is instructed to generate scheduling and allocation actions in zero-shot mode. This means it does not rely on workflow-specific, labeled examples or pre-trained task policies. Instead, based solely on the context provided by the MCP and a formal optimization goal—such as minimizing energy consumption while adhering to latency constraints—the model derives valid placement, throttling, or scaling decisions for previously unknown workflows and topologies. This zero-shot capability allows the scheduler to generalize for heterogeneous resources, new workload types, and changing policy constraints while maintaining energy efficiency and SLA compliance. The MCP and zero-shot reasoning together form a feedback loop.MCP provides the LLM with structured, up-to-date context; the zero-shot LLM generates structured action tokens; a verifier executes these via MCP and returns measured energy / power values for continuous model adjustment. This combination results in a flexible, future-proof scheduler that is both energy-aware and domain-independent.
[0049] The drawings and the preceding description illustrate embodiments. Those skilled in the art will recognize that one or more of the described elements can be combined to form a single functional element. Alternatively, certain elements can be divided into several functional elements. Elements of one embodiment can be added to another. For example, the process sequences described here can be modified and are not limited to the manner described herein. Furthermore, the actions of a flowchart need not be performed in the sequence shown; nor do all actions necessarily need to be carried out. Actions that do not depend on other actions can be performed in parallel with the other actions. The scope of protection of the embodiments is in no way limited by these specific examples. Numerous variations, whether explicitly stated in the description or not, such as...Differences in structure, dimensions, and materials are possible. The scope of protection of the embodiments is at least as comprehensive as described by the following claims.
[0050] The advantages, other benefits, and problem solutions have been described above with reference to specific embodiments. However, the advantages, benefits, problem solutions, and any components that can effect or enhance an advantage, benefit, or solution are not to be construed as critical, necessary, or essential features or components of the claims. REFERENCES 100 A System For Energy-Conscious LLM-Based Workflow Planning With Dynamic Resource Allocation. 102 Workflow input interface 104 Module for Closing Large Language Models (LLM) 104a Module for Natural Language Processing 104b Module for Generating Scheduling Strategies 104c Explanation Generation Module 106 Planning Unit Based on Reinforcement Learning 106a Module for Task-Resource Allocation 106b Runtime Adjustment Module 106c Learning Module 108 Energy monitoring unit 108a Power measurement module 108b Thermal monitoring module 108c Module for Usage Tracking 110 Multi-goal optimization engine 112 Dynamic Resource Allocation Unit 112a Module for Predictive Modeling 112b Task Migration Module 112c Load distribution module 114 Performance Optimization Module 116 User interface 118 Module for Monitoring Fail-Safety 120 Security Management Module 122 Execution module 202 LLM Argumentation 204 Dynamic Scheduler 206 Optimization Engine 208 Resource and Energy Monitoring 210 Workflow input 212 users 302 Computing infrastructure 304 CPU 306 GPU 308 Accelerators 310 MCP layer 312 LLM-Based Zero Shot Planner 314 feedback 316 Plurality of the Node
Claims
[1] An energy-conscious workflow planning system based on LLM with dynamic resource allocation, consisting of: a workflow input interface configured to receive workflow-directed acyclic graphs (DAGs), energy budget constraints, system performance constraints, and natural language requests from human operators; a large language model (LLM) logic module connected to the workflow input interface and configured to analyze the workflow specifications and system constraints in natural language, generate energy-conscious planning recommendations based on the analyzed workflow specifications, and provide explainable planning rationales in natural language; a reinforcement learning-based scheduling unit connected to the LLM reasoning module and configured to: receive scheduling recommendations from the LLM reasoning agent, fine-tune task-resource assignments by dynamically adapting to runtime variations, and perform online resource redistribution under runtime variability; an energy monitoring unit configured to: continuously monitor CPU and GPU utilization in heterogeneous clusters, track power consumption and thermal limits per node, and generate energy profiles for system components; a multi-objective optimization engine configured to: perform a Pareto-optimal scheduling analysis that balances energy consumption, lead time and reliability, apply statistical and AI-supported trade-off analyses and ensure optimal resource allocation based on Pareto frontier analysis; a dynamic resource allocation unit configured to: use predictive models that incorporate LLM inferences and feedback from reinforcement learning, reassign tasks between nodes and clusters while minimizing energy consumption and improving system throughput based on the predictive models; a performance optimization module configured to optimize scheduling decisions using multi-criteria optimization analysis; and a user interface that allows human operators to override and refine planning strategies in real time based on verifiable planning reasons. [2] System according to claim 1, wherein the LLM inference module comprises: a natural language processing module configured to interpret workflow goals and system constraints; a module for generating scheduling strategies, configured to translate the interpreted goals into specific scheduling strategies; and A module for generating explanations, configured to provide reasons for recommended scheduling decisions in natural language. [3] System according to claim 1, wherein the reinforcement learning-based planning unit comprises: a task-resource mapping module configured to dynamically adjust task assignments to available resources; a runtime adaptation module configured to respond to system fluctuations and workload variations; and a learning module configured to continuously improve planning guidelines based on historical performance data and energy consumption patterns. [4] System according to claim 1, wherein the energy monitoring unit comprises: a power measurement module configured to record the power consumption of individual nodes in real time; a thermal monitoring module configured to monitor temperature thresholds and thermal limits; and a utilization tracking module configured to measure CPU and GPU utilization rates in heterogeneous clusters. [5] System according to claim 1, wherein the multi-objective optimization engine is configured to simultaneously optimize throughput time, energy consumption and system reliability, generate Pareto-optimal solutions that represent trade-offs between the optimization objectives, and select optimal scheduling configurations based on weighted priority factors defined by the human operators. [6] System according to claim 1, wherein the dynamic resource allocation unit further comprises: a predictive modeling module that combines LLM inferences with feedback from reinforcement learning; a task migration module configured to move tasks between nodes to optimize energy efficiency; and a load balancing module configured to distribute workloads across available resources while meeting performance targets. [7] System according to claim 1, further comprising: a fault tolerance monitoring module configured to detect and respond to system errors and malware threats; and a security management module configured to integrate fault and malware resilience mechanisms into the planning policy. [8] System according to claim 1, wherein the user interface is configured such that it: It interprets operator requests and policy changes; allows operators to change scheduling decisions during execution; and prioritizes certain objectives such as the use of green energy or execution deadlines based on operator instructions. [9] System according to claim 1, further comprising: an execution module configured to distribute tasks to allocated resources based on the optimized planning decisions, to monitor runtime performance in the heterogeneous clusters, and to provide feedback to the reinforcement learning-based scheduler for continuous improvement. [10] System according to claim 1, wherein the system is configured to: process workflow DAGs with associated energy budgets, performance deadlines and priority levels; generate initial scheduling policies using the LLM reasoning agent; continuously adjust the scheduling policies using the reinforcement learning-based scheduler based on real-time system feedback; and maintain optimal energy-performance trade-offs through the dynamic resource allocation mechanism and the multi-objective optimization engine.
Citation Information
Cited By
Discrete task-oriented resource scheduling optimization management system
CN122114564A
Micro-grid energy management and control system based on large language model and deep reinforcement learning
CN122203443A