Big data computing power analysis application scene management platform based on artificial intelligence
By constructing an intelligent computing power governance system, we have achieved refined and adaptive management of heterogeneous computing power environments, solved the problems of extensive resource allocation and lagging task scheduling in existing technologies, improved resource utilization and system stability, and reduced energy consumption and carbon emissions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN DONGHUANG INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for managing computing resources in artificial intelligence and big data suffer from problems such as extensive allocation of computing resources, lagging task scheduling, poor adaptability to application scenarios, and low overall system energy efficiency. In particular, they lack effective resource allocation and adaptive scheduling mechanisms in high-dimensional dynamic environments.
Construct a closed-loop intelligent computing power governance system that integrates perception, cognition, decision-making and execution. Through dynamic knowledge graphs and adaptive resource allocation using reinforcement learning, combined with a task flow control mechanism based on causal inference, achieve unified representation and precise matching of heterogeneous computing resources and diverse AI workloads.
It significantly improves resource utilization in heterogeneous computing environments, reduces tail latency of high-priority tasks, reduces unplanned downtime events caused by local overheating, improves the accuracy of task completion time prediction, and reduces carbon emission intensity and operating costs.
Smart Images

Figure CN121935006A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and big data processing technology, specifically relating to a big data computing power analysis application scenario management platform based on artificial intelligence. Background Technology
[0002] The deep integration of artificial intelligence and big data technologies is driving a profound transformation in computing resource management paradigms. Against the backdrop of accelerated digital transformation, computing power, as the core infrastructure supporting model training, data analysis, and real-time inference, directly determines the responsiveness and operational efficiency of intelligent systems through its scheduling efficiency and application adaptability. With the continuous growth in the scale of deep learning models and the diversification of business scenarios, the demand for computing resources exhibits explosive, uneven, and dynamically fluctuating characteristics. Traditional static allocation and isolated management models are no longer sufficient to meet the performance requirements of complex application scenarios.
[0003] Among these, the management of big data computing power analysis application scenarios for artificial intelligence involves the integration of multi-source heterogeneous data, intelligent orchestration of computing tasks, and dynamic perception of resource usage status. The core objective of this direction is to achieve precise matching between computing resources and business loads, improving resource utilization and ensuring the service quality of critical tasks by building a unified management view and scheduling logic. However, existing technologies exhibit systemic deficiencies when dealing with high-dimensional dynamic environments, particularly lacking effective mechanisms for cross-scenario resource allocation, real-time load prediction, and adaptive scheduling.
[0004] Existing technologies primarily rely on preset rules or simple threshold triggers for resource allocation, failing to accurately capture the time-series characteristics and correlation patterns of computing power demands in different application scenarios. The utilization of historical operational data remains at the statistical level, lacking the ability for deep feature mining and trend inference based on machine learning. Furthermore, platforms typically handle computing power monitoring, task scheduling, and scenario configuration in a fragmented manner, resulting in significant delays in decision-making and hindering feedforward resource reservation and dynamic optimization. In addition, in multi-tenant, multi-task concurrent environments, existing systems fail to adequately model the coupling relationships between priority conflicts, service quality differences, and energy efficiency constraints, leading to both resource contention and idleness. Therefore, there is an urgent need for a big data computing power analysis application scenario management platform that integrates artificial intelligence methods and possesses global situational awareness and autonomous decision-making capabilities to solve these technical challenges. Summary of the Invention
[0005] The purpose of this invention is to provide a big data computing power analysis application scenario management platform based on artificial intelligence, to solve the technical problems existing in the current integration of big data and artificial intelligence applications, such as extensive allocation of computing power resources, lagging task scheduling, poor adaptability to application scenarios, and low overall system energy efficiency. With the continuous increase in the complexity of artificial intelligence models and the exponential growth of data processing scale, traditional static partitioning and rule-driven computing power management mechanisms can no longer meet the diverse, dynamic, and high-concurrency application needs. Existing technologies generally adopt unified queue scheduling or multi-tenant isolation strategies, lacking the ability to deeply couple and model the essential characteristics of tasks with the characteristics of computing power hardware. This leads to serious resource waste in high-end GPU clusters when executing lightweight inference tasks, while high-load training tasks frequently experience blocking due to memory bandwidth bottlenecks. Furthermore, existing platforms have failed to establish a multi-objective collaborative optimization mechanism between task priority, service quality requirements, energy consumption costs, and hardware health status, resulting in high system operation and maintenance costs and insufficient sustainability.
[0006] The technical solution of this invention is to construct a closed-loop intelligent computing power governance system integrating perception, cognition, decision-making, and execution. This system achieves a unified representation of heterogeneous computing resources and diverse AI workloads by establishing a cross-level dynamic knowledge graph, and implements an adaptive resource allocation based on reinforcement learning and a task flow control mechanism based on causal inference. A unified access gateway is deployed at the system front end to receive AI task requests from different business systems. Each task request carries metadata such as task type identifier, data volume, expected completion time, accuracy requirements, and energy consumption budget. The access gateway forwards the original request to a context parsing engine, which uses a pre-trained semantic understanding model to extract the deep intent of the task and maps it to a standardized workload description vector. After all pending tasks enter the global task pool, a task profiling module performs fine-grained feature characterization, including computational density, memory access mode, communication topology, fault tolerance level, and historical execution trajectory statistical characteristics.
[0007] Furthermore, the system is equipped with a resource status awareness layer, which collects hardware operating parameters of various computing nodes in real time through distributed probes, covering CPU / GPU utilization, memory usage, NVLink bandwidth usage, temperature gradient distribution, power efficiency (PUE) value, and solid-state storage IOPS fluctuations. These low-level indicators, after feature normalization and spatiotemporal alignment, are input to the resource topology modeling unit. This unit uses a graph neural network to construct a dynamically updated physical resource connection graph, where nodes represent computing units or storage devices, and edge weights reflect actual communication latency and bandwidth capacity. Simultaneously, the system incorporates a computing power semantic space mapper, embedding task profile vectors and resource topology graphs into a unified high-dimensional semantic space. A metric learning algorithm is used to calculate the task-resource matching score, serving as the basis for preliminary scheduling decisions.
[0008] Furthermore, the system core includes an intelligent scheduling decision center, comprising two parallel sub-modules: a multi-objective optimization scheduler and a risk avoidance controller. The multi-objective optimization scheduler employs a hierarchical reinforcement learning architecture. The upper-layer policy network is responsible for macro-level resource partitioning decisions, determining whether tasks should be allocated to high-performance computing zones, energy-efficiency priority zones, or hybrid elastic zones. The lower-layer action network executes specific node binding operations within the selected zones. The reward function is designed as a weighted combination, comprehensively considering four indicators: average task response time, overall cluster energy efficiency ratio, critical node overheating frequency, and Service Level Agreement (SLA) achievement rate. A dynamic weight adjustment mechanism is introduced to respond to changes in operational priorities at different times. The risk avoidance controller operates independently outside the scheduling process, employing a risk prediction model based on Bayesian causal networks to assess the probability of cascading failures that a specific scheduling scheme may trigger. Examples include the impact path of a node temperature surge on the reliability of neighboring devices, or the impact effect of a large-scale AllReduce operation on the network switch buffer. When the predicted risk value exceeds a preset safety threshold, the controller generates a suppression signal and feeds it back to the scheduler, forcibly replanning the resource allocation path.
[0009] Furthermore, the system implements runtime dynamic optimization capabilities. Lightweight agents deployed on each computing node collect performance inversion data during actual task execution, including microarchitecture-level metrics such as kernel function execution cycle, cache hit deviation, and PCIe transmission efficiency loss. This feedback data is aggregated into the online learning engine to continuously update the task execution time prediction model and resource consumption estimation function, forming a closed-loop calibration mechanism from planning to verification to correction. For long-running continuous learning tasks, the system employs an incremental migration mechanism. While ensuring model convergence, it dynamically adjusts the batch size, gradient synchronization frequency, and checkpoint saving interval based on real-time resource conditions, thereby maintaining an optimal balance between performance and stability.
[0010] Furthermore, the system integrates a green computing management unit, which is directly connected to the data center infrastructure management system (BMS) to obtain rack-level cooling air temperature, CRAC unit output status, and grid electricity price fluctuation curves. Combined with a chip-level thermodynamic simulation model, the system can predict the temperature rise trend of hotspot areas within the next 15 minutes and trigger fan speed adjustments or task migration plans in advance. Under the time-of-use pricing mechanism, batch processing tasks will be directed to be executed during off-peak hours, while simultaneously activating energy storage devices to reduce peak grid load. The system also supports carbon footprint tracking, calculating and recording the environmental impact index throughout the entire lifecycle based on the actual power consumption of each task and the real-time carbon emission factor of the local grid.
[0011] Preferably, the semantic understanding model adopts a Transformer architecture, is pre-trained on a constructed large-scale task description text corpus, and enhances its ability to distinguish similar task types through contrastive learning. Preferably, the graph neural network employs a heterogeneous graph attention mechanism, which can differentiate the various interaction relationships between computing nodes, storage units, and network interfaces, and automatically learn the importance coefficients of different types of edges. Preferably, the metric learning algorithm is trained using a triplet loss function to ensure that the distance between similar task-resource combinations in the embedding space is less than the distance between different combinations by at least one boundary margin. Preferably, in the hierarchical reinforcement learning architecture, the decision cycle of the upper-layer policy network is 30 seconds, and the decision cycle of the lower-layer action network is 2 seconds; the two achieve information transmission by sharing some hidden states. Preferably, the Bayesian causal network includes an explicitly encoded library of physical failure modes, such as thermal fatigue accumulation effects and voltage noise reduction ratio deterioration paths, and achieves efficient posterior probability estimation through variational inference algorithms. Preferably, the online learning engine uses a sliding time window mechanism to maintain the execution logs for the most recent 7 days and uses a random forest regressor for nonlinear residual modeling to compensate for the systematic bias of the basic prediction model. Preferably, the incremental migration mechanism follows a monotonic constraint principle when adjusting hyperparameters, meaning the batch size is only allowed to increase or remain unchanged, avoiding training divergence due to frequent jitter. Preferably, the thermodynamic simulation model solves the three-dimensional heat conduction equation based on the finite difference method, with a mesh resolution set to 2 mm accuracy at the chip package level, and performs Kalman filtering correction every 5 minutes based on the measured temperature.
[0012] Compared with the prior art, the advantages and positive effects of the present invention are as follows: This solution achieves a paradigm shift from "resource-adapted tasks" to "semantic matching and collaboration" by constructing a joint semantic representation space for tasks and resources. This significantly improves resource utilization in heterogeneous computing environments. Experimental data shows that the average GPU utilization rate increased from the traditional 48% to 79%, while reducing tail latency of high-priority tasks by 63%. The solution introduces a dual-track decision-making mechanism of hierarchical reinforcement learning and causal risk control. While pursuing maximum performance, it proactively avoids potential systemic risks, reducing unplanned downtime events caused by localized overheating by 82%, effectively ensuring the smooth operation of large-scale clusters. Long-term stable operation; This solution establishes an end-to-end closed-loop feedback optimization system, continuously correcting cognitive biases in the scheduling model by absorbing real execution data, enabling the accuracy of task completion time prediction to converge from the initial 71% to 94% within two weeks, significantly enhancing the predictability of system behavior; This solution deeply integrates the concept of green computing, incorporating energy consumption, heat dissipation, and grid load into a unified optimization framework, achieving a 37% reduction in average carbon emission intensity per task and a 22% reduction in total annual electricity expenditure while meeting performance targets, providing a practical and feasible technical path for building sustainable AI infrastructure. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the overall technical solution architecture of the big data computing power analysis application scenario management platform based on artificial intelligence proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of the adaptive resource allocation and causal inference task flow control mechanism in this invention. Detailed Implementation
[0014] Please refer to Figure 1 and Figure 2 This invention provides a big data computing power analysis application scenario management platform based on artificial intelligence, comprising: A unified access gateway is used to receive AI task requests from different business systems. Each task request carries metadata such as task type identifier, data volume, expected completion time, accuracy requirements, and energy consumption budget. The context parsing engine is used to perform semantic parsing on the raw task requests forwarded by the access gateway, extract the deep intent of the task, and map it into a standardized workload description vector. The task profile generation module is used to perform fine-grained feature characterization on the tasks to be processed that enter the global task pool, and generate a multi-dimensional task profile that includes computing density, memory access mode, communication topology, fault tolerance level and historical execution trajectory statistics. The resource status awareness layer is used to collect hardware operating parameters of various computing nodes in real time through distributed probes, covering CPU / GPU utilization, video memory usage, NVLink bandwidth usage, temperature gradient distribution, power efficiency (PUE) value, and solid-state storage IOPS fluctuation. The resource topology modeling unit is used to construct a dynamically updated physical resource connection graph after the low-level indicators collected by the resource status perception layer are processed by feature normalization and spatiotemporal alignment. The computational power semantic space mapper is used to embed task profile vectors and resource topology graphs into a unified high-dimensional semantic space, and calculate task-resource matching score through metric learning algorithm; The intelligent scheduling decision center is used to generate the final resource allocation scheme based on task-resource matching score and the collaborative output of multi-objective optimization scheduler and risk avoidance controller. The multi-objective optimization scheduler adopts a hierarchical reinforcement learning architecture. The upper-layer policy network is responsible for macro-level resource partitioning decisions, while the lower-layer action network executes specific node binding operations within the selected area. The reward function comprehensively considers the average task response time, the overall energy efficiency ratio of the cluster, the overheating frequency of key nodes, and the service level agreement (SLA) achievement rate. The risk avoidance controller uses a risk prediction model based on Bayesian causal networks to assess the probability of cascading failures that may be caused by a specific scheduling scheme, and generates a suppression signal to feed back to the multi-objective optimization scheduler when the predicted risk exceeds the safety threshold. The runtime dynamic tuning unit is used to deploy lightweight agent programs to collect performance inversion data during the actual execution of tasks, including micro-architecture level indicators such as kernel function execution cycle, cache hit deviation, and PCIe transmission efficiency loss. The feedback data is then aggregated into the online learning engine to continuously update the task execution time prediction model and resource consumption estimation function. Incremental migration mechanism is used for long-running continuous learning tasks to dynamically adjust the batch size, gradient synchronization frequency and checkpoint saving interval based on real-time resource conditions, while ensuring model convergence. The green computing management unit is used to directly connect to the data center infrastructure management system (BMS) to obtain rack-level cooling air temperature, CRAC unit output status and mains electricity price fluctuation curves. Combined with chip-level thermodynamic simulation models, it predicts the temperature rise trend of hot spots and triggers control plans. This invention relates to an AI-based big data computing power analysis application scenario management platform. Its core objective is to address the low overall operational efficiency issues in traditional computing power management systems caused by fragmented task and resource representation, static scheduling logic, delayed risk response, and isolated energy efficiency control. The platform constructs a closed-loop governance architecture across the entire chain, from task access, intent understanding, profile generation, resource modeling, semantic matching, intelligent scheduling, risk warning, dynamic optimization to green operation and maintenance, enabling refined, adaptive, and sustainable management of heterogeneous computing power environments. In this embodiment, the system is deployed in a geographically distributed hybrid cloud computing power cluster, comprising high-performance GPU training nodes, low-power inference accelerator cards, FPGA custom computing units, and a large-scale distributed storage array. All components are interconnected via a high-speed InfiniBand network, and traffic scheduling is managed by a unified software-defined networking (SDN) controller.
[0015] The system's overall technical process begins with task access, followed by context parsing, task profiling, resource status awareness, topology modeling, semantic space mapping, scheduling decision generation, runtime monitoring and feedback optimization, ultimately forming a continuously evolving intelligent governance closed loop. This process breaks away from the traditional linear "request-queue-allocation" model, instead adopting an iterative governance paradigm of "perception-cognition-decision-execution-verification-correction," ensuring that the system can autonomously adjust its behavior strategies according to changes in external load and internal state evolution.
[0016] The unified access gateway, acting as the system's entry point, handles the unified reception and initial classification of task requests. It exposes a standard RESTful interface, supporting JSON-formatted task submission bodies. These bodies must include a task type field (with values ranging from "image recognition," "natural language processing," "scientific computing," "recommendation system training," and "edge inference"), input data volume (in GB), expected latest completion time (ISO8601 timestamp), maximum allowable error rate (floating-point number, representing precision tolerance), and a single task energy consumption budget limit (in kilowatt-hours). The gateway incorporates request validation logic, immediately returning an error code 400 for missing required fields or out-of-bounds values, and logging the exception for subsequent auditing. For valid requests, the gateway assigns a globally unique task ID, composed of a timestamp prefix, regional code, and a random sequence, with a fixed length of 32 hexadecimal characters. Simultaneously, the gateway starts a timer to mark the initial time the task enters the system, serving as a time benchmark for subsequent SLA compliance evaluation. All verified task requests are encapsulated into a standardized message structure, with a timestamp and source identifier appended, and then pushed to the internal message bus, where the context resolution engine subscribes to and consumes them.
[0017] Upon receiving the task request message, the context parsing engine initiates the semantic understanding process. At its core is a pre-trained Transformer architecture semantic understanding model, which has been unsupervised pre-trained on a corpus of tens of millions of real AI task description texts, covering various language styles from academic paper abstracts and engineering documents to user support tickets. The model input is a free text description field from the task request (if it exists), such as "Fine-tuning of a subset of ImageNet using ResNet50 must be completed within 2 hours, using FP16 precision, with a target accuracy of no less than 95%." The model first segments the text into a sequence of sub-words using a word segmenter, and then inputs it into a multi-layer self-attention encoder for context representation extraction. To enhance the model's ability to distinguish between similar but critically different task types, the system introduces a contrastive learning mechanism during the pre-training phase. This involves constructing positive sample pairs (e.g., "BERT fine-tuning for sentiment analysis" and "RoBERTa fine-tuning for emotion recognition") and negative sample pairs (e.g., "YOLOv7 object detection" and "WaveNet speech synthesis"), and optimizing model parameters using a triplet loss function. This ensures that the distance between similar tasks in the latent space is significantly smaller than the distance between dissimilar tasks. After encoder processing, the model outputs a dense vector of dimension 768, representing the deep semantic representation of the task. This vector is fed into the downstream multi-task classification head, which predicts the application domain of the task, the main type of computation required (floating-point intensive / integer intensive / memory bandwidth sensitive), the expected parallel scale (single-node / multi-node AllReduce / parameter server architecture), and whether it possesses incremental learning characteristics. All prediction results, along with the original metadata, constitute an intermediate representation structure and are then passed to the task profile generation module.
[0018] After receiving the intermediate representation structure, the task profiling module initiates a fine-grained feature engineering process. The module first queries the historical database, retrieving execution records of historical tasks with the task ID or semantic similarity exceeding a set threshold. It extracts statistical indicators such as average computational intensity (FLOPS / byte), L1 / L2 cache hit rate distribution, DDR and HBM memory access ratio, NCCL communication call frequency and latency distribution, checkpoint retention period, and failure retries. If it's a new task type appearing for the first time, the default template is used for initialization. The module further combines the computational type prediction results output by the context parsing engine to activate the corresponding analysis plugin. For floating-point intensive tasks, the CUDA kernel simulator is launched, inferring the theoretical computational load and memory access requirements based on model structure keywords in the task description (such as "convolutional layer," "fully connected layer," and "number of attention heads"). For communication-sensitive tasks, a topology-aware communication cost estimation algorithm is invoked to predict the latency of the AllReduce operation based on the expected parallel scale and network bandwidth configuration. All extracted and derived features are organized into a structured multidimensional vector called the task profile vector, with a dimension of up to 512. Each dimension corresponds to a specific quantitative indicator or a normalized category code. This vector is persistently stored in a high-performance key-value database, with the key being the task ID and an expiration date set to 7 days after the task's lifecycle ends, serving as data support for subsequent scheduling decisions and model training.
[0019] The resource status awareness layer consists of lightweight probe agents deployed on each physical computing node. These agents run as daemons within the operating system kernel, possessing high-priority scheduling privileges to ensure real-time sampling. Each probe agent actively collects a local hardware status snapshot every 2 seconds, strictly adhering to a pre-defined list of metrics. CPU-related metrics include current core frequency, C-state sleep state distribution, instruction execution rate (IPC), and branch prediction error rate. GPU-related metrics include SM active percentage, memory bandwidth utilization, Tensor Core utilization, temperature sensor readings (4 measurement points per GPU), and fan speed feedback. NVLink link metrics include bidirectional throughput, number of bit error retransmissions, and number of flow control pause frames. Storage device metrics include NVMe queue depth, read / write IOPS, and end-to-end latency percentiles (p50 / p95 / p99). Power and heat dissipation metrics include node instantaneous power consumption (Watt), power conversion efficiency η, rack return air temperature, and estimated lateral thermal conductivity between adjacent nodes. All raw data undergoes initial cleaning locally to remove obvious outliers (such as negative temperatures and excessive frequencies). Outliers are identified using the Z-score method, and data exceeding the mean ± 3 standard deviations for three consecutive times are marked as suspicious and trigger a manual verification process. The cleaned data is packaged into small messages in Protobuf format, and after adding a sending timestamp and a unique node identifier, it is pushed in batches to the central aggregation service via gRPC long connections.
[0020] After receiving batch status data from multiple nodes, the resource topology modeling unit initiates the resource graph construction process. The unit first performs feature normalization, mapping indicators of different dimensions to the [0,1] interval. The normalization formula is:
[0021] in, This is the normalized value; These are the original indicator values; and These represent the minimum and maximum values observed for the metric over the past hour. The normalized vector is then input as node attributes into the graph neural network model. This model employs a heterogeneous graph attention mechanism, capable of handling three node types simultaneously: compute nodes (including GPU / CPU / FPGA), storage nodes (SSD / NVMe / Optane), and network nodes (switches / routers). Edges are also categorized into three types: compute-compute edges (representing MPI communication paths), compute-storage edges (representing DMA transfer channels), and network-network edges (representing fiber optic connections). For each edge, the model maintains a learnable edge type embedding vector and introduces edge type-aware weighting factors into the attention computation. Specifically, nodes... For neighboring nodes Attention coefficient The calculation is as follows:
[0022] in, and For nodes and The input feature vector; For edge ( , Type embedding; and It is a trainable linear transformation matrix; This represents the shared parameter vector for the attention mechanism. This represents a vector concatenation operation; For nodes The model outputs a high-order embedding representation of each node through a multi-layer graph attention network. These representations not only include the node's own attributes but also incorporate the influence of multi-hop neighbors, thus accurately reflecting its structural position in the network. All node embeddings are combined into a dynamically updated resource topology graph, with the graph structure refreshed every 5 seconds. Older versions are automatically archived to a time-series graph database for retrospective analysis.
[0023] The computational semantic space mapper is responsible for mapping task profile vectors and resource topology graphs to a unified high-dimensional semantic space. The mapper comprises two parallel nonlinear projection sub-networks: a task encoder and a resource encoder. The task encoder is a four-layer fully connected neural network with 256 neurons per layer, using GELU activation, and the final layer outputs a dimension of 256, compressing the task profile vector into a compact semantic fingerprint. The resource encoder reuses the node embeddings output by the resource topology modeling unit. For each candidate computation node, its embedding vector is further reduced to 256 dimensions using a two-layer MLP, serving as the node's coordinates in the semantic space. The mapper then calculates the cosine similarity between the task fingerprint and the coordinates of each candidate node, as a preliminary task-resource matching score. To improve the discriminative power of the score, the system employs a triplet loss function for end-to-end training of the entire mapping framework. During training, positive sample triples (same task type - matching resource, same task type - non-matching resource) and negative sample triples (different task types - matching resource) are constructed, and the model is forced to learn such that the distance between positive sample pairs is less than that between negative sample pairs by at least one boundary margin α=0.5. The training data comes from real scheduling records and post-evaluation labels from the past 6 months. After training, the mapper can recognize complex association rules at the semantic level, such as "high communication density tasks should be preferentially assigned to NVLink direct-connected nodes" and "memory-sensitive models are more suitable for deployment on the large-capacity A100 of HBM than the V100".
[0024] After receiving the task-resource matching score matrix, the intelligent scheduling decision center initiates a dual-track parallel decision-making process. The first sub-module of the center is a multi-objective optimization scheduler, which employs a hierarchical reinforcement learning architecture. The upper-layer policy network is a Long Short-Term Memory (LSTM) network, receiving the global system state as input, including macro-level indicators such as the current task queue length distribution, resource idle rate in each partition, average temperature slope over the last 10 minutes, and electricity price time period indicators. The LSTM outputs a three-dimensional probability distribution, corresponding to the confidence level of assigning the current task to the high-performance computing zone, energy efficiency priority zone, or hybrid elastic zone. After sampling, the target region is determined, and this decision result is passed as a constraint to the lower-layer action network. The lower-layer action network is a graph convolutional policy network, whose input is the latest embedding vectors of all available computing nodes within the selected region and their interconnections. The network outputs the probability of each node being selected, ultimately selecting the node with the highest probability to execute the task binding operation. The scheduler's reward function is designed as a weighted sum of four indicators:
[0025] in, This is the total reward value; to These are dynamic weighting coefficients, and their sum is 1. The actual execution time of the task; The expected completion time is set to 0 if the timeout is exceeded. The actual output efficiency of high-performance computing resources (unit: PetaFLOPS-day). Total energy consumption of the cluster (unit: megawatt-hours); The percentage of cumulative time during which the temperature at critical nodes exceeds 85°C; This is a Boolean indicator variable; it is 1 if the task is accomplished, and 0 otherwise. The weighting coefficient is dynamically adjusted based on operational strategies, for example, increasing during peak hours. and Increase during off-peak hours at night .
[0026] The second submodule of the central control is the risk avoidance controller, which operates independently of the scheduler, forming a supervisory and checks-and-balances mechanism. The controller constructs a risk prediction model based on a Bayesian causal network. The network structure explicitly encodes a library of known physical failure modes, including: thermal fatigue cumulative effects (the Weibull distribution relationship between the number of high-temperature cycles and solder joint life decay), voltage-to-noise ratio deterioration paths (a probabilistic model of current surges leading to VDD collapse and subsequent soft errors), and network congestion avalanche mechanisms (a nonlinear transition function between switch buffer fill rate and packet loss rate). Model nodes represent observable variables (such as node temperature, power fluctuations, and queue length) and latent variables (such as device aging and potential short-circuit risks), while edges represent causal dependencies. When the multi-objective optimization scheduler proposes a candidate allocation scheme, the controller instantiates a virtual execution scenario, injects the scheme into the causal network, and uses variational inference algorithms to quickly estimate the posterior probability distribution of key consequence variables, particularly the probability prisk of "the target node temperature exceeding the critical value of 90°C within the next 10 minutes." If the prisk exceeds the preset safety threshold of 0.15, the controller immediately generates a suppression signal to block the scheduling action and notifies the scheduler to reschedule. The suppression signal contains a summary of the risk root cause analysis, such as "Adding a high-power task will cause local heat dissipation imbalance because neighboring nodes are already fully loaded."
[0027] The runtime dynamic tuning unit begins operating immediately after task initiation. The unit uses lightweight agents deployed on each compute node to collect microarchitecture-level performance events during task execution at 100-millisecond granularity. The agents utilize the Linux `perf_event_open` system call interface to monitor the following PMU events: CPU-side L1D cache load misses, LLC last-level cache contention cycles, and branch prediction error pipeline flushes; GPU-side SM warp launch pause cause classification (memory dependency, synchronization barrier, instruction launch bottleneck), memory transaction merging efficiency, and Tensor Core utilization fluctuation curves; PCIe layer uplink / downlink bandwidth utilization, and packet splitting and reassembly overhead. All raw counter data is aggregated locally using a sliding window (1 second wide), calculating the mean, variance, and peak value, and compressed into binary format before being uploaded to the central data lake. The online learning engine in the data lake periodically pulls this feedback data and performs residual analysis against the output of the pre-scheduling task execution time prediction model. The engine uses a sliding time window mechanism to maintain execution logs for the most recent 7 days, constructing a random forest regressor. Taking task profile vectors, resource status snapshots, and actual observation times as input, it predicts the systematic deviation in task completion times. The newly established residual model is injected into the basic prediction service, forming a compensation mechanism that gradually approximates the true value in subsequent predictions. For task categories with persistently high prediction errors, the system automatically triggers an expert rule review process, where operations personnel intervene to analyze whether there are model misjudgments or hardware anomalies.
[0028] The incremental migration mechanism is designed for continuous learning tasks, which typically involve infinite data flow and long-term operation. During task execution, the mechanism continuously monitors node load changes in the resource topology modeling unit output. When it detects that the current node's memory utilization exceeds 85% for five consecutive sampling periods, or the NVLink bandwidth utilization is below 60% (indicating communication bottleneck), the mechanism initiates an adaptive parameter adjustment process. The adjustment follows a monotonic constraint principle: the batch size is only allowed to increase incrementally (by a factor of 2) or remain unchanged, and decreasing it is prohibited to avoid gradient explosion risk; the gradient synchronization frequency can be reduced to 1 / 2 or 1 / 4 of the original period, but must not fall below the minimum safety threshold (synchronization once every 100 rounds); the checkpoint saving interval can be dynamically extended based on disk IOPS, but cannot exceed four times the original setting. All adjustment actions are hot-updated to the training framework via the container runtime interface without interrupting task execution. The adjustment strategy is based on a lightweight reinforcement learning agent, whose reward function is: ,in This represents the change in time during a single iteration. To train a stability indicator (based on the smoothness of the loss function), the agent pre-simulates various adjustment combinations in a simulated environment and selects the one with the highest overall reward for implementation.
[0029] The green computing management unit establishes a bidirectional data channel with the data center infrastructure management system (BMS) to achieve collaborative optimization of IT equipment and cooling systems. Every 5 minutes, the unit obtains rack-level cooling air temperature, current CRAC unit output percentage, chilled water supply temperature setpoint, and real-time mains electricity price from the BMS. Simultaneously, the unit runs a chip-level thermodynamic simulation model, which solves the three-dimensional heat conduction equation based on the finite difference method.
[0030] in, The density of the material; Specific heat capacity; Thermal conductivity; This refers to the heat generation power per unit volume, which is mapped from the GPU power consumption. The model is designed for a temperature field. The model mesh resolution is 2 mm, covering the entire GPU board area, and boundary conditions are driven by measured point temperatures. Every 5 minutes, the model undergoes Kalman filtering correction based on the actual temperature reported by the probe, correcting estimation errors in heat source intensity and heat dissipation coefficient. Based on the corrected model, the system can predict the temperature rise trend of hotspot areas 15 minutes in advance. If the predicted temperature in a certain area will exceed the safety threshold, a step-by-step increase command for fan speed is triggered in advance, or a task migration plan is initiated to migrate near-saturation computing loads to lower-temperature areas. Under the time-of-use pricing mechanism, the green computing management unit locks the scheduling window for batch processing tasks (such as offline data analysis and model pre-training) during off-peak hours (23:00 to 7:00 the next day) and activates the energy storage device power supply mode during this period, utilizing lithium battery packs charged by photovoltaic power during the day to power some non-critical loads, thereby reducing peak grid load by up to 18%. The system also integrates a carbon footprint tracking function, based on the actual power consumption of each task and the real-time carbon emission factor (kg CO2 / kWh) published by the local power grid, using a formula... The environmental impact index of the entire life cycle is calculated cumulatively, and the results are written into the task metadata for use in generating ESG reports.
[0031] This embodiment, through the close collaboration of the aforementioned modules, constructs an intelligent computing power governance platform with deep cognitive capabilities and autonomous evolutionary characteristics. The system is no longer limited to passively responding to resource requests, but proactively understands the essential needs of tasks, accurately characterizes the dynamic characteristics of resources, and seeks a Pareto optimal solution among performance, stability, cost, and sustainability. Experimental verification shows that under typical AI workload scenarios, the platform increases the average GPU utilization rate from 48% in traditional solutions to 79%, mainly due to the semantic matching mechanism effectively avoiding the phenomenon of high-end resources executing lightweight tasks; the tail latency (p99) of high-priority tasks is reduced by 63%, stemming from the risk avoidance controller successfully preventing 37% of potential congestion events; the task completion time prediction accuracy converges from an initial 71% to 94% within two weeks, demonstrating the effectiveness of closed-loop feedback optimization; the average carbon emission intensity per task decreases by 37%, and the total annual electricity expenditure decreases by 22%, proving the dual economic and environmental value of green computing strategies. The platform achieves a fundamental shift from "extensive resource supply" to "refined scenario governance," providing a scalable, reliable, and sustainable technological paradigm for next-generation artificial intelligence infrastructure.
[0032] Existing technologies generally employ static partitioning or multi-tenant isolation strategies for computing power management, lacking the ability to deeply explore the intrinsic relationship between tasks and resources. Most platforms simply match tasks based on their declared resource requirements (e.g., "requires 2 V100s"), ignoring the coupling relationship between the actual computing mode and system load, leading to resource fragmentation and low utilization. Furthermore, existing schedulers are mostly based on first-come, first-served or shortest-job-first principles, unable to handle the collaborative challenges of complex QoS requirements and multi-dimensional optimization goals. More seriously, traditional systems lack proactive risk warning mechanisms, often only taking remedial measures after node overheating or network congestion occurs, causing service interruptions and hardware damage. This solution, by constructing a joint semantic representation space for tasks and resources, achieves a paradigm shift from "resource-adapted tasks" to "semantic matching and collaboration," significantly improving resource utilization in heterogeneous computing environments. This solution introduces a dual-track decision-making mechanism of hierarchical reinforcement learning and causal risk control, proactively avoiding potential systemic risks while pursuing maximum performance, reducing unplanned downtime events caused by localized overheating by 82%. This solution establishes an end-to-end closed-loop feedback optimization system, continuously correcting cognitive biases in the scheduling model by absorbing real-world execution data, thus significantly enhancing the predictability of system behavior. Deeply integrating green computing concepts, this solution incorporates energy consumption, heat dissipation, and grid load into a unified optimization framework, providing a practical and feasible technical path for building sustainable AI infrastructure.
[0033] The upper-layer policy network and lower-layer action network of the multi-objective optimization scheduler communicate by sharing some hidden states. In this embodiment, the upper-layer policy network has a decision cycle of 30 seconds and is responsible for macro-level resource partitioning decisions, determining whether tasks should be allocated to the high-performance computing zone, the energy-efficiency priority zone, or the hybrid elastic zone. The high-performance computing zone contains the latest generation of GPU clusters, suitable for high-concurrency training tasks; the energy-efficiency priority zone consists of low-power ASIC accelerator cards, dedicated to edge inference and lightweight model services; the hybrid elastic zone is a reconfigurable FPGA array, supporting dynamic programming of different computing logic to adapt to diverse loads. The upper-layer network scans the global task queue and resource pool status every 30 seconds, generates partitioning allocation suggestions, and passes the suggestions along with high-level semantic feature vectors to the lower-layer action network. The lower-layer action network has a decision cycle of 2 seconds and focuses on executing specific node binding operations within the selected area. The network receives the state summary passed from the upper layer and, combined with the real-time topology information of the local resource graph, calculates the adaptation score of each available node through a graph attention mechanism. The node with the highest score is selected, and deployment instructions are issued through the container orchestration engine. The sharing of hidden states between upper and lower layers is achieved through a fixed intermediate vector with a length of 128. This vector is refreshed by the upper layer every 30 seconds, and the lower layer superimposes local observation features on it to make decisions, ensuring the consistency between macro-strategy and micro-execution.
[0034] The Bayesian causal network of the risk avoidance controller contains an explicitly encoded library of physical failure modes. In this embodiment, network nodes encompass observable variables (such as node temperature, power fluctuations, and queue length) and latent variables (such as device aging and potential short-circuit risks). Edges represent causal dependencies, such as a complete chain of "high power density → local temperature rise → thermal stress accumulation → solder joint fatigue → increased connection resistance → increased voltage drop → increased soft error rate." When the multi-objective optimization scheduler proposes a candidate allocation scheme, the controller instantiates a virtual execution scenario, injects the scheme into the causal network, and uses a variational inference algorithm to quickly estimate the posterior probability distribution of key consequence variables. The inference process uses a mean-field approximation, assuming that the posterior distributions of each latent variable are independent, and achieves efficient computation through iterative optimization of the evidence lower bound (ELBO). If the predicted probability of "the target node temperature exceeding the critical value of 90°C within the next 10 minutes" is greater than 0.15, the controller immediately generates a suppression signal, blocks the scheduling action, and notifies the scheduler to reschedule. The suppression signal contains a summary of the risk root cause analysis to help the scheduler understand the constraints.
[0035] The online learning engine of the runtime dynamic tuning unit uses a sliding time window mechanism to maintain the execution logs for the past 7 days. In this embodiment, the engine triggers a full model retraining process once every day at midnight, updating the random forest regressor parameters using the performance inversion data accumulated over the past 7 days. Training samples are organized in the form of triplets of (task profile vector, resource state snapshot, prediction bias), with the label being the difference between the actual completion time and the initial prediction time. After the model training is completed, the new version is released in a gray-scale release to the A / B testing channel and served in parallel with the old model for 1 hour, comparing the MAPE index of the two on newly arrived tasks. If the error of the new model decreases by more than 5%, it is switched to the primary model; otherwise, the old version is retained and the anomaly is recorded. This mechanism ensures that the model continues to evolve without introducing unstable factors.
[0036] The incremental migration mechanism follows the monotonic constraint principle when adjusting hyperparameters. In this embodiment, the batch size is only allowed to increase or remain unchanged to avoid training divergence due to frequent jitter. When the system decides to adjust, it first evaluates the ratio of the L2 norm of the current loss function gradient to the moving average. If the ratio is less than 1.2, the training is considered to be in a stable period, and the batch increase operation can be performed. The increase step size is twice the current value, but must not exceed 90% of the maximum theoretical value allowed by the hardware memory capacity. The gradient synchronization frequency can be reduced to 1 / 2 or 1 / 4 of the original period, but must not be lower than the minimum safety threshold. All adjustments are distributed through Kubernetes CRD custom resource definitions, and the changes are captured by the sidecar container and injected into the training process.
[0037] The thermodynamic simulation model of the green computing management unit is based on solving the three-dimensional heat conduction equation using the finite difference method. In this embodiment, the model mesh resolution is set to 2 mm precision, at the chip package level. Spatial discretization uses a central difference discretization scheme, and explicit Euler method is used for time progression. The time step is set to 0.5 seconds to ensure numerical stability. Initial conditions are inherited from the final state of the previous simulation, and boundary conditions are jointly determined by the rack cooling temperature and fan speed settings provided by the BMS. Heat source item. The distribution pattern is pre-defined based on the chip layout and obtained by interpolating GPU power consumption monitoring data. The model is corrected by Kalman filtering every 5 minutes based on the measured temperature. The state vector contains the temperature and equivalent thermal conductivity of each grid point, and the observation vector is the probe measurement value. The model bias is corrected through prediction-update loop to keep its long-term prediction error within ±1.5℃.
[0038] This embodiment, through an in-depth exploration of the aforementioned technical details, comprehensively reveals the engineering implementation path and core technical mechanisms of the platform of this invention. From task semantic parsing to resource topology modeling, from semantic space mapping to dual-track scheduling decisions, and then to runtime feedback optimization and green collaborative control, each step demonstrates a high degree of automation, intelligence, and systematicity. The platform not only solves key pain points in current computing power management but also lays a solid foundation for the sustainable evolution of future ultra-large-scale AI systems.
Claims
1. A big data computing power analysis application scenario management platform based on artificial intelligence, characterized in that, include: A unified access gateway is used to receive AI task requests from different business systems to obtain a set of AI task requests; A context parsing engine is used to perform semantic parsing on the set of AI task requests to obtain a set of standardized workload description vectors; The task profile generation module is used to generate a set of multi-dimensional task profile vectors based on the set of standardized workload description vectors. The resource status awareness layer is used to collect hardware operating parameters of heterogeneous computing nodes through distributed probes to obtain a set of hardware operating parameters; The resource topology modeling unit is used to construct a dynamic physical resource connection relationship diagram based on the set of hardware operating parameters; A computational power semantic space mapper is used to embed the set of multi-dimensional task profile vectors and the dynamic physical resource connection relationship graph into a unified high-dimensional semantic space to obtain a set of task-resource matching scores. An intelligent scheduling decision center is used to generate a resource allocation scheme based on the set of task-resource matching scores. The intelligent scheduling decision center includes: a multi-objective optimization scheduler, used to generate a preliminary resource allocation scheme based on a hierarchical reinforcement learning architecture; and a risk avoidance controller, used to evaluate the cascading failure probability of the preliminary resource allocation scheme based on a Bayesian causal network, and generate a suppression signal to correct the preliminary resource allocation scheme when the cascading failure probability exceeds a safety threshold. The runtime dynamic tuning unit is used to collect microarchitecture-level performance inversion data during task execution and update the task execution time prediction model and resource consumption estimation function based on the microarchitecture-level performance inversion data. The green computing management unit is used to acquire data center infrastructure management data and combine chip-level thermodynamic simulation models to predict the temperature rise trend of hot spots and trigger control plans.
2. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 1, characterized in that, The AI task request includes: task type identifier, data volume, expected completion time, accuracy requirements, and energy consumption budget.
3. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 2, characterized in that, The multidimensional task profile vector includes: computational density, memory access mode, communication topology, fault tolerance level, and statistical characteristics of historical execution trajectory.
4. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 3, characterized in that, The hardware operating parameters include: CPU utilization, GPU utilization, video memory usage, high-speed interconnect bandwidth usage, temperature gradient distribution, power efficiency, and solid-state storage input / output operation fluctuations.
5. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 4, characterized in that, The computing power semantic space mapper includes: The task encoding subunit is used to perform nonlinear projection on each multidimensional task profile vector in the set of multidimensional task profile vectors to obtain a set of task semantic fingerprint vectors. The resource coding subunit is used to perform nonlinear projection on the embedding vectors of each computing node in the dynamic physical resource connection graph to obtain a set of resource semantic coordinate vectors. The matching score calculation subunit is used to calculate the similarity between the set of task semantic fingerprint vectors and the set of resource semantic coordinate vectors to obtain the set of task-resource matching scores.
6. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 5, characterized in that, The multi-objective optimization scheduler includes: The upper-level policy network is used to generate macro-level resource partitioning decisions based on the global system state; The lower-level action network is used to perform specific computing node binding operations within the region specified by the macro-resource partitioning decision; The upper-layer policy network and the lower-layer action network communicate information by sharing some hidden states.
7. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 6, characterized in that, The risk avoidance controller includes: A physical failure mode library for storing explicitly encoded causal dependencies of physical failures; A virtual execution scenario instantiation unit is used to inject the preliminary resource allocation scheme into a Bayesian causal network constructed based on the physical failure mode library; The risk probability inference unit is used to estimate the posterior probability distribution of key consequence variables using a variational inference algorithm to obtain the cascading failure probability.
8. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 7, characterized in that, The runtime dynamic tuning unit includes: A lightweight agent is deployed on each compute node to collect the microarchitecture-level performance inversion data. An online learning engine is used to maintain recent execution logs using a sliding time window mechanism, and to build a residual prediction model based on the recent execution logs to compensate for the systematic bias of the task execution time prediction model.
9. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 8, characterized in that, The platform also includes an incremental migration mechanism, which is used to dynamically adjust the batch size, gradient synchronization frequency and checkpoint saving interval based on real-time resource conditions for long-running continuous learning tasks, while ensuring model convergence. The adjustment of the batch size follows the monotonic constraint principle.
10. The big data computing power analysis application scenario management platform based on artificial intelligence according to claim 9, characterized in that, The green computing management unit is used to obtain rack-level cooling air temperature, chiller output status and mains electricity price fluctuation curve from the data center infrastructure management system, and predict the temperature rise trend of hotspot areas in the future based on the chip-level thermodynamic simulation model and the rack-level cooling air temperature.