Adaptive heterogeneous server system based on ai prediction and dynamic hardware reconfiguration and resource management method thereof

By combining modular heterogeneous hardware unit pools, reconfigurable interconnect Fabric, AI prediction and intelligent heat dissipation system, the hardware configuration and heat dissipation strategy are dynamically adjusted, solving the problems of resource waste and low heat dissipation efficiency in traditional servers, and achieving efficient resource utilization and energy management.

CN121187737BActive Publication Date: 2026-02-27四川华鲲振宇智能科技有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511725434.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-27
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

Traditional server hardware architectures are difficult to adapt flexibly to diverse workloads, resulting in resource waste and performance bottlenecks, and heat dissipation methods are difficult to manage in a refined manner.

Method used

It adopts a modular heterogeneous hardware unit pool, a reconfigurable interconnect fabric, an AI workload perception and prediction engine, an intelligent heat dissipation system and a resource orchestrator, and dynamically adjusts hardware configuration and heat dissipation strategies through AI prediction to achieve real-time optimization.

Benefits of technology

It improves server resource utilization, reduces energy consumption, adapts to diverse workload requirements, and solves the problems of lagging resource configuration and low heat dissipation efficiency in traditional servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121187737B_ABST
    Figure CN121187737B_ABST
Patent Text Reader

Abstract

The application discloses an adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction and a resource management method thereof, relates to the technical field of server hardware architecture and resource management, and discloses the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction and the resource management method thereof. Through the dynamic adjustment capability of a modular heterogeneous hardware unit pool and a reconfigurable interconnection fabric, in combination with the accurate prediction of an AI prediction engine on resource demand and power consumption, real-time optimization configuration and heat dissipation management of hardware resources are achieved, the server resource utilization rate can be improved, the energy consumption can be reduced, and the diversified work load demand can be adapted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of server hardware architecture and resource management, in particular to an adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction and a resource management method thereof. BACKGROUND

[0002] With the rise of the Internet and cloud computing, higher requirements are put forward for the computing power, storage capacity, network bandwidth and scalability of servers. To cope with large-scale and diversified business demands, resource pooling technology has been developed. Computing, storage and network resources are pooled at the logical level and scheduled and allocated through software-defined methods. In this process, if a certain type of resource of a server is exhausted while other resources are still available, the server cannot undertake new tasks that require that resource, resulting in resource waste. For tasks that require specific hardware configurations, traditional servers are difficult to adapt flexibly. In recent years, with the explosive growth of AI, big data and high-performance computing applications, workloads have shown highly dynamic and diverse computing patterns. The fixed hardware architecture of traditional servers is difficult to efficiently meet these demands, resulting in either excessive redundancy to cope with peak values or slow response and performance bottlenecks.

[0003] In order to improve hardware flexibility, existing technologies such as combined infrastructure aim to decouple computing, storage, network and other resources and combine them on demand through high-speed interconnection technology. Most of the above solutions rely on passive scheduling by upper-layer software, lack deep understanding and forward-looking prediction capabilities for workloads, and the real-time and intelligent degree of resource reconstruction needs to be improved. At the same time, with the increase in hardware power density, the problem of heat dissipation is becoming increasingly prominent, and traditional heat dissipation methods are difficult to achieve fine and low-energy consumption management.

[0004] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0005] The main purpose of the present application is to provide an adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction and a resource management method thereof, aiming to improve server resource utilization, reduce energy consumption, and realize dynamic hardware reconstruction and intelligent heat dissipation management.

[0006] To achieve the above purpose, the present application provides an adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction, which comprises:

[0007] A modular heterogeneous hardware unit pool containing online hot replaceable computing units, storage units and network interface units;

[0008] A reconfigurable interconnection Fabric connecting the modular heterogeneous hardware unit pool;

[0009] An AI workload perception and prediction engine connected to the pool of modular heterogeneous hardware units for collecting system telemetry data and generating resource demand prediction data and power consumption distribution prediction data;

[0010] An intelligent cooling system connected to the AI workload perception and prediction engine and the pool of modular heterogeneous hardware units;

[0011] A resource orchestrator connected to the AI workload perception and prediction engine, the reconfigurable interconnection Fabric and the intelligent cooling system, respectively;

[0012] A unified hardware abstraction layer connected to the resource orchestrator and the pool of modular heterogeneous hardware units;

[0013] The resource orchestrator is configured to receive resource intent request data of an upper-layer application and generate hardware configuration scheme data based on the resource demand prediction data;

[0014] The reconfigurable interconnection Fabric is configured to dynamically adjust connection topology, bandwidth and communication protocol among hardware units according to the hardware configuration scheme data;

[0015] The intelligent cooling system is configured to perform cooling control on hardware units according to the power consumption distribution prediction data.

[0016] In an embodiment, the AI workload perception and prediction engine comprises:

[0017] A data collection unit configured to collect CPU utilization, memory bandwidth, storage IOPS and network throughput data from an operating system and hardware sensors;

[0018] A feature processing unit connected to the data collection unit and configured to perform cleaning, noise reduction and normalization processing on the data collected by the data collection unit to generate state feature data;

[0019] A prediction model unit connected to the feature processing unit and configured to apply a long short-term memory neural network to process the state feature data to output the resource demand prediction data and the power consumption distribution prediction data.

[0020] In an embodiment, the reconfigurable interconnection Fabric comprises:

[0021] An optoelectronic hybrid interconnection structure including electrical links and optical interconnection links;

[0022] An array of switch chips connected to the electrical links and the optical interconnection links;

[0023] An exchange control unit is connected to the exchange chip array and the resource orchestrator, and is configured to adjust a connection path within a preset time according to the hardware configuration scheme data.

[0024] In an embodiment, the intelligent heat dissipation system comprises:

[0025] A temperature sensor network is arranged on surfaces of the computing units and the storage units.

[0026] A partitioned liquid cooling unit comprises liquid cooling branches with independently adjustable flow rates.

[0027] An air flow guiding device comprises an array of adjustable speed fans.

[0028] A heat dissipation control unit is connected to the temperature sensor network, the partitioned liquid cooling unit and the air flow guiding device, and is configured to control liquid cooling flow distribution and fan rotation speed based on the power consumption distribution prediction data.

[0029] In an embodiment, the resource orchestrator comprises:

[0030] An intent analysis module is configured to convert resource intent request data into resource demand parameter data.

[0031] A resource evaluation module is connected to the unified hardware abstraction layer, and is configured to obtain current available hardware resource state data.

[0032] An optimization decision module is connected to the intent analysis module and the AI workload perception and prediction engine, and is configured to generate the hardware configuration scheme data based on the resource demand parameter data, the resource demand prediction data and the available hardware resource state data.

[0033] In addition, to achieve the above-mentioned purpose, the present application further provides a resource management method applied to the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction, and the resource management method comprises:

[0034] The resource orchestrator receives resource intent request data of an upper layer application.

[0035] The AI workload perception and prediction engine collects system telemetry data, performs cleaning and normalization processing, and generates state feature data.

[0036] Based on the state feature data, a preset prediction model generates resource demand prediction data and power consumption distribution prediction data.

[0037] The resource orchestrator generates hardware configuration scheme data based on the resource intent request data and the resource demand prediction data.

[0038] The control reconfigurable interconnection fabric dynamically adjusts the connection topology, bandwidth and communication protocol between hardware units according to the hardware configuration scheme data;

[0039] The control intelligent heat dissipation system adjusts the liquid cooling flow and fan speed according to the power consumption distribution prediction data.

[0040] In an embodiment, the step of generating resource demand prediction data and power consumption distribution prediction data by a preset prediction model based on the state feature data comprises:

[0041] extracting periodic features and burst features from the state feature data;

[0042] inputting the periodic features and burst features into a long short-term memory neural network;

[0043] calculating the current time memory state by a memory cell of the long short-term memory neural network;

[0044] generating resource demand prediction data and power consumption distribution prediction data for a future preset time period by an output gate.

[0045] In an embodiment, the step of calculating the current time memory state by a memory cell of the long short-term memory neural network comprises:

[0046] calculating an input gate state based on the previous time hidden state and the current input feature;

[0047] calculating a forget gate state based on the previous time hidden state and the current input feature;

[0048] updating the memory cell state in combination with the forget gate state and the input gate state, and outputting the updated memory cell state.

[0049] In an embodiment, the step of generating hardware configuration scheme data by the resource orchestrator based on the resource intent request data and the resource demand prediction data comprises:

[0050] parsing the computing power demand parameter and the video memory demand parameter in the resource intent request data;

[0051] obtaining the current available hardware resource state data through a unified hardware abstraction layer;

[0052] calculating an optimal hardware combination scheme based on the computing power demand parameter, the video memory demand parameter, the available hardware resource state data and the resource demand prediction data;

[0053] generating hardware configuration scheme data containing module activation instructions and connection topology instructions.

[0054] In an embodiment, the step of calculating the optimal hardware combination scheme based on the computing power requirement parameter, the graphics memory requirement parameter, the available hardware resource state data, and the resource requirement prediction data comprises:

[0055] setting an optimization target function with performance priority or energy efficiency priority;

[0056] calculating a module combination scheme that meets the computing power requirement parameter and the graphics memory requirement parameter;

[0057] determining the bandwidth allocation ratio among the selected modules;

[0058] verifying that the module combination scheme meets the constraint conditions of the resource intent request, and outputting the verified hardware combination scheme.

[0059] The adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction and the resource management method thereof provided in the present application can realize real-time optimization configuration and heat dissipation management of hardware resources by the dynamic adjustment capability of the modularized heterogeneous hardware unit pool and the reconfigurable interconnection Fabric, in combination with the accurate prediction of resource requirements and power consumption by the AI prediction engine, thereby improving the server resource utilization rate, reducing energy consumption, and adapting to the diversified workload requirements. BRIEF DESCRIPTION OF DRAWINGS

[0060] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0062] Figure 1 The structural schematic diagram provided for an embodiment of the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction of the present application;

[0063] Figure 2 The structural schematic diagram provided for another embodiment of the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction of the present application;

[0064] Figure 3 The structural schematic diagram provided for still another embodiment of the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction of the present application;

[0065] Figure 4 The structural schematic diagram provided for still another embodiment of the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction of the present application;

[0066] Figure 5 Structure diagram provided by an embodiment of the adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction of the present application;

[0067] Figure 6 Flow diagram provided by an embodiment of the resource management method of the present application;

[0068] Figure 7 Figure 6 Detailed flow diagram of step S300 in the method;

[0069] Figure 8 Figure 6 Detailed flow diagram of step S400 in the method;

[0070] Figure 9 Figure 8 Detailed flow diagram of step S430 in the method.

[0071] Explanation of reference numerals:

[0072] 100, adaptive heterogeneous server system based on AI prediction and dynamic hardware reconstruction; 110, modular heterogeneous hardware unit pool; 120, reconfigurable interconnection Fabric; 130, AI workload perception and prediction engine; 140, intelligent cooling system; 150, resource orchestrator; 160, unified hardware abstraction layer; 131, data acquisition unit; 132, feature processing unit; 133, prediction model unit; 121, optoelectronic hybrid interconnection structure; 122, switching chip array; 123, switching control unit; 141, temperature sensor network; 142, partitioned liquid cooling unit; 143, air flow guiding device; 144, cooling control unit; 151, intention analysis module; 152, resource evaluation module; 153, optimization decision module.

[0073] The purposes, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0074] The technical solutions in the present application will be described clearly and completely in the present application with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. The components of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0075] ​​​It should be understood that like numerals and letters refer to like items throughout the drawings, and once an item is defined in one drawing, it is not necessary to further define and explain it in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0076] In the prior art, the server hardware architecture has long been in a fixed configuration mode, and resource expansion and adjustment rely on the replacement or addition of physical components. The interconnection topology and communication protocol between internal components of the traditional server are fixed at the time of factory shipment, and cannot be dynamically adjusted according to real-time load changes. When facing sudden computing tasks or heterogeneous workloads, such a static architecture is prone to low resource utilization, with some hardware units being in an idle state for a long time, while other units may experience performance degradation due to resource bottlenecks. At the same time, the cooling system usually adopts a global unified control strategy, which cannot accurately cool according to the power consumption changes in different areas, resulting in energy waste.

[0077] In order to solve the above problems, it is necessary to design a server system that can dynamically perceive the characteristics of workloads and actively adjust the hardware configuration. The inventors have observed that the periodic fluctuations and sudden growth of workloads are predictable, and if resource demand prediction is combined with hardware reconstruction capability, the optimal hardware combination can be configured in advance. In addition, the modular design of hardware units can support online hot replacement, and the dynamic adjustment of the interconnection structure can eliminate the fixed connection constraints between resources. By introducing a hierarchical control mechanism, the resource orchestration decision and the underlying hardware operation are decoupled, which can achieve fast response and flexible expansion.

[0078] Based on this, the embodiment of the present application provides an adaptive heterogeneous server system 100 based on AI prediction and dynamic hardware reconstruction, referring to Figure 1 , the adaptive heterogeneous server system 100 based on AI prediction and dynamic hardware reconstruction comprises:

[0079] A modular heterogeneous hardware unit pool 110, which includes online hot replaceable computing units, storage units and network interface units;

[0080] A reconfigurable interconnection Fabric 120 connected to the modular heterogeneous hardware unit pool 110;

[0081] An AI workload perception and prediction engine 130 connected to the modular heterogeneous hardware unit pool 110, for collecting system telemetry data and generating resource demand prediction data and power consumption distribution prediction data;

[0082] An intelligent cooling system 140 connected to the AI workload perception and prediction engine 130 and the modular heterogeneous hardware unit pool 110;

[0083] The resource orchestrator 150 is connected to the AI workload perception and prediction engine 130, the reconfigurable interconnection fabric 120, and the intelligent cooling system 140, respectively.

[0084] The unified hardware abstraction layer 160 is connected to the resource orchestrator 150 and the modular heterogeneous hardware unit pool 110.

[0085] The resource orchestrator 150 is configured to receive resource intention request data of an upper-layer application and generate hardware configuration scheme data based on the resource requirement prediction data.

[0086] The reconfigurable interconnection fabric 120 is configured to dynamically adjust the connection topology, bandwidth, and communication protocol between hardware units according to the hardware configuration scheme data.

[0087] The intelligent cooling system 140 is configured to perform cooling control on the hardware units according to the power consumption distribution prediction data.

[0088] In this embodiment, the modular heterogeneous hardware unit pool 110 refers to a resource collection composed of independently packaged hardware components supporting hot plugging, which can be implemented by using standard-size computing acceleration cards, storage modules, and network adapters, and each unit is connected through a unified interface and an interconnection structure. The reconfigurable interconnection structure refers to a communication network supporting dynamic adjustment of physical connection relationships, which can be implemented by using optoelectronic hybrid interconnection components and programmable switch chip arrays 122 to realize topology reconfiguration by changing the port mapping relationship of the switch chip. The AI workload perception and prediction engine 130 refers to an intelligent analysis module with time series data processing capability, which can be implemented by using an LSTM neural network model integrated with a data acquisition interface to predict future resource demand trends by analyzing historical load data. The intelligent cooling system 140 refers to a cooling device with partition control capability, which can be implemented by using a multi-branch liquid cooling circuit and an adjustable-speed fan array to allocate cooling resources according to the real-time power consumption of the hardware units.

[0089] In this embodiment, the system first continuously collects the running state data of each hardware unit through the sensor network during operation, including processor utilization, memory bandwidth occupancy, and other indicators. After cleaning and normalization processing, the collected data is input into the prediction model. The model identifies the load change rule by analyzing the time sequence characteristics and outputs the resource demand prediction results for the future period. The resource orchestrator 150 matches the prediction results with the resource requests of the upper-layer application to generate a scheme containing hardware unit activation instructions and interconnection configuration parameters. The reconfigurable interconnection structure completes the physical connection establishment or removal between specified units within milliseconds according to the scheme requirements, while adjusting the link bandwidth allocation ratio. The intelligent cooling system 140 synchronously receives the power consumption prediction data and adjusts the cooling liquid flow rate and fan speed in the area where the corresponding hardware unit is located in advance to ensure that the cooling capacity and power consumption distribution are matched in real time.

[0090] In this embodiment, the scheme can adjust the hardware combination in real time during operation through modular design and dynamic interconnection technology. The scheme can reduce invalid cooling energy consumption through prediction-driven partitioned cooling control. In addition, the scheme combines prediction data and real-time state for forward-looking resource allocation, significantly shortening the response delay. In this way, the on-demand dynamic combination of hardware resources is realized, effectively improving the execution efficiency of heterogeneous computing tasks. The system can automatically optimize the hardware connection topology according to the load characteristics, eliminating the communication bottleneck in traditional architectures. The prediction-driven cooling control mechanism reduces the redundant energy consumption of the cooling system and avoids the risk of local overheating. The collaborative work of the resource orchestrator 150 and the unified hardware abstraction layer 160 enables the upper-layer application to obtain optimal resource configuration without perceiving the details of the underlying hardware.

[0091] In a feasible implementation, with reference to Figure 2 , the AI workload perception and prediction engine 130 includes:

[0092] A data collection unit 131 for collecting CPU utilization, memory bandwidth, storage IOPS, and network throughput data from the operating system and hardware sensors;

[0093] A feature processing unit 132 connected to the data collection unit 131 for performing cleaning, noise reduction, and normalization processing on the data collected by the data collection unit 131 to generate state feature data;

[0094] A prediction model unit 133 connected to the feature processing unit 132 for applying a long short-term memory neural network to process the state feature data and output resource demand prediction data and power consumption distribution prediction data.

[0095] In this embodiment, the data acquisition unit 131 refers to a monitoring agent deployed in the operating system kernel layer and the hardware interface layer, which can be implemented based on the observability framework of eBPF technology, and real-time capture of resource usage indicators is realized through hooking system call interfaces and hardware performance counters. The feature processing unit 132 refers to a computing module with time series data preprocessing function, which can be implemented by combining sliding window mean filtering with Z-score standardization algorithm, and is used to eliminate sensor noise interference and unify the dimension. The prediction model unit 133 includes a preset prediction model, which is a machine learning model based on a recurrent neural network architecture, and can be implemented by a bidirectional LSTM network containing 128 hidden nodes, which captures the time series dependence of the workload through the memory unit.

[0096] In this embodiment, during operation, the data acquisition unit 131 continuously obtains original monitoring data from the CPU performance monitoring unit, the memory controller, the NVMe driver and the network card DMA engine. The feature processing unit 132 performs sliding average filtering on the original data with a fixed time window to eliminate transient peak noise, and then normalizes the multi-dimensional heterogeneous data to map indicators of different dimensions to a unified numerical interval. The preprocessed state feature data is input into the prediction model unit 133, and the long short-term memory neural network dynamically adjusts the memory unit state through the forgetting gate and the input gate mechanism, combines the historical load features and the current input features, and outputs the resource demand prediction and the corresponding power consumption distribution prediction in the future time period.

[0097] In this embodiment, the feature processing unit 132 realizes deep fusion of multi-source heterogeneous data, and the prediction model unit 133 utilizes the time series modeling capability of the long short-term memory neural network to accurately identify periodic load fluctuations and burst task features. Compared with the traditional ARIMA prediction method, the timeliness and accuracy of resource demand prediction are significantly improved, and the resource mismatch problem caused by passive response to workload changes in traditional server systems is solved. The feature processing unit 132 eliminates sensor noise and dimension differences to ensure the reliability of the prediction model input data; the long short-term memory neural network captures the time series features of the workload to realize forward-looking resource demand prediction and provide accurate decision basis for dynamic hardware reconstruction, thereby improving hardware resource utilization and reducing invalid energy consumption.

[0098] In a feasible implementation manner, referring to Figure 3 The reconfigurable interconnection Fabric 120 includes:

[0099] The optoelectronic hybrid interconnection structure 121 includes electrical links and optical interconnection links;

[0100] The switch chip array 122 connects the electrical links and the optical interconnection links;

[0101] The exchange control unit 123 is connected to the exchange chip array 122 and the resource orchestrator 150, and is used to adjust the connection path within a preset time according to the hardware configuration scheme data.

[0102] In this embodiment, the optoelectronic hybrid interconnection structure 121 refers to a physical connection architecture constructed by using both electrical transmission media and optical transmission media, and can be specifically implemented in a manner of hybrid wiring of a silicon optical integrated module and a copper cable, carries high-bandwidth low-latency data streams through an optical link, and processes control signals and short-distance communications through an electrical link. The exchange chip array 122 refers to a distributed exchange node cluster composed of multiple programmable exchange chips, and can be specifically implemented in a combination of an exchange chip supporting a PCIe Gen5 / CXL protocol and an optical exchange module, and is used to implement data routing and protocol conversion between different hardware units. The exchange control unit 123 refers to an embedded controller with topology calculation capability, and can be specifically implemented by using an FPGA to carry a dynamic routing algorithm, and is used to analyze a hardware configuration scheme and generate exchange chip control instructions.

[0103] In this embodiment, when the resource orchestrator 150 issues a hardware configuration scheme containing a target topology and bandwidth requirement, the exchange control unit 123 first analyzes the connection relationship parameters in the scheme, calculates a shortest path set satisfying the bandwidth constraint, and then matches the target communication protocol type through a preset protocol conversion table, and sends port mapping instructions and clock synchronization signals to the exchange chip array 122. The exchange chip array 122 dynamically switches the enable states of the electrical link and the optical interconnection link according to the instructions, and completes the physical connection reconstruction across the hardware units within a millisecond level time. The optical interconnection link in the optoelectronic hybrid interconnection structure 121 is preferentially used for large-scale tensor data transmission between GPU clusters, and the electrical link is responsible for metadata interaction between the storage unit and the control unit, thereby realizing on-demand bandwidth allocation between the heterogeneous computing units.

[0104] In this embodiment, the scheme cooperates the optoelectronic hybrid interconnection structure 121 and the programmable exchange chip, not only supports dynamic reconstruction of the physical layer connection topology, but also adaptively selects the optimal transmission medium according to the workload characteristics, breaks through the bandwidth bottleneck of the traditional interconnection architecture under the premise of maintaining high reliability, realizes real-time dynamic optimization of the connection topology between the hardware units, reduces the communication delay in the heterogeneous computing task, and reduces the energy waste of the high-power module through the shunt use of the optoelectronic medium. The scheme enables the CPU, GPU, storage unit and other heterogeneous resources to form an optimal connection combination according to the task requirement, and significantly improves the system throughput in the large-scale matrix operation and distributed storage scene.

[0105] In a feasible implementation manner, referring to Figure 4 The intelligent heat dissipation system 140 includes:

[0106] A temperature sensor network 141 is deployed on the surface of the computing units and storage units.

[0107] A partitioned liquid cooling unit 142 includes liquid cooling branches with independent flow control.

[0108] An air flow guiding device 143 includes an array of variable-speed fans.

[0109] A heat dissipation control unit 144 is connected to the temperature sensor network 141, the partitioned liquid cooling unit 142, and the air flow guiding device 143, and is configured to control the liquid cooling flow distribution and fan rotation speed based on the power consumption distribution prediction data.

[0110] In this embodiment, the temperature sensor network 141 refers to temperature sensing devices distributed on the surface of the hardware units, which can be implemented by thermocouples or infrared sensor arrays, for real-time monitoring of the temperature distribution on the surface of the hardware units. The partitioned liquid cooling unit 142 refers to cooling circuits with independent flow control functions, which can be implemented by a liquid cooling pipe branch structure with electromagnetic valves, and each branch can independently adjust the flow rate of the cooling liquid. The air flow guiding device 143 refers to a heat dissipation component with variable air volume output, which can be implemented by a PWM speed-regulated fan group cooperating with a guide plate structure, and can adjust the air flow path and intensity in a directional manner. The heat dissipation control unit 144 refers to a controller integrating data processing and execution instructions, which can be implemented by an embedded microprocessor combined with a driving circuit, and can generate dynamic heat dissipation strategies according to prediction data.

[0111] In this embodiment, the temperature sensor network 141 continuously collects temperature data on the surface of the computing units and storage units, and sends the data to the heat dissipation control unit 144. The heat dissipation control unit 144 analyzes the heating trends of different hardware modules in future time periods in combination with the power consumption distribution prediction data provided by the AI workload perception and prediction engine 130. When it is predicted that a specific computing unit will enter a high-power consumption state, the heat dissipation control unit 144 increases the flow rate of the corresponding region liquid cooling branch in advance, and adjusts the fan array rotation speed to enhance local air circulation. For example, when it is predicted that the GPU module will have a load surge in the next time period, the heat dissipation control unit 144 will preferentially increase the cooling liquid flow rate in the region where the module is located, and start the adjacent fan group for auxiliary heat dissipation. This proactive heat dissipation method based on prediction can avoid the lag problem existing in traditional responsive heat dissipation.

[0112] In this embodiment, the temperature sensor network 141 is combined with prediction data to achieve on-demand allocation of heat dissipation resources, such as reducing liquid cooling flow in low-load areas to reduce energy consumption, and enhancing heat dissipation capacity in high-load areas to avoid overheating. In addition, the combination of partitioned liquid cooling and adjustable fans breaks through the limitations of a single heat dissipation mode, and the optimal heat dissipation scheme can be selected according to the thermal characteristics of different hardware units. Through the above technical solutions, the application can realize dynamic optimization and allocation of heat dissipation resources according to hardware load prediction data, reduce the overall energy consumption of the system on the premise of ensuring heat dissipation efficiency. Independent control of the partitioned liquid cooling branch can avoid unnecessary cooling liquid circulation loss, and the adjustable fan array can adjust the local air flow intensity according to the real-time thermal distribution, thereby effectively solving the problem of excessive heat dissipation or insufficient heat dissipation in traditional heat dissipation systems.

[0113] In one possible implementation, the resource orchestrator 150 comprises: Figure 5

[0114] an intent analysis module 151 for converting resource intent request data into resource demand parameter data;

[0115] a resource assessment module 152 connected to the unified hardware abstraction layer 160 for obtaining current available hardware resource state data;

[0116] an optimization decision module 153 connected to the intent analysis module 151 and the AI workload awareness and prediction engine 130 for generating the hardware configuration scheme data based on the resource demand parameter data, resource demand prediction data, and available hardware resource state data.

[0117] In this embodiment, the intent analysis module 151 refers to a logical unit for converting abstract resource demands submitted by users into quantifiable parameters, which can be implemented by using a natural language processing engine or a rule engine, for eliminating semantic ambiguity in resource requests and extracting key demand indicators. The resource assessment module 152 refers to a component for monitoring the available state of hardware resources in real time, which can be implemented by calling the API interface provided by the unified hardware abstraction layer 160, for obtaining the working state and remaining capacity of each unit in the modular hardware pool. The optimization decision module 153 refers to an operation unit for generating resource allocation strategies based on multi-dimensional data, which can be implemented by using a linear programming algorithm or a heuristic algorithm, for balancing performance and energy consumption indicators while meeting resource demands.

[0118] ​In this embodiment, when the upper-layer application submits a resource intention request containing computing power requirements and storage requirements, the intention analysis module 151 first converts it into specific quantified parameters such as CPU core number, memory capacity, etc. The resource evaluation module 152 obtains the current available computing unit list and its residual computing power value, storage unit residual space, etc. through the unified hardware abstraction layer 160 in real time. The optimization decision module 153 combines the future resource demand trend data provided by the AI prediction engine, and calculates the hardware combination scheme that meets the current demand and adapts to future changes through the preset optimization algorithm. For example, when it is predicted that memory-intensive tasks will surge in the next three hours, the optimization decision module 153 preferentially selects hardware units with high memory bandwidth and reserves expansion capacity.

[0119] In this embodiment, by dynamically integrating real-time resource status and prediction data, the hardware configuration can be actively adjusted to adapt to load changes, and the number of reconfigurations is reduced and the response delay is reduced through predictive decision-making, realizing dynamic optimization configuration of hardware resources, and solving the problem of low utilization caused by rigid resource allocation in traditional servers. By converting abstract requirements into executable parameters and integrating prediction data, performance bottlenecks or energy waste caused by insufficient or redundant resource configuration are avoided. The optimization decision module 153 considers real-time status and future trends to ensure that the hardware combination scheme meets current task requirements and has expansion capability to cope with load fluctuations.

[0120] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the adaptive heterogeneous server system 100 based on AI prediction and dynamic hardware reconstruction of the present application. More forms of simple transformation based on this technical concept are within the protection scope of the present application.

[0121] The present application also provides a resource management method applied to the adaptive heterogeneous server system 100 based on AI prediction and dynamic hardware reconstruction, which refers to Figure 6 , and the resource management method comprises steps S100-S600, wherein:

[0122] Step S100, receiving resource intention request data of an upper-layer application through a resource orchestrator 150;

[0123] Step S200, collecting system telemetry data through an AI workload perception and prediction engine 130, performing cleaning and normalization processing, and generating state feature data;

[0124] Step S300, generating resource demand prediction data and power consumption distribution prediction data based on the state feature data through a preset prediction model;

[0125] At step S400, the resource orchestrator 150 generates hardware configuration scheme data based on the resource intention request data and the resource demand prediction data.

[0126] At step S500, the control reconfigurable interconnection Fabric 120 dynamically adjusts the connection topology, bandwidth and communication protocol between hardware units according to the hardware configuration scheme data.

[0127] At step S600, the control intelligent cooling system 140 adjusts the liquid cooling flow and fan speed according to the power consumption distribution prediction data.

[0128] In this embodiment, the resource intention request data refers to the abstract demand description of the upper layer application for hardware resources, which can be implemented by using a configuration file in JSON format or YAML format, and contains computing power demand parameters, video memory demand parameters and task priority information, which is used to guide the resource orchestrator 150 to generate an adaptive hardware configuration scheme. The system telemetry data refers to a set of original indicators reflecting the running state of the hardware, which can be collected by using the Prometheus monitoring tool to collect CPU utilization, memory bandwidth, storage IOPS and network throughput data, to provide input basis for the prediction model. The state feature data refers to the standardized data sequence after preprocessing, which can be processed by using a sliding window algorithm to denoise and normalize the original data, to eliminate the dimension difference and extract the time series correlation features. The preset prediction model refers to a prediction module based on a machine learning algorithm, which can be implemented by using a long short-term memory neural network model, to capture the periodicity and burstiness of the workload by using a memory cell state update mechanism. The hardware configuration scheme data refers to a set of deployment instructions of hardware resource combination, which can be implemented by using a TOSCA template description module to describe the activation instruction, the interconnection topology configuration parameter and the communication protocol switching command, to drive the reconfigurable interconnection Fabric 120 to perform dynamic adjustment.

[0129] In this embodiment, after the resource intention request data is received, the system telemetry data is acquired in real time by the data acquisition unit 131, and is cleaned and normalized by the feature processing unit 132 to form the state feature data. The data is input into the long short-term memory neural network model, and the resource demand and power consumption distribution prediction results in the future time window are calculated by using the memory cell state update mechanism. The resource orchestrator 150 generates the hardware configuration scheme containing the module activation instruction and the interconnection topology parameter in combination with the prediction results and the current available hardware resource state. The reconfigurable interconnection Fabric 120 switches the connection path of the electrical link and the optical interconnection link according to the scheme, adjusts the bandwidth allocation ratio and the communication protocol type. At the same time, the intelligent cooling system 140 adjusts the cooling liquid flow distribution by using the partitioned liquid cooling branch and controls the air volume distribution of the adjustable speed fan array based on the power consumption distribution prediction data, to realize the dynamic matching of the cooling resources.

[0130] It can be understood that the traditional resource scheduling method relies on static hardware configuration and passive response mechanism, and cannot predict the trend of workload change, resulting in lag of resource allocation and low heat dissipation efficiency. The method actively predicts resource demand and power consumption distribution through an AI prediction model, and realizes real-time matching of resource supply and demand in combination with a dynamic hardware reconstruction mechanism. For example, when the prediction model detects that GPU computing-intensive tasks will occur in the future, the resource arranger 150 activates the idle GPU module in advance and establishes a high-bandwidth optical interconnection link, while the traditional scheme needs to wait for the arrival of the task before starting scheduling, causing response delay. In addition, the heat dissipation system adjusts the liquid cooling flow in advance based on the prediction data to avoid local hot spots, while the traditional heat dissipation method only relies on current temperature feedback for lag adjustment. In this way, the present application can realize the on-demand dynamic combination of hardware resources, reduce the idle or performance bottleneck caused by configuration mismatch. At the same time, the heat dissipation system makes prospective adjustment based on prediction data to avoid the risk of frequency reduction or downtime caused by local overheating, and improves system stability. The method shortens the response time of resource scheduling, so that the hardware configuration adjustment is synchronized with the arrival of the task, and improves the processing efficiency of heterogeneous computing tasks.

[0131] In a possible implementation, with reference to Figure 7 , the step S300 includes steps S310-S340, wherein:

[0132] Step S310, extracting periodic features and burst features from state feature data;

[0133] Step S320, inputting the periodic features and burst features into a long short-term memory neural network;

[0134] Step S330, calculating the current time memory state through the memory unit of the long short-term memory neural network;

[0135] Step S340, generating resource demand prediction data and power consumption distribution prediction data for a future preset time period through an output gate.

[0136] In this embodiment, the periodic feature refers to a data pattern in which the workload presents regular fluctuations over time, which can be realized by Fourier transform or autocorrelation function analysis to identify the repetitive regularity of the computing resource demand. The bursty feature refers to a non-stationary data pattern caused by a burst event or unpredictable task, which can be realized by sliding window statistics or anomaly detection algorithm to capture the instantaneous change of the resource demand. The memory cell of the long short-term memory neural network refers to a recurrent neural network structure with a gating mechanism, which can realize long-term dependence modeling of time series information through input gate, forget gate and output gate, and is used to fuse historical state and current input features. The input gate state refers to a gating signal that controls the entry of new information into the memory cell, which can be calculated by combining sigmoid function and tanh function, and is used to filter the effective information in the current input features. The forget gate state refers to a gating signal that controls the retention degree of historical information, which can be calculated by sigmoid function, and is used to determine the decay ratio of historical information in the memory cell.

[0137] In this embodiment, the periodic feature and the bursty feature are simultaneously input into the long short-term memory neural network, and the input gate state and the forget gate state are dynamically adjusted based on the hidden state at the previous moment and the current input feature. The memory cell state selectively forgets historical information through the forget gate and updates the current effective information through the input gate. After the updated memory cell state is processed by the output gate, the resource demand prediction data and the corresponding power consumption distribution prediction data of each hardware unit in the future preset time period are generated. Therefore, the system can predict the trend of resource demand changes in advance according to the dynamic characteristics of the workload.

[0138] In this embodiment, through the gating mechanism of the long short-term memory neural network, the long-term dependence relationship in the time series feature and the short-term fluctuation feature can be adaptively captured, thereby improving the accuracy and robustness of the prediction results, so that the scheme can accurately predict the complex change pattern of resource demand in the heterogeneous computing scenario, and provide reliable data support for dynamic hardware reconstruction. The identification of the periodic feature can optimize the resource reservation strategy and reduce the risk of resource fragmentation; the capture of the bursty feature can shorten the response delay of abnormal load and avoid local hardware overload. At the same time, the joint output of the power consumption distribution prediction data and the resource demand prediction data enables the cooling system to adjust the cooling strategy in advance, realizing the collaborative optimization of energy consumption and performance.

[0139] In a feasible implementation, the step of calculating the memory state at the current moment by the memory cell of the long short-term memory neural network includes: calculating the input gate state based on the hidden state at the previous moment and the current input feature; calculating the forget gate state based on the hidden state at the previous moment and the current input feature; updating the memory cell state in combination with the forget gate state and the input gate state, and outputting the updated memory cell state.

[0140] In this embodiment, the input gate state refers to a weight parameter for controlling the update of the current input feature on the memory cell state, which can be specifically implemented by combining a sigmoid activation function with matrix multiplication operation, for filtering effective features and suppressing noise interference. The forget gate state refers to a weight parameter for controlling the retention degree of the memory cell state at the last time, which can be specifically implemented by combining a sigmoid function with element-wise multiplication, for dynamically adjusting the decay rate of historical information. The memory cell state refers to an intermediate variable for storing historical features and current features, which can be specifically implemented by linear transformation and hyperbolic tangent function, for carrying the time step-dependent relationship across time.

[0141] In this embodiment, in the resource demand prediction process, the input gate state determines which new features to add to the memory cell by analyzing the relevance of the current system load features and the historical state. The forget gate state decides to retain or discard part of the historical information by evaluating the relevance of the historical state and the current prediction target. The update process of the memory cell state generates an intermediate state containing complete time sequence information by superimposing new and old features after gating filtering. For example, when the system load presents periodic fluctuations, the forget gate can reduce the weight of the old state at irrelevant time points, and the input gate enhances the transmission of key features in the current period, thereby improving the dynamic adaptation capability of the prediction model.

[0142] In this embodiment, by dynamically adjusting the parameters of the input gate and the forget gate, the scheme can adaptively capture the long-term periodicity and short-term burstiness of the load data, avoiding prediction bias caused by fixed weights. For example, in the case of sudden high load, the input gate can quickly enhance the weight of the current feature, while the forget gate reduces the influence of old historical data, thereby achieving more accurate short-term prediction. In this way, the present application solves the prediction lag problem caused by the static model in the prior art, which cannot dynamically adjust the feature weight. By differentiating the historical information and the current input through the gating mechanism, the time sequence correlation of the resource demand prediction can be effectively improved, providing more accurate decision basis for the dynamic reconstruction of hardware resources, thereby reducing the risk of resource mismatch and improving the system response speed.

[0143] In a feasible implementation manner, referring to Figure 8 , the step S400 includes steps S410-S440, wherein:

[0144] Step S410, parsing the computing power demand parameter and the video memory demand parameter in the resource intent request data;

[0145] Step S420, obtaining the current available hardware resource state data through the unified hardware abstraction layer 160;

[0146] Step S430, based on the computing power requirement parameter, the video memory requirement parameter, the available hardware resource state data and the resource demand prediction data, the optimal hardware combination scheme is calculated;

[0147] Step S440, the hardware configuration scheme data containing the module activation instruction and the connection topology instruction is generated.

[0148] In this embodiment, the computing power requirement parameter refers to the quantitative index of the computing capacity of the upper application, which can be realized by floating point operation frequency or instruction throughput parameter, and is used to determine the minimum computing resource required for task execution. The video memory requirement parameter refers to the requirement of the application to the memory capacity of the graphics processor, which can be realized by the video memory capacity or the video memory bandwidth parameter, and is used to ensure the data storage requirement of the graphics processing task. The unified hardware abstraction layer 160 refers to the software interface layer shielding the difference of the bottom hardware, which can be realized by standardized API or device driver framework, and is used to obtain the running state and availability data of each unit in the modular hardware pool in real time. The optimal hardware combination scheme refers to the minimum resource consumption configuration meeting the computing power and video memory requirement, which can be realized by integer linear programming or heuristic algorithm, and is used to balance the performance and energy efficiency under the resource constraint condition. The module activation instruction refers to the instruction set for controlling the power-on state of the hardware unit, which can be realized by power management protocol or hot plug control signal, and is used to dynamically enable or disable the specific computing or storage module. The connection topology instruction refers to the configuration parameter for defining the physical connection relationship between the hardware units, which can be realized by link routing table or switch matrix configuration parameter, and is used to establish a communication path meeting the bandwidth and delay requirement.

[0149] In this embodiment, when the upper application submits a resource request containing the computing power and video memory requirement, the resource orchestrator 150 first extracts the quantitative index from the request, for example, it needs to complete 100 billion floating point operations per second and is equipped with 16GB video memory. Then, through the unified hardware abstraction layer 160, the idle or releasable modules in the current hardware pool are queried, such as available GPU acceleration card, FPGA computing unit or high bandwidth memory module. Based on the real-time resource state and the future load trend provided by the prediction model, the optimal algorithm is used to generate a hardware combination meeting the current and expected requirement, for example, two GPU cards equipped with 8GB video memory are selected for parallel computing, and the data channel is established through the high-speed interconnection link. The finally generated configuration scheme contains the module activation instruction, such as waking up the GPU unit in the low power consumption state, and the connection topology instruction, such as adjusting the communication bandwidth between the selected GPU and the computing node to PCIe 4.0x16 mode.

[0150] In some embodiments, the computing power requirement parameter can be further refined into single-precision floating-point performance or mixed-precision computing capability indicators, and the graphics memory requirement parameter can include a graphics memory bandwidth threshold or a cache hit rate requirement. The unified hardware abstraction layer 160 can dynamically update the real-time state of each module in the hardware pool through a device tree structure or a resource description framework. The calculation process of the optimal hardware combination scheme can introduce a multi-objective optimization algorithm, such as minimizing power consumption or heat dissipation cost while meeting the computing power requirement. The module activation instruction can be combined with a power management strategy, such as performing a soft shutdown operation on a hardware unit that has been idle for more than a preset time. The connection topology instruction can include link redundancy configuration, such as establishing a backup channel for a critical data transmission path.

[0151] In this embodiment, the scheme actively plans hardware combinations and dynamically adjusts connection topologies by integrating real-time resource states and prediction data, such as reserving expansion slots in advance when predicting an increase in future graphics memory requirements. The scheme also realizes flexible combination of hardware resources through module activation and topology reconstruction, such as dynamically switching between direct connection or bridging mode between GPUs and storage units according to task characteristics, solving the problem of low resource utilization caused by fixed hardware configuration in traditional servers, and realizing dynamic resource combination based on real-time requirements and prediction data. By accurately matching computing power and graphics memory requirements, the performance bottleneck caused by insufficient resource allocation or energy waste caused by excessive configuration is avoided. The module activation and topology reconstruction mechanism enables hardware resources to be combined on demand, effectively supporting flexible deployment of diverse heterogeneous computing tasks.

[0152] In a feasible implementation, step S430 includes steps S431-S434, wherein:

[0153] Step S431, set the optimization objective function with performance priority or energy efficiency priority;

[0154] Step S432, calculate a module combination scheme that meets the computing power requirement parameter and the graphics memory requirement parameter;

[0155] Step S433, determine the bandwidth allocation ratio between the selected modules;

[0156] Step S434, verify that the module combination scheme meets the constraint conditions of the resource intent request, and output the verified hardware combination scheme.

[0157] In this embodiment, the optimization objective function refers to a mathematical expression used to balance system performance and energy consumption. Specifically, the computational performance indicators and power consumption indicators can be combined using a linear weighting method, such as setting the task processing delay and the energy consumption per unit computing power as optimization variables, and adjusting the weight coefficients to achieve different optimization strategies. The module combination scheme refers to selecting a set of physical devices from the hardware resource pool that meet the computing requirements. Specifically, the computing units and storage units that meet the computing power and memory requirements can be selected through a greedy algorithm or an integer programming model. The bandwidth allocation ratio refers to the bandwidth occupancy ratio of the communication channels between modules in the interconnection structure. Specifically, a dynamic bandwidth allocation protocol can be used to adjust the link resources according to the task data flow characteristics, such as allocating more optical interconnection channels for high-throughput tasks. The constraint condition verification refers to checking whether the hardware combination meets the physical resource limitations and quality of service requirements. Specifically, the availability of the selected modules can be ensured through a resource reservation mechanism, and the communication delay can be tested to ensure that it is below the preset threshold.

[0158] In this embodiment, when receiving the resource demand prediction data, first, the target function type is selected according to the optimization strategy defined by the upper layer application. If the performance priority mode is selected, the core goal is to improve the task processing speed, and high-frequency computing units are activated and low-delay interconnection links are allocated first. If the energy efficiency priority mode is selected, the hardware combination with the optimal energy efficiency ratio is selected by balancing the computing unit power consumption and the heat dissipation energy consumption. Subsequently, based on the current available hardware resource state data, all module combinations that meet the computing power and memory requirements are traversed, and the score of each combination under the target function is calculated. After determining the preliminary scheme, bandwidth resources are allocated for data transmission between modules according to the task communication mode, such as allocating high-bandwidth optical links for CPU-GPU combinations that need to exchange data frequently. Finally, it is verified whether the selected scheme meets the constraint conditions in the resource intent request, including the online state of the hardware unit, the capacity limitation of the heat dissipation system, and the maximum load capacity of the interconnection structure, to ensure the executability of the scheme.

[0159] It can be understood that the traditional resource scheduling method generally only performs static allocation based on the current resource surplus, and cannot dynamically adjust the hardware combination and interconnection topology. For example, when facing a sudden AI inference task, the existing scheme may refuse to serve due to insufficient GPU resources, while the present scheme activates the standby computing unit in advance through the prediction model, and reconfigures the interconnection link to realize resource expansion. In addition, the existing technology lacks dynamic optimization of bandwidth allocation, which can easily cause interconnection bottlenecks, while the present scheme effectively improves the communication efficiency between heterogeneous modules by adjusting the proportion of electrical links and optical links in real time.

[0160] Through the technical solution, the application can automatically select the optimal hardware configuration according to the workload characteristics, avoid over-provisioning of resources while meeting performance requirements. Through dynamic bandwidth allocation and constraint verification, the feasibility of the hardware combination scheme under the physical limitations of heat dissipation capacity, interconnection load and the like is ensured, thereby improving the overall utilization of heterogeneous computing resources. Compared with the traditional static resource configuration mode, the scheme significantly enhances the adaptability of the system to diversified and dynamic workloads.

[0161] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0162] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. An adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration, characterized in that, The system includes: Modular heterogeneous hardware unit pool, including online hot-swappable computing units, storage units and network interface units; The reconfigurable interconnect fabric connects the modular heterogeneous hardware unit pool; The AI ​​workload perception and prediction engine is connected to the modular heterogeneous hardware unit pool to collect system telemetry data and generate resource demand prediction data and power consumption distribution prediction data. The intelligent heat dissipation system connects the AI ​​workload perception and prediction engine and the modular heterogeneous hardware unit pool; The resource orchestrator is connected to the AI ​​workload perception and prediction engine, the reconfigurable interconnect fabric, and the intelligent cooling system, respectively. A unified hardware abstraction layer connects the resource orchestrator and the modular heterogeneous hardware unit pool; The resource orchestrator is used to receive resource intent request data from upper-layer applications and generate hardware configuration scheme data based on the resource demand prediction data, including: Parse the computing power requirement parameters and video memory requirement parameters in the resource intent request data; Obtain the status data of currently available hardware resources through the unified hardware abstraction layer; Based on computing power demand parameters, video memory demand parameters, available hardware resource status data, and resource demand prediction data, calculate the optimal hardware combination scheme; Generate hardware configuration scheme data containing module activation instructions and connection topology instructions; The steps for calculating the optimal hardware combination scheme based on computing power requirement parameters, video memory requirement parameters, available hardware resource status data, and resource requirement prediction data include: Set an optimization objective function that prioritizes performance or energy efficiency; Calculate the module combination scheme that meets the computing power and video memory requirements; Determine the bandwidth allocation ratio among the selected modules; The verification module combination scheme satisfies the constraints of the resource intent request, and the verified hardware combination scheme is output. The reconfigurable interconnect fabric is used to dynamically adjust the connection topology, bandwidth and communication protocol between hardware units according to the hardware configuration scheme data. The intelligent heat dissipation system is used to perform heat dissipation control on the hardware unit based on the power consumption distribution prediction data.

2. The adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration as described in claim 1, characterized in that, The AI ​​workload perception and prediction engine includes: The data acquisition unit is used to collect data on CPU utilization, memory bandwidth, storage IOPS, and network throughput from the operating system and hardware sensors. The feature processing unit, connected to the data acquisition unit, is used to perform cleaning, noise reduction and normalization processing on the data acquired by the data acquisition unit to generate state feature data. The prediction model unit, connected to the feature processing unit, is used to process the state feature data using a long short-term memory neural network and output the resource demand prediction data and power consumption distribution prediction data.

3. The adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration as described in claim 1, characterized in that, The reconfigurable interconnect fabric includes: A hybrid optoelectronic interconnect structure, comprising electrical links and optical interconnect links; A switching chip array connects the electrical link and the optical interconnect link; A switching control unit, connected to the switching chip array and the resource orchestrator, is used to adjust the connection path within a preset time according to the hardware configuration scheme data.

4. The adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration as described in claim 1, characterized in that, The intelligent heat dissipation system includes: A temperature sensor network is deployed on the surfaces of computing and storage units; The partitioned liquid cooling unit includes independently adjustable liquid cooling branches; An airflow guiding device, comprising an adjustable-speed fan array; The heat dissipation control unit, connected to the temperature sensor network, the partitioned liquid cooling unit, and the airflow guiding device, is used to control the liquid cooling flow distribution and fan speed based on the power consumption distribution prediction data.

5. The adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration as described in claim 1, characterized in that, The resource orchestrator includes: The intent parsing module is used to convert resource intent request data into resource requirement parameter data; The resource assessment module, connected to the unified hardware abstraction layer, is used to obtain the status data of currently available hardware resources; The optimization decision module connects the intent parsing module and the AI ​​workload perception and prediction engine, and is used to generate the hardware configuration scheme data based on the resource requirement parameter data, resource requirement prediction data and available hardware resource status data.

6. A resource management method, applied to an adaptive heterogeneous server system based on AI prediction and dynamic hardware reconfiguration as described in any one of claims 1 to 5, characterized in that, The resource management method includes: Receive resource intent request data from upper-layer applications through the resource orchestrator; The system collects telemetry data through an AI workload perception and prediction engine, performs cleaning and normalization processing, and generates state feature data. Based on the state characteristic data, resource demand prediction data and power consumption distribution prediction data are generated through a preset prediction model. The resource orchestrator generates hardware configuration scheme data based on the resource intent request data and resource demand prediction data; The Control Reconfigurable Interconnect Fabric dynamically adjusts the connection topology, bandwidth, and communication protocols between hardware units based on the hardware configuration scheme data. The intelligent cooling system adjusts the liquid cooling flow rate and fan speed based on the power consumption distribution prediction data.

7. The resource management method as described in claim 6, characterized in that, The steps of generating resource demand prediction data and power consumption distribution prediction data based on the state feature data and using a preset prediction model include: Extract periodic and burst features from state feature data; Input periodic features and burst features into a long short-term memory neural network; The memory state at the current moment is calculated using the memory units of a long short-term memory neural network; The output gate generates resource demand forecast data and power consumption distribution forecast data for a preset time period in the future.

8. The resource management method as described in claim 7, characterized in that, The step of calculating the current memory state using the memory units of a long short-term memory neural network includes: Calculate the input gate state based on the hidden state of the previous time step and the current input features; Calculate the forget gate state based on the hidden state of the previous time step and the current input features; Update the memory cell state by combining the forget gate state and the input gate state, and output the updated memory cell state.

9. The resource management method as described in claim 6, characterized in that, The step of the resource orchestrator generating hardware configuration scheme data based on the resource intent request data and resource demand prediction data includes: Parse the computing power requirement parameters and video memory requirement parameters in the resource intent request data; Obtain the status data of currently available hardware resources through the unified hardware abstraction layer; Based on computing power demand parameters, video memory demand parameters, available hardware resource status data, and resource demand prediction data, calculate the optimal hardware combination scheme; Generate hardware configuration scheme data containing module activation instructions and connection topology instructions.

10. The resource management method as described in claim 9, characterized in that, The steps for calculating the optimal hardware combination scheme based on computing power requirement parameters, video memory requirement parameters, available hardware resource status data, and resource requirement prediction data include: Set an optimization objective function that prioritizes performance or energy efficiency; Calculate the module combination scheme that meets the computing power and video memory requirements; Determine the bandwidth allocation ratio among the selected modules; The verification module combination scheme satisfies the constraints of the resource intent request, and outputs the verified hardware combination scheme.

Citation Information

Patent Citations

  • Instruction optimization scheduling method based on large model and related device

    CN119473560A

  • GPU heterogeneous cluster scheduling method and system oriented to large model training and reasoning

    CN120448134A