Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

143 results about "GPU cluster" patented technology

A GPU cluster is a computer cluster in which each node is equipped with a Graphics Processing Unit (GPU). By harnessing the computational power of modern GPUs via General-Purpose Computing on Graphics Processing Units (GPGPU), very fast calculations can be performed with a GPU cluster.

GPU heterogeneous cluster scheduling method and system oriented to large model training and reasoning

The invention relates to the technical field of cluster scheduling, and provides a GPU heterogeneous cluster scheduling method and system oriented to large model training and reasoning, which constructs a set of complete cluster scheduling system by integrating multi-source information such as hardware features, running states and historical task data and applying technologies such as a clustering algorithm, a fuzzy comprehensive evaluation method and reinforcement learning. Comprehensive, intelligent and dynamic management and scheduling of GPU cluster resources are realized, the cluster scheduling system can significantly improve the execution efficiency of GPU heterogeneous clusters in large model training and reasoning tasks, the resource utilization rate is improved, the energy consumption is reduced, and the stability and adaptability of the system are enhanced. And an efficient and reliable solution is provided for large-scale deep learning application.
Owner:NEWLIXON TECH CO LTD

GPU computing power resource scheduling method and system

The invention relates to the technical field of data analysis, and discloses a GPU computing power resource scheduling method and system, and the method comprises the steps: collecting node hardware parameters and dynamic load indexes of a GPU cluster to construct a multi-dimensional resource feature vector of the GPU cluster, and constructing a resource portrait of the GPU cluster; establishing a node health degree scoring model of the GPU cluster, and generating a health degree score of a cluster node corresponding to the GPU cluster; analyzing a video memory demand of the GPU task request, and calculating an intensive identifier and a communication dependency relationship; determining the SLA weight of the GPU task request, calculating the resource shortage sensitivity of the GPU task request based on the video memory demand, and calculating the target task priority of the GPU task request in combination with the SLA weight; and determining a resource scheduling node group requested by the GPU task in the resource portrait, generating resource scheduling parameters of the resource scheduling node group, and executing scheduling of computing power resources of the GPU cluster based on the resource scheduling parameters. According to the method, the scheduling efficiency of the GPU computing power resources can be improved.
Owner:SHENZHEN DIXI YUNLIAN TECH CO LTD

Heterogeneous GPU resource management scheduling method

The invention provides a heterogeneous GPU resource management scheduling method, and relates to the technical field of GPU resource allocation, heterogeneous equipment management and unified abstract modeling are carried out, GPU resources of different architectures are registered to a container arrangement platform, and a unified abstract layer is constructed to shield bottom layer hardware differences; gPU cluster optimization management based on a multi-dimensional real-time monitoring and intelligent scheduling strategy is carried out, GPU operation indexes are collected, priorities are dynamically calibrated for tasks, and task performance portraits are constructed; scheduling decision making is carried out through multi-strategy cooperation, and optimal GPU resources are distributed for tasks; carrying out fine-grained resource allocation, carrying out space or time segmentation on the GPU, and dynamically adjusting resource allocation according to a load state; aPI conversion of cross-architecture tasks is realized through a unified runtime library, and task execution data is collected to feed back an optimization scheduling model; automatic detection, isolation and task migration of GPU faults are carried out, and unified monitoring and alarm are provided.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Large model training-oriented GPU (Graphics Processing Unit) cluster computing power optimization architecture

A GPU cluster computing power optimization architecture oriented to large model training is characterized in that a computing power resource management module is used for monitoring and managing the computing power resource state of each GPU in a GPU cluster in real time, a task scheduling module is used for allocating GPU resources according to the requirements of training tasks, and a computing power optimization module is used for dynamically optimizing the computing power of the GPU cluster in the task execution process. And the result feedback module is used for collecting and analyzing the optimized computing power use condition so as to adjust a subsequent task scheduling strategy. According to the method, the computing power resource state of the GPU cluster is monitored and managed in real time, the GPU resources are dynamically allocated according to the demand of the training task, and the computing power of the GPU cluster is dynamically optimized in the task execution process, so that the computing power utilization efficiency of the GPU cluster is effectively improved, and the training cost is reduced.
Owner:XIANGTAN UNIV

GPU cluster scheduling strategy optimization system based on deep reinforcement learning

The invention discloses a GPU cluster scheduling strategy optimization system based on deep reinforcement learning, which comprises a simulation environment layer, a reinforcement learning layer and a strategy evaluation layer, and is characterized in that the simulation environment layer is used for providing a real and credible GPU cluster scheduling environment, reproducing a core mechanism of actual cluster scheduling and simultaneously providing controllable experiment conditions; comparative evaluation and iterative optimization of strategies are facilitated; the reinforcement learning layer converts a GPU cluster scheduling problem into a Markov decision process based on a deep neural network, and obtains an optimal scheduling strategy by applying a reinforcement learning strategy; and the strategy evaluation layer counts each key index, compares and analyzes a plurality of scheduling strategies, and visually displays an analysis result to realize optimization of the GPU cluster scheduling strategy. The system solves the time sequence short view problem of a traditional scheduler, so that the scheduling decision can consider the influence on the future, global optimization instead of local optimization is realized, and the overall resource utilization efficiency is improved.
Owner:UNIV OF SCI & TECH OF CHINA

Cloud platform-based computing power resource dynamic scheduling and monitoring method

The invention relates to a computing power resource dynamic scheduling and monitoring method based on a cloud platform. The method is suitable for an intelligent scheduling scene of a high-performance GPU cluster. The method comprises seven steps of task portrait modeling, GPU node state acquisition, resource trend prediction, SLA tracking, scheduling scoring and deployment, operation monitoring and task migration, and SLA feedback optimization. According to the system, task semantics are represented by constructing task vectors, node health states, topology affinity, SLA historical performance conditions and resource prediction risks are fused, a multi-factor adjustable scheduling scoring mechanism is constructed, and second-level perception and task thermal migration of high-temperature nodes are achieved. Compared with a traditional Kubernetes static scheduling scheme, the method has the advantages that the GPU utilization rate, the task SLA achievement rate and the system stability are remarkably improved, the learning ability, the self-adaptive ability and the high availability are achieved, and the method is an intelligent scheduling closed-loop system oriented to AI reasoning and training scenes.
Owner:北京娱广科技有限公司

Computing power resource fusion method based on distributed flow pipeline

The invention relates to the technical field of distributed computing and computing power scheduling, in particular to a computing power resource fusion method based on a distributed flow pipeline, which comprises the following steps of: disassembling a user computing power request into a flow pipeline unit for packaging computing logic, an input / output interface and a resource demand label; meanwhile, computing power types, real-time load rates, memory occupancy rates, network round-trip delays, geographic positioning and energy consumption data of cloud edge end nodes are collected, a hierarchical topology network framework is constructed based on the collected data, node computing power available values are calculated, an inter-node transmission cost matrix is generated, a fault probability prediction model is constructed, and a fault probability prediction model is constructed. And finally, analyzing a data dependency relationship of the pipeline unit through a four-dimensional joint decision engine, executing dynamic mapping, and preferentially mapping the high-computing-power demand unit to a GPU cluster node, so as to realize non-interruption reconstruction during pipeline topology operation. And the global computing power resource utilization rate, the task operation efficiency and the service stability are improved.
Owner:LANZHOU YUNFAN ZHILIAN TECH CO LTD +2

Cooling device special for computer GPU cluster server

The invention discloses a special cooling device for a computer GPU (Graphic Processing Unit) cluster server, which comprises a box body, a support frame, a circulating pipe, a cooling box, a first piston cylinder, a liquid outlet pipe, an injection pipe, a liquid inlet pipe, a heat dissipation mechanism, a cooling mechanism, a detection mechanism and an auxiliary mechanism, meanwhile, a filter plate can prevent dust in air from entering, a detection mechanism monitors the temperature in the box body in real time by utilizing the thermal expansion characteristic of alcohol, and a signal is transmitted to an external controller through a pressure sensor so that a worker can master the operation temperature state of the server in real time; the cooling mechanism automatically adjusts the flow speed and flow of cooling liquid according to temperature changes so as to meet the heat dissipation requirements of the server under different loads, the auxiliary mechanism conducts rolling sweeping on a filter plate through a cleaning roller, smooth air circulation is ensured, the efficient and intelligent cooling effect is achieved, and the service life of the server is prolonged. And the operation stability and the service life of the GPU cluster server are improved.
Owner:GUANGDONG FENGFUZHILENG TECHNOLOGY CO LTD

Multi-core GPU interconnection architecture and self-adaptive cache allocation method thereof

The invention belongs to the technical field of semiconductors, and discloses a multi-core GPU interconnection architecture and a self-adaptive cache allocation method thereof.The multi-core GPU interconnection architecture comprises a plurality of GPU clusters and a global cluster, the global cluster comprises a global router, a global scheduler, a directory, a memory and a plurality of GPU clusters, each GPU cluster is composed of a plurality of GPU cores, and the GPU cores are arranged in the global router. Each GPU core grain comprises a plurality of independent primary data caches, secondary data caches, a local router and a network interface, and the network interface judges whether a current request is processed by the GPU core grains in a GPU cluster or not according to whether a target address in a target node belongs to the GPU cluster or not, so that a high-level multi-core grain interconnection architecture supporting cross-cluster communication is formed. According to the method, performance is optimized by adapting to different access modes, private and shared cache modes are switched in real time, single data access delay is reduced, the overall system efficiency is improved, and the method is used for a multi-core-particle heterogeneous system with a plurality of GPUs.
Owner:NANJING UNIV OF POSTS & TELECOMM

Server GPU (Graphics Processing Unit) computing power distribution method and system and server

The invention provides a server GPU computing power allocation method and system and a server, belongs to the field of server resource management, and solves the problems of low utilization rate and inflexible decision of traditional allocation resources. The method comprises the steps that a central agent constructs a directed weighted game diagram of tasks and a GPU cluster, and a global allocation strategy is solved based on Nash equilibrium; a local agent distributes tasks to a specific GPU through reinforcement learning, intelligent migration is achieved in combination with a dynamic threshold value and anomaly detection, and a strategy is iteratively optimized through federal learning. The system comprises a central agent, a local agent, a GPU cluster and an experience pool module, and a server carries the system execution method. According to the scheme, through double-layer agent cooperation and multi-algorithm fusion, the task cluster adaptation relation is quantified, the allocation strategy is dynamically adjusted, the resource utilization rate and the task processing efficiency are remarkably improved, and the adaptivity and reliability in a complex scene are enhanced.
Owner:TIANJIN LINYUE INTELLIGENT MANUFACTURING CO LTD

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

Security sensing method and system for GPU (Graphic Processing Unit) cluster

The invention relates to the technical field of network security awareness, in particular to a security awareness method and system for a GPU cluster, and the method comprises the steps: collecting multi-source security situation data, and providing basic information for subsequent security analysis; by constructing and updating a cluster security situation map in real time, information timeliness is updated and ensured in real time, and powerful support is provided for risk assessment; a dynamic risk assessment model is utilized to calculate vertex and edge scores, high-risk vertexes and key risk conduction paths are identified, and a basis is provided for resource isolation and scheduling; the targeted instruction is generated according to the security risk score and the preset security policy library, different security risks are effectively handled, and safe and stable operation of the cluster is guaranteed; by converting the instruction into a specific control command and distributing and executing the specific control command, effective implementation of security measures is ensured, and the cluster security protection capability is improved; by setting a finite-state machine monitoring instruction execution state, a recovery mechanism is automatically triggered, and continuous and safe operation of the cluster is guaranteed.
Owner:中科云达(北京)科技有限公司

High-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) coprocessing architecture of audio frequency integrated signal processor

The invention discloses a high-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) cooperative processing architecture of an audio frequency integrated signal processor, belonging to the technical field of audio frequency integrated signal processing, the architecture is based on a dynamic priority scheduling model and realizes efficient collaboration of a CPU and a GPU through hierarchical resource management and protocol level optimization, a hardware layer adopts a multi-GPU cluster and a distributed storage node, and the CPU-GPU cooperative processing architecture has the advantages that the multi-GPU cluster and the distributed storage node are integrated; high-concurrency task processing is supported; the transmission layer fuses RapidIO and an Ethernet protocol, and adapts to a short frame control signal and a long packet data stream respectively; and the application layer calculates task resource demands through dynamic priority weights, and ensures low time delay of key tasks in combination with a normalized allocation algorithm. According to the architecture, in an audio signal processing scene, the task scheduling efficiency is improved by 40%, the short frame transmission delay is as low as 0.5, the throughput of a long data stream reaches 100 Gbps, and the requirements for real-time performance and calculation precision in a complex acoustic environment can be met.
Owner:CHINA SHIP DEV & DESIGN CENT +1

Depth learning task node allocation method and system for executing time-aware computing power network heterogeneous GPU (Graphics Processing Unit) cluster

The invention discloses an execution time aware computing power network heterogeneous GPU cluster deep learning task node allocation method and system. The method comprises the following steps: firstly, based on a deep learning task, extracting and preprocessing task features and available node features; secondly, a sampler equally divides new tasks without historical data to available nodes, and each node performs mixed sampling on the tasks until all the tasks estimate execution time data; taking execution time data as a training set, taking the task features and the node features as a test set, and using a regression decision tree model to predict the execution time of the task on each node; performing task allocation on each node by using a cost search algorithm and a short job total JCT priority strategy; and finally, periodically monitoring node resources released in the cluster to obtain an optimal node allocation result. According to the method, task delay and total task JCT are remarkably reduced, cluster node resource changes are monitored in real time, and the resource utilization rate is increased.
Owner:HANGZHOU DIANZI UNIV +1

GPU cluster, redundancy optimization method of GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a GPU cluster, a redundancy optimization method of the GPU cluster, electronic equipment, a storage medium and a computer program product, the GPU cluster comprises a plurality of GPU cabinets, and each GPU cabinet comprises a plurality of GPU nodes, a GPU extension frame and an interconnection module; for any GPU cabinet, each GPU node in the GPU cabinet comprises a plurality of physical GPUs, and a GPU expansion frame in the GPU cabinet comprises a plurality of standby virtual GPUs; the interconnection module in the GPU cabinet is used for carrying out GPU interconnection on a plurality of GPU nodes and a plurality of GPU expansion frames in the GPU cabinet; and the plurality of standby virtual GPUs in the GPU extension frame in the GPU cabinet are used for providing redundant computing resources for the plurality of physical GPUs in each GPU node in the GPU cabinet. According to the embodiment of the invention, the GPU card-level redundancy capability can be effectively realized.
Owner:MOORE THREADS TECH CO LTD

Model training method and device based on heterogeneous GPU cluster and storage medium

The invention relates to a model training method and device based on a heterogeneous GPU cluster and a storage medium. The method comprises the steps of obtaining hardware index data of each GPU in the heterogeneous GPU cluster; obtaining a plurality of operation types corresponding to the deep learning algorithm in the to-be-trained model, and measuring operation performance data of operation corresponding to each operation type executed by each GPU; acquiring communication bandwidth multidimensional features and a load state sensing strategy of each GPU; according to the hardware index data, the operation performance data, the communication bandwidth multi-dimensional features and the load state sensing strategy corresponding to each GPU, constructing a multi-dimensional operation performance matrix of each GPU; according to the multi-dimensional operation performance matrix corresponding to each GPU, the GPU is allocated to each structure layer in the to-be-trained model for model training to obtain a target model, the training efficiency of the model is improved, and the training time of the model is shortened.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Image processing unit (GPU) cluster communication method and electronic device

The embodiment of the invention provides an image processing unit (GPU) cluster communication method and an electronic device, and the method comprises the steps: splitting a GPU cluster, and obtaining at least one algorithm isomorphic group and at least one algorithm heterogeneous group; the protocol distribution communication and the full collection communication are performed through the isomorphic GPUs in each algorithm isomorphic group, and the full protocol communication is performed through the heterogeneous GPUs in each algorithm heterogeneous group. Therefore, through the embodiment of the invention, the problems that the communication traffic between heterogeneous GPUs is increased and extra data copy overhead is introduced due to the difference of communication libraries between different GPUs in a heterogeneous GPU cluster environment can be solved.
Owner:ZTE CORP

Asymmetric segmentation scheduling system and method in heterogeneous GPU cluster

The invention provides an asymmetric segmentation scheduling system and method in a heterogeneous GPU cluster, and the method comprises the steps: S1, checking the features of an incoming request, and grouping the incoming request into different request buckets according to the token length; s2, in a heterogeneous GPU cluster environment, optimizing a large language model reasoning instance by adopting a double-layer strategy; and S3, calculating the matching degree of the request buckets and the big language model reasoning instances, and scheduling different request buckets to the big language model reasoning instance with the highest matching degree according to a calculation result. According to the method provided by the invention, the model layer can be asymmetrically segmented according to the computing power and the video memory capacity of each GPU on the premise of satisfying the model parallelism degree constraint, the dynamic balance of the execution duration between stages is realized, assembly line cavitation bubbles are fundamentally reduced, and the overall throughput rate is improved.
Owner:SHANGHAI JIAOTONG UNIV

GPU (Graphics Processing Unit) program optimization method for parallel environment

The invention provides a GPU program optimization method in a parallel environment, a plurality of target models are configured in a GPU cluster for parallel training, and the method comprises the following steps: performing strategy configuration and memory bottleneck pre-judgment on parallel training configuration of the target models; carrying out throughput optimization through parallel strategy combination and instruction-level performance monitoring by utilizing a pre-judgment result; kernel re-optimization is carried out on an instruction level bottleneck appearing in the optimization process, so that the model training efficiency and the GPU instruction execution efficiency are synchronously improved.
Owner:无锡九方科技有限公司

GPU resource intelligent scheduling method and system based on reinforcement learning

The invention provides an intelligent GPU resource scheduling method and system based on reinforcement learning, and relates to the field of cloud computing resource scheduling, and the method comprises the steps: collecting resource state data of a GPU cluster and attribute data of a to-be-scheduled task, carrying out the preprocessing, building a reinforcement learning model, and training the reinforcement learning model through historical scheduling records; inputting the resource state data of the current GPU cluster and the task attribute data of the to-be-scheduled task into the trained reinforcement learning model to obtain a preliminary scheduling scheme; generating a candidate scheduling scheme set based on the preliminary scheduling scheme, and selecting a final scheduling scheme from the candidate scheduling scheme set by adopting a multi-objective optimization algorithm; performing resource slice configuration on the target GPU node in the GPU cluster according to the final scheduling scheme, and starting the to-be-scheduled task in the task execution container; and updating reinforcement learning model parameters by using empirical data. According to the invention, adaptive evolution and multi-target balance optimization of the scheduling strategy are realized.
Owner:WUHAN SINGLE CLOUD NETWORK TECH CO LTD

Fault network card isolation method and device applied to GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a fault network card isolation method and device applied to a GPU cluster, electronic equipment, a storage medium and a computer program product, and the method comprises the steps: determining whether a fault network card exists on any GPU node in the GPU cluster or not according to the state information of each network card in the GPU node; when it is determined that fault network cards exist on the GPU node and the number of the fault network cards is smaller than or equal to a preset threshold value, an NCCL topological file of the GPU node is updated, and the NCCL topological file of the GPU node is used for indicating the topological relation between available GPU devices and available network cards included in the GPU node; and according to the NCCL topology file of the GPU node, isolating a fault network card in the GPU node when carrying out load scheduling on the GPU node. According to the embodiment of the invention, the fault isolation granularity can be effectively refined to a single network card in the GPU cluster.
Owner:MOORE THREADS TECH CO LTD

GPU hardware health prediction and active maintenance method and system

The invention provides a GPU (Graphics Processing Unit) hardware health prediction and active maintenance method and system, and belongs to the technical field of computing hardware maintenance and fault prediction. Based on the GPU health state data, a health degree comprehensive score of the GPU is calculated by using a preset health degree evaluation model, and a key health index in a future set time period is predicted by using a time sequence model. Calculating the fault probability of a preset type of fault risk by using a preset fault risk model based on the health degree comprehensive score of the GPU and / or the predicted key health index; and comparing the calculated fault risk probability with a preset threshold value, and executing a main preset maintenance action when the fault risk probability exceeds the preset threshold value. Through multi-source data acquisition and time sequence prediction, the potential fault risk of the GPU is found in advance, the maintenance action is actively executed, task interruption and data loss are reduced, and the reliability and availability of a GPU cluster are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Large model reasoning method for operator-level distributed scheduling

The invention discloses a large model reasoning method for operator-level distributed scheduling. The large model reasoning method comprises the steps of S1, operator decoupling and dependency modeling; s2, carrying out heterogeneous hardware capability portraying; and S3, dynamically adjusting the strategy. Operator-level fine-grained scheduling is realized in large model reasoning, and the computing power utilization rate of the heterogeneous GPU cluster is maximized.
Owner:BEIJING MOMENT UNLIMITED TECHNOLOGY CO LTD

Real-time fault tolerance and task migration system and method for process GPU (Graphics Processing Unit)

The invention provides a process GPU-oriented real-time fault tolerance and task migration system and method, and relates to the technical field of distributed computation.The method comprises the steps that a training task state in a GPU video memory is captured, difference data is generated through an incremental difference algorithm, and the difference data is stored in a multi-level storage system after being subjected to lightweight compression; meanwhile, GPU node health indexes are collected in real time, node states are analyzed through a time sequence prediction model, and a migration process is triggered when fault risks are predicted or preemption early warning is received; during migration, the task is paused, a final increment check point is generated, check point data is quickly transmitted to a target healthy GPU node through an RDMA network, a complete context state is reconstructed, task execution is recovered, and real-time fault tolerance and seamless migration of the GPU task are achieved. According to the method, the continuous operation reliability and the resource utilization efficiency of the GPU cluster in the long-time training task are improved.
Owner:HANHOU (BEIJING) TECH CO LTD

Power grid flexibility potential prediction and evaluation system based on deep learning

The invention discloses a power grid flexibility potential prediction and evaluation system based on deep learning, and belongs to the technical field of power system dispatching and artificial intelligence crossing. The system comprises a hardware layer, a software layer and an application layer, wherein the hardware layer realizes data acquisition and calculation support through a multi-source terminal, an edge node and a GPU cluster; in the software layer, the data preprocessing subsystem processes multi-source heterogeneous data by adopting an adaptive feature selection mechanism, the space-time fusion prediction subsystem improves the prediction precision based on GNN-Attention-LSTM dual-mode architecture in combination with transfer learning, and the multi-dimensional intelligent evaluation subsystem dynamically optimizes the index weight through an MLP model. And the reinforcement learning decision subsystem generates a scheduling strategy by using a PPO algorithm. According to the method, the problems of high prediction error, subjective evaluation and disjoint decision in the traditional technology are solved, the short-term prediction MAE is less than or equal to 7.8%, the renewable energy consumption rate is greatly improved, the adjustment cost is reduced, and the method is suitable for power grid dispatching optimization under high-proportion new energy access.
Owner:XINJIANG NEW ENERGY RES INST

GPU cluster resource allocation method and device, equipment and storage medium

The invention discloses a GPU cluster resource allocation method and device, equipment and a storage medium, and relates to the technical field of GPU cluster resource allocation. According to the method, a hardware performance description vector is generated by fusing static hardware parameters and dynamic micro-benchmark test data, the limitation that the GPU performance is represented only by depending on the static parameters is broken through, and the actual operation capacity of different architecture GPUs in a heterogeneous cluster is accurately matched; task feature vectors are generated by extracting task calculation features and resource demands, and quantitative description of reasoning task resource consumption features is achieved; information of hardware, tasks and load dimensions is integrated through a machine learning model, performance degradation characteristics during multi-task parallel are effectively captured, and the accuracy of execution time prediction is improved; scheduling decision operation including candidate node screening, comprehensive cost evaluation and resource reservation backfilling is executed in combination with the predicted execution time, and the task execution efficiency and cluster load balancing are both considered.
Owner:HANGZHOU DIANZI UNIV +2

GPU cluster and communication method

The invention discloses a GPU cluster and a communication method, the GPU cluster comprises at least two servers, two paths of accessories are configured in each server, each path of accessory comprises a plurality of GPUs, and in the accessories, the plurality of GPUs communicate through a first communication link; in each server, the GPUs on the two paths of accessories are in one-to-one correspondence, each group of corresponding GPUs communicate through the second communication link, GPU model selection does not need to be limited, and low-cost general GPUs can be used; two independent annular communication links are formed between the at least two servers; and in any link, data flows into the server through the high-speed network link, alternately flows through the second communication link and the first communication link so as to transmit the data to each GPU, and flows out to the next server through the high-speed network link until communication of all servers on the link is completed, so that improvement of the GPU cluster communication bandwidth is realized.
Owner:BAIYANG TIMES (BEIJING) TECH CO LTD

Heterogeneous general GPU task optimization scheduling method and system based on adaptation degree matrix

The invention discloses a heterogeneous general GPU task optimization scheduling method and system based on an adaptation degree matrix, and belongs to the technical field of computing power resource scheduling in high-performance computation.The method comprises the steps that the adaptation degree matrix between different task types and card types is constructed based on multi-type tasks and task characteristics and GPU hardware capacity in a multi-architecture GPU cluster, and the adaptation degree matrix is used for scheduling the task types and the card types; the matching degree of the task and the GPU is quantified; designing a scheduling optimization problem based on the adaptation degree matrix; and solving the scheduling optimization problem through a genetic algorithm to realize optimal resource scheduling. According to the heterogeneous general GPU task optimization scheduling method and system based on the adaptation degree matrix, optimal scheduling is achieved by matching task characteristics with GPU hardware capacity, and a new scheme is provided for efficient scheduling of heterogeneous computing tasks.
Owner:UNIV OF SCI & TECH BEIJING

GPU cluster task scheduling method and device, equipment and medium

The invention discloses a GPU cluster task scheduling method and device, equipment and a medium. The method comprises the following steps: constructing a task dependency graph in response to a to-be-scheduled task; obtaining a matched scheduling strategy according to the task characteristics of the to-be-scheduled task, marking a node type for each computing node in the task dependency graph, and determining an available GPU resource pool according to the scheduling strategy and the node type; generating all candidate GPU allocation schemes, and calculating an end-to-end execution delay of each candidate GPU allocation scheme by using a preset delay calculation model; and according to each end-to-end execution time delay and the scheduling strategy, selecting a target GPU allocation scheme from all the candidate GPU allocation schemes, so that a cluster resource manager schedules a target GPU cluster to execute the to-be-scheduled task according to the target GPU allocation scheme. According to the method, a closed loop from task analysis, resource screening to scheme optimization is formed, and the accuracy, efficiency and adaptability of GPU cluster task scheduling are improved.
Owner:HUBEI SILANG WANWEI COMPUTING EQUIPMENT MANUFACTURING CO LTD

Efficient collective communication method and device for heterogeneous GPU cluster

The invention discloses a heterogeneous GPU cluster-oriented efficient collective communication method and device, and the method comprises the steps: obtaining a target communication operator set and a physical topology corresponding to a heterogeneous GPU cluster, and determining an initial state constraint and a target state constraint; the scheduling time cost of each communication scheme is determined, and the communication scheme corresponding to the minimum scheduling time cost in the multiple scheduling time costs is selected as the efficient collective communication mode of the heterogeneous GPU cluster. According to the method, the communication scheme corresponding to the minimum scheduling time cost of the heterogeneous GPU cluster is determined to serve as the efficient collective communication mode of the heterogeneous GPU cluster based on the actual conditions of topological structure difference, link bandwidth imbalance and the like in the heterogeneous GPU cluster environment; it is ensured that the communication process is not limited by the bandwidth bottleneck, and system resource utilization efficiency and parallel training performance are prevented from being limited.
Owner:NORTHEASTERN UNIV CHINA