Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

108 results about "GPU cluster" patented technology

A GPU cluster is a computer cluster in which each node is equipped with a Graphics Processing Unit (GPU). By harnessing the computational power of modern GPUs via General-Purpose Computing on Graphics Processing Units (GPGPU), very fast calculations can be performed with a GPU cluster.

Heterogeneous GPU resource management scheduling method

The invention provides a heterogeneous GPU resource management scheduling method, and relates to the technical field of GPU resource allocation, heterogeneous equipment management and unified abstract modeling are carried out, GPU resources of different architectures are registered to a container arrangement platform, and a unified abstract layer is constructed to shield bottom layer hardware differences; gPU cluster optimization management based on a multi-dimensional real-time monitoring and intelligent scheduling strategy is carried out, GPU operation indexes are collected, priorities are dynamically calibrated for tasks, and task performance portraits are constructed; scheduling decision making is carried out through multi-strategy cooperation, and optimal GPU resources are distributed for tasks; carrying out fine-grained resource allocation, carrying out space or time segmentation on the GPU, and dynamically adjusting resource allocation according to a load state; aPI conversion of cross-architecture tasks is realized through a unified runtime library, and task execution data is collected to feed back an optimization scheduling model; automatic detection, isolation and task migration of GPU faults are carried out, and unified monitoring and alarm are provided.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Cloud platform-based computing power resource dynamic scheduling and monitoring method

The invention relates to a computing power resource dynamic scheduling and monitoring method based on a cloud platform. The method is suitable for an intelligent scheduling scene of a high-performance GPU cluster. The method comprises seven steps of task portrait modeling, GPU node state acquisition, resource trend prediction, SLA tracking, scheduling scoring and deployment, operation monitoring and task migration, and SLA feedback optimization. According to the system, task semantics are represented by constructing task vectors, node health states, topology affinity, SLA historical performance conditions and resource prediction risks are fused, a multi-factor adjustable scheduling scoring mechanism is constructed, and second-level perception and task thermal migration of high-temperature nodes are achieved. Compared with a traditional Kubernetes static scheduling scheme, the method has the advantages that the GPU utilization rate, the task SLA achievement rate and the system stability are remarkably improved, the learning ability, the self-adaptive ability and the high availability are achieved, and the method is an intelligent scheduling closed-loop system oriented to AI reasoning and training scenes.
Owner:北京娱广科技有限公司

Computing power resource fusion method based on distributed flow pipeline

The invention relates to the technical field of distributed computing and computing power scheduling, in particular to a computing power resource fusion method based on a distributed flow pipeline, which comprises the following steps of: disassembling a user computing power request into a flow pipeline unit for packaging computing logic, an input / output interface and a resource demand label; meanwhile, computing power types, real-time load rates, memory occupancy rates, network round-trip delays, geographic positioning and energy consumption data of cloud edge end nodes are collected, a hierarchical topology network framework is constructed based on the collected data, node computing power available values are calculated, an inter-node transmission cost matrix is generated, a fault probability prediction model is constructed, and a fault probability prediction model is constructed. And finally, analyzing a data dependency relationship of the pipeline unit through a four-dimensional joint decision engine, executing dynamic mapping, and preferentially mapping the high-computing-power demand unit to a GPU cluster node, so as to realize non-interruption reconstruction during pipeline topology operation. And the global computing power resource utilization rate, the task operation efficiency and the service stability are improved.
Owner:LANZHOU YUNFAN ZHILIAN TECH CO LTD +2

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

Security sensing method and system for GPU (Graphic Processing Unit) cluster

The invention relates to the technical field of network security awareness, in particular to a security awareness method and system for a GPU cluster, and the method comprises the steps: collecting multi-source security situation data, and providing basic information for subsequent security analysis; by constructing and updating a cluster security situation map in real time, information timeliness is updated and ensured in real time, and powerful support is provided for risk assessment; a dynamic risk assessment model is utilized to calculate vertex and edge scores, high-risk vertexes and key risk conduction paths are identified, and a basis is provided for resource isolation and scheduling; the targeted instruction is generated according to the security risk score and the preset security policy library, different security risks are effectively handled, and safe and stable operation of the cluster is guaranteed; by converting the instruction into a specific control command and distributing and executing the specific control command, effective implementation of security measures is ensured, and the cluster security protection capability is improved; by setting a finite-state machine monitoring instruction execution state, a recovery mechanism is automatically triggered, and continuous and safe operation of the cluster is guaranteed.
Owner:中科云达(北京)科技有限公司

Depth learning task node allocation method and system for executing time-aware computing power network heterogeneous GPU (Graphics Processing Unit) cluster

The invention discloses an execution time aware computing power network heterogeneous GPU cluster deep learning task node allocation method and system. The method comprises the following steps: firstly, based on a deep learning task, extracting and preprocessing task features and available node features; secondly, a sampler equally divides new tasks without historical data to available nodes, and each node performs mixed sampling on the tasks until all the tasks estimate execution time data; taking execution time data as a training set, taking the task features and the node features as a test set, and using a regression decision tree model to predict the execution time of the task on each node; performing task allocation on each node by using a cost search algorithm and a short job total JCT priority strategy; and finally, periodically monitoring node resources released in the cluster to obtain an optimal node allocation result. According to the method, task delay and total task JCT are remarkably reduced, cluster node resource changes are monitored in real time, and the resource utilization rate is increased.
Owner:HANGZHOU DIANZI UNIV +1

GPU cluster, redundancy optimization method of GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a GPU cluster, a redundancy optimization method of the GPU cluster, electronic equipment, a storage medium and a computer program product, the GPU cluster comprises a plurality of GPU cabinets, and each GPU cabinet comprises a plurality of GPU nodes, a GPU extension frame and an interconnection module; for any GPU cabinet, each GPU node in the GPU cabinet comprises a plurality of physical GPUs, and a GPU expansion frame in the GPU cabinet comprises a plurality of standby virtual GPUs; the interconnection module in the GPU cabinet is used for carrying out GPU interconnection on a plurality of GPU nodes and a plurality of GPU expansion frames in the GPU cabinet; and the plurality of standby virtual GPUs in the GPU extension frame in the GPU cabinet are used for providing redundant computing resources for the plurality of physical GPUs in each GPU node in the GPU cabinet. According to the embodiment of the invention, the GPU card-level redundancy capability can be effectively realized.
Owner:MOORE THREADS TECH CO LTD

Asymmetric segmentation scheduling system and method in heterogeneous GPU cluster

The invention provides an asymmetric segmentation scheduling system and method in a heterogeneous GPU cluster, and the method comprises the steps: S1, checking the features of an incoming request, and grouping the incoming request into different request buckets according to the token length; s2, in a heterogeneous GPU cluster environment, optimizing a large language model reasoning instance by adopting a double-layer strategy; and S3, calculating the matching degree of the request buckets and the big language model reasoning instances, and scheduling different request buckets to the big language model reasoning instance with the highest matching degree according to a calculation result. According to the method provided by the invention, the model layer can be asymmetrically segmented according to the computing power and the video memory capacity of each GPU on the premise of satisfying the model parallelism degree constraint, the dynamic balance of the execution duration between stages is realized, assembly line cavitation bubbles are fundamentally reduced, and the overall throughput rate is improved.
Owner:SHANGHAI JIAOTONG UNIV

GPU (Graphics Processing Unit) program optimization method for parallel environment

The invention provides a GPU program optimization method in a parallel environment, a plurality of target models are configured in a GPU cluster for parallel training, and the method comprises the following steps: performing strategy configuration and memory bottleneck pre-judgment on parallel training configuration of the target models; carrying out throughput optimization through parallel strategy combination and instruction-level performance monitoring by utilizing a pre-judgment result; kernel re-optimization is carried out on an instruction level bottleneck appearing in the optimization process, so that the model training efficiency and the GPU instruction execution efficiency are synchronously improved.
Owner:无锡九方科技有限公司

GPU resource intelligent scheduling method and system based on reinforcement learning

The invention provides an intelligent GPU resource scheduling method and system based on reinforcement learning, and relates to the field of cloud computing resource scheduling, and the method comprises the steps: collecting resource state data of a GPU cluster and attribute data of a to-be-scheduled task, carrying out the preprocessing, building a reinforcement learning model, and training the reinforcement learning model through historical scheduling records; inputting the resource state data of the current GPU cluster and the task attribute data of the to-be-scheduled task into the trained reinforcement learning model to obtain a preliminary scheduling scheme; generating a candidate scheduling scheme set based on the preliminary scheduling scheme, and selecting a final scheduling scheme from the candidate scheduling scheme set by adopting a multi-objective optimization algorithm; performing resource slice configuration on the target GPU node in the GPU cluster according to the final scheduling scheme, and starting the to-be-scheduled task in the task execution container; and updating reinforcement learning model parameters by using empirical data. According to the invention, adaptive evolution and multi-target balance optimization of the scheduling strategy are realized.
Owner:WUHAN SINGLE CLOUD NETWORK TECH CO LTD

Fault network card isolation method and device applied to GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a fault network card isolation method and device applied to a GPU cluster, electronic equipment, a storage medium and a computer program product, and the method comprises the steps: determining whether a fault network card exists on any GPU node in the GPU cluster or not according to the state information of each network card in the GPU node; when it is determined that fault network cards exist on the GPU node and the number of the fault network cards is smaller than or equal to a preset threshold value, an NCCL topological file of the GPU node is updated, and the NCCL topological file of the GPU node is used for indicating the topological relation between available GPU devices and available network cards included in the GPU node; and according to the NCCL topology file of the GPU node, isolating a fault network card in the GPU node when carrying out load scheduling on the GPU node. According to the embodiment of the invention, the fault isolation granularity can be effectively refined to a single network card in the GPU cluster.
Owner:MOORE THREADS TECH CO LTD

GPU hardware health prediction and active maintenance method and system

The invention provides a GPU (Graphics Processing Unit) hardware health prediction and active maintenance method and system, and belongs to the technical field of computing hardware maintenance and fault prediction. Based on the GPU health state data, a health degree comprehensive score of the GPU is calculated by using a preset health degree evaluation model, and a key health index in a future set time period is predicted by using a time sequence model. Calculating the fault probability of a preset type of fault risk by using a preset fault risk model based on the health degree comprehensive score of the GPU and / or the predicted key health index; and comparing the calculated fault risk probability with a preset threshold value, and executing a main preset maintenance action when the fault risk probability exceeds the preset threshold value. Through multi-source data acquisition and time sequence prediction, the potential fault risk of the GPU is found in advance, the maintenance action is actively executed, task interruption and data loss are reduced, and the reliability and availability of a GPU cluster are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Large model reasoning method for operator-level distributed scheduling

The invention discloses a large model reasoning method for operator-level distributed scheduling. The large model reasoning method comprises the steps of S1, operator decoupling and dependency modeling; s2, carrying out heterogeneous hardware capability portraying; and S3, dynamically adjusting the strategy. Operator-level fine-grained scheduling is realized in large model reasoning, and the computing power utilization rate of the heterogeneous GPU cluster is maximized.
Owner:BEIJING MOMENT UNLIMITED TECHNOLOGY CO LTD

Real-time fault tolerance and task migration system and method for process GPU (Graphics Processing Unit)

The invention provides a process GPU-oriented real-time fault tolerance and task migration system and method, and relates to the technical field of distributed computation.The method comprises the steps that a training task state in a GPU video memory is captured, difference data is generated through an incremental difference algorithm, and the difference data is stored in a multi-level storage system after being subjected to lightweight compression; meanwhile, GPU node health indexes are collected in real time, node states are analyzed through a time sequence prediction model, and a migration process is triggered when fault risks are predicted or preemption early warning is received; during migration, the task is paused, a final increment check point is generated, check point data is quickly transmitted to a target healthy GPU node through an RDMA network, a complete context state is reconstructed, task execution is recovered, and real-time fault tolerance and seamless migration of the GPU task are achieved. According to the method, the continuous operation reliability and the resource utilization efficiency of the GPU cluster in the long-time training task are improved.
Owner:HANHOU (BEIJING) TECH CO LTD

Power grid flexibility potential prediction and evaluation system based on deep learning

The invention discloses a power grid flexibility potential prediction and evaluation system based on deep learning, and belongs to the technical field of power system dispatching and artificial intelligence crossing. The system comprises a hardware layer, a software layer and an application layer, wherein the hardware layer realizes data acquisition and calculation support through a multi-source terminal, an edge node and a GPU cluster; in the software layer, the data preprocessing subsystem processes multi-source heterogeneous data by adopting an adaptive feature selection mechanism, the space-time fusion prediction subsystem improves the prediction precision based on GNN-Attention-LSTM dual-mode architecture in combination with transfer learning, and the multi-dimensional intelligent evaluation subsystem dynamically optimizes the index weight through an MLP model. And the reinforcement learning decision subsystem generates a scheduling strategy by using a PPO algorithm. According to the method, the problems of high prediction error, subjective evaluation and disjoint decision in the traditional technology are solved, the short-term prediction MAE is less than or equal to 7.8%, the renewable energy consumption rate is greatly improved, the adjustment cost is reduced, and the method is suitable for power grid dispatching optimization under high-proportion new energy access.
Owner:XINJIANG NEW ENERGY RES INST

GPU cluster resource allocation method and device, equipment and storage medium

The invention discloses a GPU cluster resource allocation method and device, equipment and a storage medium, and relates to the technical field of GPU cluster resource allocation. According to the method, a hardware performance description vector is generated by fusing static hardware parameters and dynamic micro-benchmark test data, the limitation that the GPU performance is represented only by depending on the static parameters is broken through, and the actual operation capacity of different architecture GPUs in a heterogeneous cluster is accurately matched; task feature vectors are generated by extracting task calculation features and resource demands, and quantitative description of reasoning task resource consumption features is achieved; information of hardware, tasks and load dimensions is integrated through a machine learning model, performance degradation characteristics during multi-task parallel are effectively captured, and the accuracy of execution time prediction is improved; scheduling decision operation including candidate node screening, comprehensive cost evaluation and resource reservation backfilling is executed in combination with the predicted execution time, and the task execution efficiency and cluster load balancing are both considered.
Owner:HANGZHOU DIANZI UNIV +2

GPU cluster task scheduling method and device, equipment and medium

The invention discloses a GPU cluster task scheduling method and device, equipment and a medium. The method comprises the following steps: constructing a task dependency graph in response to a to-be-scheduled task; obtaining a matched scheduling strategy according to the task characteristics of the to-be-scheduled task, marking a node type for each computing node in the task dependency graph, and determining an available GPU resource pool according to the scheduling strategy and the node type; generating all candidate GPU allocation schemes, and calculating an end-to-end execution delay of each candidate GPU allocation scheme by using a preset delay calculation model; and according to each end-to-end execution time delay and the scheduling strategy, selecting a target GPU allocation scheme from all the candidate GPU allocation schemes, so that a cluster resource manager schedules a target GPU cluster to execute the to-be-scheduled task according to the target GPU allocation scheme. According to the method, a closed loop from task analysis, resource screening to scheme optimization is formed, and the accuracy, efficiency and adaptability of GPU cluster task scheduling are improved.
Owner:HUBEI SILANG WANWEI COMPUTING EQUIPMENT MANUFACTURING CO LTD

Efficient collective communication method and device for heterogeneous GPU cluster

The invention discloses a heterogeneous GPU cluster-oriented efficient collective communication method and device, and the method comprises the steps: obtaining a target communication operator set and a physical topology corresponding to a heterogeneous GPU cluster, and determining an initial state constraint and a target state constraint; the scheduling time cost of each communication scheme is determined, and the communication scheme corresponding to the minimum scheduling time cost in the multiple scheduling time costs is selected as the efficient collective communication mode of the heterogeneous GPU cluster. According to the method, the communication scheme corresponding to the minimum scheduling time cost of the heterogeneous GPU cluster is determined to serve as the efficient collective communication mode of the heterogeneous GPU cluster based on the actual conditions of topological structure difference, link bandwidth imbalance and the like in the heterogeneous GPU cluster environment; it is ensured that the communication process is not limited by the bandwidth bottleneck, and system resource utilization efficiency and parallel training performance are prevented from being limited.
Owner:NORTHEASTERN UNIV CHINA

Tail delay optimization job scheduling method and system based on heterogeneous GPU cluster

The invention relates to the technical field of computers, and discloses a tail delay optimization job scheduling method and system based on a heterogeneous GPU cluster. The method comprises the following steps: constructing a cluster physical topological graph containing link delay, bandwidth and hop count; predicting a job communication demand graph through static code analysis and a graph neural network; executing topology matching scheduling, and preferentially deploying a high communication task pair on a high-speed interconnection link in a node or a low-delay link on the same rack; and tail delay is monitored during operation, and cost-aware local rescheduling is triggered. The system comprises a physical topology modeling module, a communication demand prediction module, an affinity scheduling module and a dynamic rescheduling module. Through active prevention and closed-loop optimization, job tail delay is reduced, and heterogeneous cluster service quality and resource utilization efficiency are improved.
Owner:SHENZHEN XINSAIKE SCI&TECH DEV CO LTD

A GPU resource intelligent scheduling method and system based on reinforcement learning

The application provides a GPU resource intelligent scheduling method and system based on reinforcement learning, relates to the field of cloud computing resource scheduling, and comprises the following steps: collecting resource state data of a GPU cluster and attribute data of a to-be-scheduled task, pre-processing, constructing a reinforcement learning model, and training the reinforcement learning model by using historical scheduling records; inputting the resource state data of the current GPU cluster and the task attribute data of the to-be-scheduled task into the trained reinforcement learning model to obtain a preliminary scheduling scheme; generating a candidate scheduling scheme set based on the preliminary scheduling scheme, selecting a final scheduling scheme from the candidate scheduling scheme set by using a multi-objective optimization algorithm; performing resource slice configuration on a target GPU node in the GPU cluster according to the final scheduling scheme, and starting the to-be-scheduled task in a task execution container; and updating the parameters of the reinforcement learning model by using experience data. The application realizes adaptive evolution and multi-objective balanced optimization of a scheduling strategy.
Owner:WUHAN SINGLE CLOUD NETWORK TECH CO LTD

GPU-accelerated gene data analysis method and system

This invention relates to the field of gene data analysis technology, and discloses a gene data analysis method and system based on GPU acceleration. The method involves: sorting the resources of each GPU card in a GPU cluster to obtain a first resource sorting queue and calculating the target GPU memory usage of sub-regions in each genome; simultaneously constructing a subtask priority queue; matching the GPU cards for each subtask in the subtask priority queue and constructing a subtask binding list for each GPU card; constructing a Pileup tensor for each bound subtask on each GPU card, performing convolutional classification on the Pileup tensor, and outputting the probabilities of different genotypes. This invention overcomes the technical problem of existing static allocation strategies being completely unaware of the real-time load status of GPU cards, ensuring that subtask allocation decisions are always based on the latest resource status of the cluster, systematically eliminating load skew problems from the scheduling layer.
Owner:SHENZHEN BLUE YIDIAN TECH CO LTD

Electromagnetic scattering calculation method based on heterogeneous GPU cluster

The embodiment of the invention provides an electromagnetic scattering calculation method based on a heterogeneous GPU cluster. The method is applied to the field of computational electromagnetism, and comprises the following steps: constructing a hierarchical bounding volume geometric acceleration structure of a target object and adding an edge index, and meanwhile, completely copying and distributing data of the acceleration structure to a video memory of each computational node in a heterogeneous GPU cluster; generating a random ray path starting from an emission source, modeling an electromagnetic scattering process as a ray path set in a probability space, and generating a ray direction through an importance sampling strategy; determining a static scheduling strategy for the ray batch based on the calculation cost of pre-sampling and geometric enhancement, and carrying out non-uniform distribution on task tiles according to the heterogeneous GPU calculation power; and monitoring the execution state of the heterogeneous GPU cluster in real time, and dynamically adjusting the distribution strategy of the residual ray tasks based on execution feedback. According to the method, the calculation overhead can be remarkably reduced while the calculation precision is ensured, and the efficiency and expandability of complex target electromagnetic scattering calculation are improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

GPU (Graphics Processing Unit) fault detection and recovery method, equipment and medium

The invention discloses a GPU (Graphics Processing Unit) fault detection and recovery method, which relates to the technical field of GPU detection, is used for solving the problem of inaccurate detection in the prior art, and comprises the following steps: collecting operation indexes of a plurality of GPU nodes in a GPU cluster; constructing graph structure data based on the operation indexes; inputting the graph structure data into a trained graph neural network model to obtain a health degree evaluation result corresponding to the GPU node; in response to the health degree evaluation result indicating abnormity, selecting and executing corresponding adaptive recovery measures from the candidate recovery actions according to the severity of the abnormity; and recording an execution effect, and dynamically optimizing a selection strategy of a subsequent recovery measure based on historical execution effect data. The invention further discloses electronic equipment and a computer storage medium. According to the method, multi-dimensional indexes are fused through graph structure data, the GPU health degree is automatically evaluated by using a graph neural network, and hierarchical recovery and strategy learning are triggered based on an evaluation result.
Owner:HANGZHOU CHENGFENGLAI DIGITAL TECH CO LTD

Intelligent calculation center cloud scheduling system

The invention provides an intelligent calculation center cloud scheduling system, and relates to the technical field of artificial intelligence and big data, and the system comprises a GPU cluster which is formed by the mixed deployment of at least a first type of GPU cards and a second type of GPU cards; and the scheduler is connected with the GPU cluster, a plurality of scheduling strategies are configured in the scheduler, and the scheduler is used for carrying out hybrid scheduling on the first type of GPU cards and the second type of GPU cards based on the scheduling strategies. The method has the beneficial effects that multi-strategy hybrid scheduling of different types of GPU cards is effectively realized, the use efficiency of computing power resources is improved, the operation cost is reduced, the operation efficiency of a system is improved, and the system risk is reduced.
Owner:SHANGHAI INTELLIGENT COMPUTING TECHNOLOGY CO LTD +1

Data processing method and device, equipment and storage medium

The invention discloses a data processing method and device, equipment and a storage medium, and belongs to the technical field of computers. The method specifically comprises the following steps: acquiring a hardware resource use condition and a draft historical acceptance rate; wherein the hardware resource use condition is determined on the basis of the operation condition of the GPU cluster system executing reasoning of the large language model, and the draft historical acceptance rate is determined on the basis of the reasoning result of the large language model; obtaining a speculation parameter adjustment strategy of a speculation decoding algorithm by utilizing a reasoning acceleration strategy prediction model based on the hardware resource use condition and the draft historical acceptance rate; and based on the speculation parameter adjustment strategy, performing optimization processing on reasoning of the large language model.
Owner:GUANGDONG UCAP INTERNET INFORMATION TECH

A GPU cluster-oriented computing power resource dynamic scheduling method

The application provides a GPU cluster-oriented computing power resource dynamic scheduling method, comprising the following steps: collecting current power quotas and memory occupation scales of each GPU computing card of a GPU cluster, and extracting running features of a non-sensitive task to obtain first resource situation data; inputting the first resource situation data into a random forest regression model previously established on a GPU cluster management node to output an estimated remaining duration of the non-sensitive task; adding the estimated remaining duration to a current system time, and combining historical running records of the non-sensitive task in the GPU cluster to determine a memory release time of the non-sensitive task; and calculating a time interval between the memory release time of the non-sensitive task and an expected arrival time of a to-be-scheduled sensitive task to obtain a first time delay result.
Owner:SHENZHEN HUMENG TECH CO LTD

Communication method and device for GPU cluster, electronic equipment and medium

The invention provides a communication method and device for a GPU cluster, electronic equipment and a medium, and relates to the technical field of artificial intelligence, in particular to the technical field of distributed computing, GPU cluster communication optimization and the like. According to the implementation scheme, the method comprises the steps of obtaining network performance data among a plurality of GPU nodes in a GPU cluster and hardware specification data of each GPU node in the plurality of GPU nodes; generating a hybrid logic communication topology for the plurality of GPU nodes based on the network performance data and the hardware specification data, the hybrid logic communication topology being a combination of at least two of a tree topology, a ring topology and a bus topology; determining a communication protocol adopted by a communication link in the hybrid logic communication topology based on the network performance data; and communicating among the plurality of GPU nodes according to the generated hybrid logic communication topology and communication protocol.
Owner:BAIDU (CHINA) CO LTD

GPU (Graphics Processing Unit) computing power resource scheduling method, system, equipment and medium

The invention provides a GPU computing power resource scheduling method, system and device and a medium. According to the method, task information of at least one to-be-scheduled task is obtained, a corresponding task portrait is constructed based on the task information, resource state information of each computing node in a GPU cluster is obtained, a corresponding node state is constructed based on the resource state information, and then the scheduling priority of the to-be-scheduled task is determined according to the task portrait and the node state. And then, according to the task portrait, the node state and the scheduling priority, determining a target GPU resource combination used for executing the to-be-scheduled task, so that the task scheduling efficiency is improved under the conditions of large-scale operation of the GPU cluster, various task types and dynamic change of resource requirements. And selecting a target GPU resource combination which can improve the overall resource utilization rate, reduce inter-task interference and shorten task completion time while ensuring that the high-priority tasks are executed in time.
Owner:杭州羿贝科技有限公司

Heterogeneous GPU cluster scheduling method based on two deadline perception

The invention relates to a scheduling method which schedules deep learning jobs through a primitive-dual framework and a dynamic programming algorithm in a heterogeneous GPU cluster and can meet different deadline requirements of the jobs. According to the method, on-line analysis and historical data are combined in a heterogeneous GPU cluster environment, so that the overhead of obtaining the job throughput is reduced; according to the method, job scheduling and job placement are decoupled, the complexity of job scheduling is reduced, meanwhile, a simple and effective job placement scheme is designed, and the degree of resource fragmentation is effectively reduced; a scheduling target is modeled in a scheduling module, reward value functions are respectively designed for strict deadline operation and soft deadline operation, an optimization problem is constructed by taking maximization of an overall reward value as the scheduling target, and strict deadline requirements and soft deadline requirements of the operation are considered at the same time; and reconstructing the target optimization function into an integer linear programming problem, and solving by using a primitive-dual framework and a dynamic programming algorithm to ensure the theoretical effectiveness of the solution.
Owner:CHENGDU UNIV OF INFORMATION TECH