Training method and device based on GPU, electronic equipment, storage medium and computer program product

By dividing a single physical GPU into computing instances and communication instances, parallel execution of computing tasks and communication tasks is achieved, solving the problem of low physical GPU utilization in GPU cluster training and improving training efficiency.

CN120653444APending Publication Date: 2025-09-16MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510789820.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In GPU cluster training, traditional methods result in low physical GPU utilization and fail to fully realize high efficiency.

Method used

A single physical GPU is divided into GPU computing instances and GPU communication instances. Both instances share the same physical memory address space and execute computing and communication tasks in parallel.

Benefits of technology

It improves the utilization of a single physical GPU and enhances the efficiency of GPU cluster training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653444A_ABST
    Figure CN120653444A_ABST
Patent Text Reader

Abstract

The invention relates to a GPU-based training method and device, electronic equipment, a storage medium and a computer program product, and the method comprises the steps: carrying out the task segmentation of a target training task, and determining a calculation task and a communication task corresponding to the target training task; the computing task is sent to a GPU computing instance, the communication task is sent to a GPU communication instance, the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space; and calling the physical GPU to execute the calculation task by using the GPU calculation instance, and calling the physical GPU to execute the communication task by using the GPU communication instance in parallel. According to the embodiment of the invention, the utilization rate of a single physical GPU can be improved, so that the training efficiency of GPU cluster training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a GPU-based training method and device, electronic device, storage medium, and computer program product. Background Art

[0002] In GPU cluster training, traditional methods typically use a single physical GPU to serially execute the computational and communication tasks corresponding to the training task. This monopolization of the physical GPU during the communication task results in relatively low physical GPU utilization, thus failing to fully utilize the physical GPU throughout the training task. Summary of the Invention

[0003] The present disclosure provides a GPU-based training method and device, an electronic device, a storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, a GPU-based training method is provided, comprising: splitting a target training task, determining a computing task and a communication task corresponding to the target training task; sending the computing task to a GPU computing instance, and sending the communication task to a GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space; using the GPU computing instance to call the physical GPU to execute the computing task, and in parallel using the GPU communication instance to call the physical GPU to execute the communication task.

[0005] In one possible implementation, the method further includes: receiving GPU kernel driver configuration parameters and GPU resource allocation ratio parameters; based on the GPU kernel driver configuration parameters, creating a first GPU device and a second GPU device corresponding to the physical GPU on the device side in the GPU kernel driver, and creating the GPU computing instance corresponding to the first GPU device and the GPU communication instance corresponding to the second GPU device on the driver side in the GPU kernel driver; based on the GPU resource allocation ratio parameters, allocating the GPU resources of the physical GPU in proportion to the GPU computing instance and the GPU communication instance.

[0006] In one possible implementation, the method further includes: in a GPU kernel driver, initializing the GPU computing instance and the GPU communication instance, wherein the initialized GPU computing instance and the GPU communication instance share the physical memory address space corresponding to the physical GPU, and the GPU computing instance and the GPU communication instance respectively correspond to different virtual memory address spaces.

[0007] In a possible implementation, the method further includes: creating a connection context in a GPU user driver, wherein the connection context is used to establish a connection with the GPU computing instance and the GPU communication instance.

[0008] In a possible implementation, the method further includes: creating a memory space context in a GPU user driver according to the memory size of the physical GPU, wherein the memory space context is used to perform memory management on the GPU computing instance and the GPU communication instance.

[0009] In a possible implementation, the method further includes: creating, in a GPU user driver, a computing task context corresponding to the GPU computing instance and a communication task context corresponding to the GPU communication instance.

[0010] In one possible implementation, sending the computing task to the GPU computing instance and sending the communication task to the GPU communication instance include: using the computing task context to send the computing task to the GPU computing instance; using the communication task context to send the communication task to the GPU communication instance.

[0011] According to one aspect of the present disclosure, a GPU-based training device is provided, including: a task splitting module, used to split a target training task and determine a computing task and a communication task corresponding to the target training task; a sending module, used to send the computing task to a GPU computing instance, and send the communication task to a GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space; a training module, used to use the GPU computing instance to call the physical GPU to execute the computing task, and in parallel use the GPU communication instance to call the physical GPU to execute the communication task.

[0012] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0013] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.

[0014] According to one aspect of the present disclosure, a computer program product is provided, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, wherein the computer program implements the above method when executed by a processor.

[0015] In the embodiment of the present disclosure, the same physical GPU is divided into two GPU instances: a GPU computing instance and a GPU communication instance, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space. Therefore, the target training task of a single user program can be divided into computing tasks and communication tasks, and the computing tasks are sent to the GPU computing instance and the communication tasks are sent to the GPU communication instance. Then, the GPU computing instance is used to call the physical GPU to execute the computing tasks, and the GPU communication instance is used in parallel to call the physical GPU to execute the communication tasks. This effectively enables a single user program to use different GPU instances of a single physical GPU to execute computing tasks and communication tasks in parallel, thereby improving the utilization rate of a single physical GPU and thus improving the training efficiency of GPU cluster training.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0018] Figure 1 A flowchart of a GPU-based training method according to an embodiment of the present disclosure is shown.

[0019] Figure 2 A diagram showing the architecture of a training device for GPU-based training according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A block diagram of a GPU-based training device according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0023] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0024] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0025] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0026] In GPU cluster training, traditional methods typically use a single physical GPU to serially execute the computational and communication tasks corresponding to the training task. This monopolization of the physical GPU during the communication task results in relatively low physical GPU utilization, thus failing to fully utilize the physical GPU throughout the training task.

[0027] MIG (Multi-Instance) is a GPU space division and computing power reuse technology that allows a single physical GPU to be divided into multiple independent GPU instances. Each GPU instance has its own independent computing power and memory space, and each GPU instance can be assigned to different workloads. This technology enables a single physical GPU to serve multiple users or applications simultaneously.

[0028] Although MIG technology allows a single physical GPU to be divided into multiple independent GPU instances, due to the dependencies between the computational and communication tasks of a single user program and the independent memory spaces of each GPU instance, it is impossible to simultaneously dispatch the computational and communication tasks of a single user program to different GPU instances for parallel execution.

[0029] The disclosed embodiments provide a GPU-based training method that effectively enables a single user program to utilize different GPU instances of a single physical GPU to execute computing and communication tasks in parallel, thereby improving the utilization of a single physical GPU and, in turn, the efficiency of GPU cluster training. The following describes the GPU-based training method according to the disclosed embodiments in detail.

[0030] Figure 1 A flowchart of a GPU-based training method according to an embodiment of the present disclosure is shown. The method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor calling a computer-readable instruction stored in a memory. Alternatively, the method can be executed by a server. Figure 1 As shown, the method includes:

[0031] In step S11, the target training task is segmented to determine the computing task and communication task corresponding to the target training task.

[0032] In a GPU cluster training scenario, the entire training task is first broken down into multiple subtasks, each of which is independently executed on a physical GPU. Each physical GPU performs subtasks, including computing tasks and communicating with other physical GPUs in the GPU cluster. These subtasks are used to communicate and synchronize data within the GPU cluster at different stages of the training task, ensuring data consistency across all physical GPUs within the cluster.

[0033] For a user program corresponding to a single physical GPU, the subtask assigned to the user program can be considered as the target training task corresponding to the physical GPU. The user program here can be the training program corresponding to a single physical GPU in a GPU cluster scenario.

[0034] After receiving the target training task, the user program will split the target training task and determine the computing task and communication task corresponding to the target training task.

[0035] In step S12, the computing task is sent to the GPU computing instance, and the communication task is sent to the GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space.

[0036] A single physical GPU is divided into two GPU instances: a GPU compute instance and a GPU communication instance. The GPU compute instance and the GPU communication instance are set to share the same physical memory address space. Therefore, after the computation task of a single user program is sent to the GPU compute instance and the communication task is sent to the GPU communication instance, the dependency relationship between the computation task and the communication task can be ensured, effectively enabling a single user program to establish a connection with different GPU instances in a single physical GPU.

[0037] In step S13, the GPU computing instance is used to call the physical GPU to perform the computing task, and the GPU communication instance is used to call the physical GPU to perform the communication task in parallel.

[0038] The GPU computing instance calls the physical GPU to perform computing tasks, and in parallel uses the GPU communication instance to call the physical GPU to perform communication tasks, thereby effectively sending the computing tasks and communication tasks of a single user program to different GPU instances corresponding to a single physical GPU for parallel execution, thereby improving the utilization rate of a single physical GPU.

[0039] In the embodiment of the present disclosure, the same physical GPU is divided into two GPU instances: a GPU computing instance and a GPU communication instance, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space. Therefore, the target training task of a single user program can be divided into computing tasks and communication tasks, and the computing tasks are sent to the GPU computing instance and the communication tasks are sent to the GPU communication instance. Then, the GPU computing instance is used to call the physical GPU to execute the computing tasks, and the GPU communication instance is used in parallel to call the physical GPU to execute the communication tasks. This effectively enables a single user program to use different GPU instances of a single physical GPU to execute computing tasks and communication tasks in parallel, thereby improving the utilization rate of a single physical GPU and thus improving the training efficiency of GPU cluster training.

[0040] In one possible implementation, the method further includes: receiving GPU kernel driver configuration parameters and GPU resource allocation ratio parameters; creating, on the device side in the GPU kernel driver, a first GPU device and a second GPU device corresponding to the physical GPU based on the GPU kernel driver configuration parameters, and creating, on the driver side in the GPU kernel driver, a GPU computing instance corresponding to the first GPU device and a GPU communication instance corresponding to the second GPU device; and allocating, on the basis of the GPU resource allocation ratio parameters, the GPU resources of the physical GPU in proportion to the GPU computing instance and the GPU communication instance.

[0041] By calling the GPU kernel driver configuration file or using the GPU kernel driver configuration tool, enter the GPU kernel driver configuration parameters and GPU resource allocation ratio parameters.

[0042] Based on the GPU kernel driver configuration parameters, the device side in the GPU kernel driver is initialized, and two GPU devices (GPU devices) corresponding to the physical GPU are created on the device side: a first GPU device and a second GPU device; and the driver side in the GPU kernel driver is initialized, and two GPU instances corresponding to the physical GPU are created on the driver side: a GPU computing instance and a GPU communication instance.

[0043] Figure 2 FIG. 1 shows a training device architecture diagram based on a GPU for training according to an embodiment of the present disclosure. Figure 2 As shown in the figure, the training device includes a hardware layer, a kernel driver layer, a user driver layer, and a user layer. The hardware layer deploys the hardware's physical GPU, the kernel driver layer deploys the GPU kernel driver corresponding to the physical GPU, the user driver layer deploys the GPU user driver corresponding to the physical GPU, and the user layer runs the user program for GPU-based training.

[0044] like Figure 2 As shown, after the GPU kernel driver is initialized, two GPU devices are created on the device side of the GPU kernel driver: a first GPU device (first GPU Device) and a second GPU device (second GPU Device); two GPU instances are created on the driver side of the GPU kernel driver: a GPU compute instance (Compute Instance) and a GPU communication instance (Communication Instance).

[0045] The first GPU device is used to record relevant information of the GPU computing instance, including: relevant resources of the GPU computing instance (register resources, interrupt resources, etc.), and the second GPU device is used to record relevant information of the GPU communication instance, including: relevant resources of the GPU communication instance (register resources, interrupt resources, etc.).

[0046] In one example, a target GPU device may be abstracted between the first GPU device, the second GPU device, and the physical GPU to achieve better management of a single physical GPU device, which is not specifically limited in the embodiments of the present disclosure.

[0047] Based on the GPU resource allocation ratio parameter, the GPU resources of the physical GPU are proportionally allocated to the GPU computing instance and the GPU communication instance. The GPU resources can be computing resources.

[0048] In one example, based on the GPU resource allocation ratio parameter, the computing power space of the physical GPU is proportionally divided into two parts and allocated to the GPU computing instance and the GPU communication instance respectively.

[0049] The specific ratio value of the GPU resource allocation ratio parameter can be flexibly adjusted according to actual conditions, and this disclosure does not make any specific restrictions on this.

[0050] like Figure 2 As shown, the computing space of the physical GPU device in the hardware layer includes computing cores 1-8. Based on the GPU resource allocation ratio parameter, the computing space is divided into two parts: the first computing space: computing cores 1-6, and the second computing space: computing cores 7-8. Furthermore, the first computing space is allocated to the GPU compute instance, and the second computing space is allocated to the GPU communication instance.

[0051] In one possible implementation, the method further includes: in the GPU kernel driver, initializing a GPU computing instance and a GPU communication instance, wherein the initialized GPU computing instance and GPU communication instance share a physical memory address space corresponding to the physical GPU, and the GPU computing instance and GPU communication instance respectively correspond to different virtual memory address spaces.

[0052] After the above GPU kernel driver initialization is executed, the GPU computing instance and the GPU communication instance are initialized so that the GPU computing instance and the GPU communication instance are in an available state.

[0053] Since the computing tasks and communication tasks obtained by dividing the target training task are dependent on each other, the initialized GPU computing instance and GPU communication instance share the physical memory address space corresponding to the physical GPU to ensure that the GPU computing instance and GPU communication instance can access the same physical address, thereby ensuring the timing of tasks and data consistency.

[0054] like Figure 2 As shown, the GPU computing instance and the GPU communication instance share the entire physical memory space (Video Random Access Memory VRAM) corresponding to the physical GPU.

[0055] In addition, since the GPU computing instance and the GPU communication instance are two independent physical structures, the GPU computing instance and the GPU communication instance correspond to different virtual memory address spaces to ensure independent parallel processing between computing tasks and communication tasks.

[0056] In a possible implementation, the method further includes: creating a connection context in a GPU user driver, wherein the connection context is used to establish a connection with the GPU computing instance and the GPU communication instance.

[0057] Create a connection context in the GPU user driver to establish connections with the GPU compute instance and GPU communication instance in the GPU kernel driver.

[0058] In one example, the runtime in the user driver layer obtains the number of GPU instances initialized in the kernel driver layer, and then creates a connection context in the GPU user driver in the user driver layer to establish a connection with the GPU instance.

[0059] like Figure 2 As shown in Figure 1, a connection context is created in the GPU user driver.

[0060] In one possible implementation, the method further includes: creating a memory space context in the GPU user driver according to the memory size of the physical GPU, wherein the memory space context is used to perform memory management on the GPU computing instance and the GPU communication instance.

[0061] Since the GPU computing instance and GPU communication instance need to share the same physical memory address space, a memory space context is created in the GPU user driver based on the memory size of the physical GPU to enable the GPU computing instance and GPU communication instance to share the entire physical memory space corresponding to the physical GPU, thereby achieving unified physical memory allocation and management for the GPU computing instance and GPU communication instance.

[0062] In one example, the runtime in the user driver layer creates a memory space context in the GPU user driver according to the memory size of the physical GPU.

[0063] like Figure 2 As shown in Figure 1, a memory space context (Memory Context) is created in the GPU user driver.

[0064] In a possible implementation, the method further includes: creating a computing task context corresponding to the GPU computing instance and a communication task context corresponding to the GPU communication instance in a GPU user driver.

[0065] In the GPU user driver, a computing task context and a communication task context corresponding to the GPU computing instance and the GPU communication instance are created to implement the delivery of computing tasks and communication tasks.

[0066] like Figure 2 As shown in FIG, a computing task context (Compute Context) and a communication task context (Communication Context) are created in the GPU user driver.

[0067] By performing these optimizations at the GPU driver level, we effectively partition the physical GPU's computing space and manage GPU instances. By abstracting a single physical GPU into two GPU instances: a GPU compute instance and a GPU communication instance, the computing space of a single physical GPU is divided into two parts, enabling independent management of the GPU compute instance and GPU communication instance, allowing them to execute their respective tasks independently.

[0068] In one possible implementation, sending a computing task to a GPU computing instance and sending a communication task to a GPU communication instance include: using a computing task context to send the computing task to the GPU computing instance; and using a communication task context to send the communication task to the GPU communication instance.

[0069] After the above optimization of the GPU driver layer, a single user program in the user layer sends the target training task to the GPU user driver in the user driver layer. The Runtime in the user driver layer divides the target training task into computing tasks and communication tasks. Figure 2 As shown in FIG, the user layer sends the target training task to the user driver layer, and the user driver layer divides the target training task into computing tasks and communication tasks.

[0070] Furthermore, by utilizing the computing task context in the GPU user driver to send computing tasks to the GPU computing instance in the GPU kernel driver, and by utilizing the communication task context in the GPU user driver to send communication tasks to the GPU communication instance in the GPU kernel driver, it is possible to effectively realize flexible scheduling of GPU computing instances and GPU communication instances, and execute computing tasks and communication tasks in parallel.

[0071] In the embodiment of the present disclosure, the same physical GPU is divided into two GPU instances: a GPU computing instance and a GPU communication instance, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space. Therefore, the target training task of a single user program can be divided into computing tasks and communication tasks, and the computing tasks are sent to the GPU computing instance and the communication tasks are sent to the GPU communication instance. Then, the GPU computing instance is used to call the physical GPU to execute the computing tasks, and the GPU communication instance is used in parallel to call the physical GPU to execute the communication tasks. This effectively enables a single user program to use different GPU instances of a single physical GPU to execute computing tasks and communication tasks in parallel, thereby improving the utilization rate of a single physical GPU and thus improving the training efficiency of GPU cluster training.

[0072] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0073] In addition, the present disclosure also provides a GPU-based training device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any GPU-based training method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method section and will not be repeated here.

[0074] Figure 3 FIG. 1 is a block diagram of a GPU-based training device according to an embodiment of the present disclosure. Figure 3 As shown, the device 30 includes:

[0075] The task segmentation module 31 is used to segment the target training task and determine the computing task and communication task corresponding to the target training task;

[0076] a sending module 32 for sending computing tasks to a GPU computing instance and sending communication tasks to a GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space;

[0077] The training module 33 is used to use the GPU computing instance to call the physical GPU to perform computing tasks, and in parallel use the GPU communication instance to call the physical GPU to perform communication tasks.

[0078] In a possible implementation, the apparatus 30 further includes:

[0079] A receiving module is used to receive GPU kernel driver configuration parameters and GPU resource allocation ratio parameters;

[0080] A first creation module is configured to create, on the device side of the GPU kernel driver, a first GPU device and a second GPU device corresponding to the physical GPU based on GPU kernel driver configuration parameters, and to create, on the driver side of the GPU kernel driver, a GPU computing instance corresponding to the first GPU device and a GPU communication instance corresponding to the second GPU device;

[0081] The resource allocation module is used to allocate the GPU resources of the physical GPU to the GPU computing instances and GPU communication instances in proportion based on the GPU resource allocation ratio parameter.

[0082] In a possible implementation, the apparatus 30 further includes:

[0083] The initialization module is used to initialize the GPU computing instance and the GPU communication instance in the GPU kernel driver. The initialized GPU computing instance and the GPU communication instance share the physical memory address space corresponding to the physical GPU, and the GPU computing instance and the GPU communication instance correspond to different virtual memory address spaces.

[0084] In a possible implementation, the apparatus 30 further includes:

[0085] The second creation module is used to create a connection context in the GPU user driver, wherein the connection context is used to establish a connection with the GPU computing instance and the GPU communication instance.

[0086] In a possible implementation, the apparatus 30 further includes:

[0087] The third creation module is used to create a memory space context in the GPU user driver according to the memory size of the physical GPU, wherein the memory space context is used to perform memory management on the GPU computing instance and the GPU communication instance.

[0088] In a possible implementation, the apparatus 30 further includes:

[0089] The fourth creation module is used to create a computing task context corresponding to the GPU computing instance and a communication task context corresponding to the GPU communication instance in the GPU user driver.

[0090] In a possible implementation, the sending module 32 is specifically configured to:

[0091] Use the computing task context to send the computing task to the GPU computing instance;

[0092] Use the communication task context to send the communication task to the GPU communication instance.

[0093] This method has a specific technical connection with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware computing efficiency or execution effect (including reducing the amount of data storage, reducing the amount of data transmission, increasing the hardware processing speed, etc.), thereby obtaining the technical effect of improving the internal performance of the computer system in accordance with the laws of nature.

[0094] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0095] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0096] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.

[0097] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0098] The electronic device may be provided as a terminal, a server, or other forms of devices.

[0099] Figure 4 FIG. 1 is a block diagram of an electronic device according to an embodiment of the present disclosure. Figure 4 , the electronic device 1900 can be provided as a server or a terminal device. Figure 4 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0100] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (Mac OS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.

[0101] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0102] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0103] Computer-readable storage media can be a tangible device that can hold and store the instructions used by the instruction execution device. Computer-readable storage media can be, for example, (but not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. Computer-readable storage media used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0104] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0105] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0106] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0107] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0108] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0109] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0111] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0112] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0113] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0114] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A GPU-based training method, characterized in that: include: Splitting the target training task to determine the computing task and communication task corresponding to the target training task; Sending the computing task to a GPU computing instance, and sending the communication task to a GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space; The GPU computing instance is used to call the physical GPU to execute the computing task, and the GPU communication instance is used in parallel to call the physical GPU to execute the communication task.

2. The method according to claim 1, characterized in that The method further comprises: Receive GPU kernel driver configuration parameters and GPU resource allocation ratio parameters; Based on the GPU kernel driver configuration parameters, create a first GPU device and a second GPU device corresponding to the physical GPU on the device side of the GPU kernel driver, and create the GPU computing instance corresponding to the first GPU device and the GPU communication instance corresponding to the second GPU device on the driver side of the GPU kernel driver; Based on the GPU resource allocation ratio parameter, the GPU resources of the physical GPU are allocated to the GPU computing instance and the GPU communication instance in proportion.

3. The method according to claim 1, characterized in that The method further comprises: In the GPU kernel driver, the GPU computing instance and the GPU communication instance are initialized, wherein the initialized GPU computing instance and the GPU communication instance share the physical memory address space corresponding to the physical GPU, and the GPU computing instance and the GPU communication instance respectively correspond to different virtual memory address spaces.

4. The method according to claim 1, wherein The method further comprises: A connection context is created in the GPU user driver, wherein the connection context is used to establish a connection with the GPU computing instance and the GPU communication instance.

5. The method according to claim 1, wherein The method further comprises: A memory space context is created in a GPU user driver according to the memory size of the physical GPU, wherein the memory space context is used to perform memory management on the GPU computing instance and the GPU communication instance.

6. The method according to claim 1, characterized in that The method further comprises: A computing task context corresponding to the GPU computing instance and a communication task context corresponding to the GPU communication instance are created in the GPU user driver.

7. The method according to claim 6, characterized in that The sending of the computing task to the GPU computing instance and the sending of the communication task to the GPU communication instance include: Using the computing task context, sending the computing task to the GPU computing instance; The communication task is sent to the GPU communication instance using the communication task context.

8. A GPU-based training device, characterized in that: include: A task segmentation module is used to segment the target training task and determine the computing task and communication task corresponding to the target training task; a sending module, configured to send the computing task to a GPU computing instance and send the communication task to a GPU communication instance, wherein the GPU computing instance and the GPU communication instance are two GPU instances corresponding to the same physical GPU, and the GPU computing instance and the GPU communication instance correspond to the same physical memory address space; A training module is used to use the GPU computing instance to call the physical GPU to perform the computing task, and in parallel use the GPU communication instance to call the physical GPU to perform the communication task.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.