System, method and host for realizing multi-core drive
By mapping each group of cores of the multi-core data processing unit to different storage areas and executing independent firmware separately, the performance degradation and failure impact of the multi-core data processing device during time-sharing multiplexing is solved, and higher performance and reliability are achieved, and isolation is performed without user sensitivity.
Patent Information
- Application Number
- CN202411998270.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When the data processing device is time-shared and multiplexed through multiple different containers, it is necessary to frequently switch the context, resulting in a degradation of the performance of the data processing device. If a certain container causes the data processing device to run a failure, it will affect other containers to use the data processing device.
Each group of cores of the data processing unit with multiple cores is mapped to different storage areas of the memory of the data processing unit, and each group of cores executes independent firmware respectively to isolate the cores of the data processing unit and the storage resources.
By isolating each group of cores and storage resources of the multi-core data processing unit, the performance and reliability of the multi-core data processing unit are improved, and the isolation of each group of cores and storage resources is achieved without the user's sense.
Smart Images

Figure CN119938181A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a system, method and host for implementing multi-core drive. Background Art
[0002] An application program (APP) running in a host often needs to start a data processing device to process data. The data processing device is, for example, a central processing unit (CPU) or a graphics processing unit (GPU).
[0003] The core of a data processing device refers to a unit in the data processing device that has data processing capabilities. The number of cores can be used to measure the data processing capabilities of the data processing device. With the advancement of technology and the continuous increase in data processing needs, data processing devices equipped with multiple cores are widely used.
[0004] A data processing device with multiple cores usually uses a firmware (FW) to schedule the multiple cores, and the device driver exposes a device to the user. When the user uses the data processing device through multiple different containers, the data processing device can be time-division multiplexed through the multiple different containers.
[0005] It should be noted that the above introduction to the technical background is only for the convenience of providing a clear and complete description of the technical solutions of the present application and for the convenience of understanding by those skilled in the art. It cannot be considered that the above technical solutions are well known to those skilled in the art simply because they are described in the background technology section of the present application. Summary of the invention
[0006] The inventors of the present application have discovered that when a data processing device is time-division multiplexed through multiple different containers, context switching needs to be performed frequently, which results in a degradation in the performance of the data processing device. Furthermore, if a container causes the data processing device to malfunction, it will affect the use of the data processing device by other containers.
[0007] In order to solve at least the above technical problems or similar technical problems, the embodiments of the present application provide a system, method, host and data processing unit for implementing multi-core drive. In the system, each group of cores of a multi-core data processing unit is mapped to different storage areas of the memory of the data processing unit, and each group of cores executes firmware independent of each other, so that each group of cores and storage resources of the data processing unit can be isolated, and the performance and reliability of the multi-core data processing unit can be improved; and a driver device unit of the host selects a data processing unit device node for each process of the host to map to a corresponding group of cores, so that the isolation of each group of cores can be achieved without the user noticing.
[0008] The embodiment of the present application provides a system for implementing multi-core driving, the system comprising a host and at least one data processing unit communicating with the host,
[0009] Each of the data processing units comprises:
[0010] A plurality of cores, wherein the plurality of cores are divided into at least two groups, and each group of cores respectively executes firmware independent of each other; and
[0011] A memory having at least two memory regions (MemRegion-N), each of which is mapped to a corresponding set of cores,
[0012] The host includes:
[0013] A driver device unit (DRMDevice), and at least two data processing unit device nodes (GPUDeviceNode),
[0014] Each of the data processing unit device nodes is mapped to one of the storage areas, and the kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for the process running on the host.
[0015] The present application also provides a method for implementing multi-core driving, the method comprising:
[0016] The multiple cores of the data processing unit are divided into at least two groups (MCn, n represents the number of cores in a group of cores), and each group of cores executes independent firmware;
[0017] Dividing the memory of the data processing unit into at least two memory regions (MemRegion-N), each of the memory regions being respectively mapped to a group of cores; and
[0018] A driver device unit (DRMDevice) and at least two data processing unit device nodes (GPUDeviceNode) are configured for a host (host) communicating with the data processing unit.
[0019] Each of the data processing unit device nodes is mapped to one of the storage areas.
[0020] The method further comprises:
[0021] A kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for a process running on the host.
[0022] The embodiment of the present application also provides a host for implementing multi-core driving, wherein the host communicates with at least one data processing unit.
[0023] Each of the data processing units comprises:
[0024] A plurality of cores, wherein the plurality of cores are divided into at least two groups (MCn, n represents the number of cores in a group of cores), and each group of cores respectively executes firmware independent of each other; and
[0025] A memory having at least two memory regions (MemRegion-N), each of which is mapped to a group of cores,
[0026] The host includes:
[0027] A driver device unit (DRMDevice), and at least two data processing unit device nodes (GPUDeviceNode),
[0028] Each of the data processing unit device nodes is mapped to a storage area of the data processing unit.
[0029] A kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for a process running on the host.
[0030] The embodiment of the present application also provides a method for implementing multi-core driving, which is applied to a host, wherein the host communicates with at least one data processing unit.
[0031] Each of the data processing units comprises:
[0032] A plurality of cores, wherein the plurality of cores are divided into at least two groups (MCn, n represents the number of cores in a group of cores), and each group of cores respectively executes firmware independent of each other; and
[0033] A memory having at least two memory regions (MemRegion-N), each of which is mapped to a group of cores,
[0034] Wherein, the method comprises:
[0035] A driver device unit (DRMDevice) and at least two data processing unit device nodes (GPUDeviceNode) are configured for the host.
[0036] Each of the data processing unit device nodes is mapped to one of the storage areas.
[0037] The method further comprises:
[0038] A kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for a process running on the host.
[0039] The beneficial effects of the embodiments of the present application are: being able to isolate each group of cores and storage resources of a multi-core data processing unit, thereby improving the performance and reliability of the multi-core data processing unit; and being able to achieve the isolation of the above-mentioned groups of cores and storage resources without the user noticing.
[0040] With reference to the following description and accompanying drawings, the specific embodiments of the present application are disclosed in detail, indicating the way in which the principles of the present application can be adopted. It should be understood that the embodiments of the present application are not limited in scope. Within the scope of the terms of the appended claims, the embodiments of the present application include many changes, modifications and equivalents.
[0041] Features described and / or illustrated with respect to one embodiment may be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or substituted for features in other embodiments.
[0042] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, integers, steps or components, but does not exclude the presence or addition of one or more other features, integers, steps or components. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0044] Figure 1 is a schematic diagram of a system for implementing multi-core drive in an embodiment of the first aspect of the present application;
[0045] Figure 2 is a schematic diagram of information stored in the data processing unit device node 125;
[0046] Figure 3 It is a schematic diagram of the composition structure of a group of nuclei;
[0047] Figure 4 is a schematic diagram of an architecture of a system for implementing multi-core drive in an embodiment of the first aspect of the present application;
[0048] Figure 5 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the second aspect;
[0049] Figure 6 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the third aspect;
[0050] Figure 7 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the fourth aspect;
[0051] Figure 8 It is a schematic diagram of an electronic device used to implement the host function. DETAILED DESCRIPTION
[0052] The foregoing and other features of the present application will become apparent through the following description with reference to the accompanying drawings. In the description and the accompanying drawings, specific embodiments of the present application are specifically disclosed, which show some embodiments in which the principles of the present application can be adopted. It should be understood that the present application is not limited to the described embodiments. On the contrary, the present application includes all modifications, variations and equivalents falling within the scope of the attached claims. Various embodiments of the present application are described below in conjunction with the accompanying drawings. These embodiments are exemplary only and are not limitations of the present application.
[0053] In the embodiments of the present application, the terms "first", "second", "upper", "lower", etc. are used to distinguish different elements in terms of title, but do not indicate the spatial arrangement or time order of these elements, etc., and these elements should not be limited by these terms. The term "and / or" includes any one and all combinations of one or more of the associated listed terms. The terms "comprising", "including", "having", etc. refer to the presence of the stated features, elements, components or components, but do not exclude the presence or addition of one or more other features, elements, components or components.
[0054] In the embodiments of the present application, the singular forms "a", "the", etc. include plural forms and should be broadly understood as "a kind" or "a type" rather than being limited to the meaning of "one"; in addition, the term "said" should be understood to include both singular and plural forms, unless the context clearly indicates otherwise. In addition, the term "according to" should be understood as "at least in part according to...", and the term "based on" should be understood as "at least in part based on...", unless the context clearly indicates otherwise.
[0055] Embodiments of the first aspect
[0056] The embodiment of the first aspect of the present application provides a system for implementing multi-core driving.
[0057] Figure 1 Schematic diagram of a system for implementing multi-core drive in an embodiment of the first aspect of the present application. Figure 1 As shown, the system 1 includes: a data processing unit 11 and a host 12.
[0058] The data processing unit 11 is, for example, a graphics processing unit (GPU), and the data processing unit 11 can be installed on a board. The data processing unit 11 can process data, for example, render image data, or perform other types of processing on data, for example, train an artificial intelligence (AI) model, etc.
[0059] The host 12 is, for example, a host of a personal computer (PC) or a host of a server. An operating system, such as a Linux operating system, runs in the host 12. The host 12 may have a user layer 121 and a kernel driver layer 122, wherein the user layer may correspond to a user mode of the operating system, and the kernel driver layer may correspond to a kernel mode of the operating system.
[0060] An application (APP) 123 may also run in the host 12. The application 123 runs in an operating system environment and may interact with a user layer 121 of the host 12. For example, the application 123 may send a task to the user layer 121. The data processing unit 11 processes the task to obtain a processing result, and the user layer 121 sends the processing result to the application 123.
[0061] exist Figure 1 In the figure, 1A indicates the host 12 side, and 1B indicates the board side on which the image processing unit 11 is installed. The board and the host 12 can exchange data via an interface 13. The interface 13 is, for example, a peripheral component interconnect express (PCIE) interface, but the present application is not limited thereto, and the interface 13 may also be other types.
[0062] exist Figure 1 In the figure, one data processing unit 11 is shown, but the present application is not limited thereto, that is, the number of data processing units 11 in the system 1 may be more than one, for example, two or three or more.
[0063] like Figure 1 As shown, each data processing unit 11 may include: multiple cores 111 and a memory 112 .
[0064] The number of the multiple cores 111 of the data processing unit 11 may be more than 2, for example, 8. The multiple cores 111 in the data processing unit 11 may be divided into at least two groups, and each group of cores may be represented as 110. There is at least 1 core in each group of cores 110, for example, there is 1 core, 2 cores, or more than 3 cores in each group of cores 110. In addition, the number of cores 111 in each group of cores 110 may be the same or different. In the present application, a group of cores may be marked as MCn-k, where n represents the number of cores in a group of cores, k represents that the group of cores is the k+1th group of cores in the data processing unit 11, n is a natural number, and k is an integer greater than or equal to 0. For example, MC1-4 (i.e., n=1, k=4) represents that there is 1 core in a group of cores 110, and the group of cores is the 5th group of cores in the data processing unit 11.
[0065] In the present application, each group of cores in a data processing unit 11 can respectively execute independent firmware, so each group of cores can be isolated from the firmware level.
[0066] In the present application, the memory 112 may be a random access memory (RAM). For example, when the data processing unit 11 is a graphics processing unit (GPU), the memory 112 may be a video RAM (VRAM). In addition, the present application is not limited thereto, and the memory 112 may also be other types.
[0067] In at least one embodiment, the memory 112 has at least two storage areas 1121, for example, the memory 112 has 8 storage areas 1121, each storage area 1121 can be represented as MemRegionN, where N can be an integer greater than or equal to 0. Each storage area 1121 can be mapped to a corresponding group of cores 110, that is, in each data processing unit 11, MemRegionN has a mapping relationship with MCn-k.
[0068] like Figure 1 As shown, in at least one embodiment, the host 12 includes at least two data processing unit device nodes 125, which can be represented as GPUDeviceNode. In at least one implementation, the data processing unit device node 125 can be configured in the kernel driver layer 122 of the host 12. Each data processing unit device node 125 can be mapped to a corresponding storage area 1121 of the data processing unit 11.
[0069] According to an embodiment of the first aspect, in system 1, each group of cores of a multi-core data processing unit 11 is mapped to a different storage area of a memory 112 of the data processing unit 11, and each group of cores executes firmware independent of each other, and each data processing unit device node 125 of the host 12 has a mapping relationship with each storage area 1121. Therefore, a mapping relationship is established between the data processing unit device node 125 of the host 12, the storage area 1121 and each group of cores, and the operation and storage resources of each group of cores are isolated from each other. In this way, when an operation failure occurs in a group of cores or several groups of cores, it will not affect the operation of other cores and the use of corresponding storage resources, thereby realizing the isolation of each group of cores and storage resources of the data processing unit, and improving the performance and reliability of the data processing unit with multiple cores.
[0070] In at least one embodiment, Figure 1 As shown, the host 12 may further include a driver device unit 126, which may be represented as DRMDevice. In at least one embodiment, the driver device unit 126 may be configured in the kernel driver layer 122 of the host 12. The driver device unit 126 may interact with the user layer 121 of the host 12 to obtain tasks, and send the tasks to the storage area 1121 corresponding to the data processing unit device node 125.
[0071] The driver unit 126 can also interact with the user layer of the host 12 to obtain information about the process. A process can run in the host 12, for example, a process can run in the user layer of the host 12. A process can correspond to a task, for example, one or more tasks occupy a process, etc. A process can also correspond to a user, for example, a user uses a process to perform a task in the host 12, etc. The driver unit 126 can have a kernel mode driver (KMD) module 1261, and the kernel mode driver (KMD) module 1261 can select a corresponding data processing unit device node (GPUDeviceNode) for each process running on the host 12.
[0072] In at least one embodiment, the method for the kernel mode driver (KMD) to select a data processing unit device node (GPUDeviceNode) for the process may include:
[0073] If the connection (Connection) corresponding to the process has been created, select the data processing unit device node (GPUDeviceNode) selected when the connection (Connection) was created last time; and / or
[0074] If the connection (Connection) corresponding to the process has not been created, the data processing unit device node (GPUDeviceNode) is selected for the process according to the following conditions:
[0075] Selecting the GPUDeviceNode with the least number of bound connections; and / or
[0076] Selecting the GPUDeviceNode with the least number of ioctl usages; and / or
[0077] Select the data processing unit device node (GPUDeviceNode) with the shortest cumulative usage time of input and output control (ioctl); and / or
[0078] According to a predetermined ordering of the data processing unit device nodes (GPUDeviceNode), the data processing unit device node (GPUDeviceNode) is selected.
[0079] In the above embodiment, connection refers to the mapping relationship between a process and a data processing unit device node (GPUDeviceNode). For example, if the connection corresponding to a process is not created, it means that no mapping relationship is established between the process and the data processing unit device node (GPUDeviceNode); if the connection corresponding to a process is created, it means that a mapping relationship has been established between the process and the data processing unit device node (GPUDeviceNode).
[0080] In at least one embodiment, Figure 1 As shown, the host 12 may further include at least one device unit 127. In at least one embodiment, the device unit 127 may be configured in a user layer 121 of the host 12.
[0081] Each device unit 127 may have a set of device files, and the set of device files in each device unit 127 may be mapped to a corresponding driver device unit 126, so that each device unit 127 has a mapping relationship with each group of cores.
[0082] Among them, the group of device files may include a first file and / or a second file, wherein the first file can be represented as / dev / dri / cardk, and the second file can be represented as dev / dri / renderK, wherein K=128+k, K is a natural number, and k is an integer greater than or equal to 0. The first file / dev / dri / cardk and the second file / dev / dri / renderK can indicate that the first file and the second file correspond to the k+1th group of cores in the data processing unit 11.
[0083] Figure 2 FIG. 1 is a schematic diagram of the information stored in the data processing unit device node 125. Figure 2 As shown, the data processing unit device node 125 may store: address information 1251 of the storage area 1121 corresponding to the data processing unit device node 125. The address information 1251 of the storage area 1121 may be represented as PhysHeap. The address information 1251 of the storage area 1121 may include: the starting address of the storage area 1121 (for example, the starting address is represented as PhysMemStartAdress) and the size of the storage area 1121 (for example, the size is represented as PhysMemSize).
[0084] like Figure 2As shown, the data processing unit device node (GPUDeviceNode) 125 also stores: information 1252 related to the storage area for firmware and tasks. The information 1252 related to the storage area for firmware and tasks can be represented as FWKernelMemCtx. Among them, the information 1252 related to the storage area for firmware and tasks includes: first information of the first storage area for storing firmware in the storage area 1121. The first information can be represented as FWMainHeap. The first information may include the starting address of the first storage area (for example, the starting address is represented as FWMainHeapStartAdress) and the size of the first storage area (for example, the size is represented as FWMainHeapSize).
[0085] The starting address of the first storage area may be a virtual address (VA), which may be read by a first processor of a group of cores corresponding to the GPUDeviceNode 125 to obtain the firmware. For the first processor, please refer to the following description.
[0086] like Figure 2 As shown, the information 1252 related to the storage area of the firmware and the task also includes: second information of a second storage area for storing tasks in the storage area 1121. The second storage area may be a shared storage space in the storage area 1121, which may be shared by different tasks. The type of the second storage area may be, for example, a kernel ring command buffer (KCCB).
[0087] Figure 3 1 is a schematic diagram of the structure of a group of cores. From a hardware perspective, a group of cores 110 may include at least one core 111; from a functional perspective, a group of cores 110 may have the following features: Figure 3 As shown in the functional module architecture, each functional module can be implemented by hardware, or by software, or by a combination of hardware and software.
[0088] like Figure 3 As shown, a group of cores 110 may include: a first processor 1101 and a first executor 1102. The first processor 1101 may execute the firmware stored in the first storage area to read the task from the second storage area and send the task to the first executor 1102; the first executor 1102 may execute the received task and perform corresponding processing, such as drawing, artificial intelligence model training, etc.
[0089] like Figure 3As shown, a group of cores 110 may further include: a memory management unit (MMU) 1103 and a reading unit 1104. The memory management unit 1103 may convert a virtual address (VA) in the first information into a physical address; the reading unit 1104 may read the firmware from the first storage area according to the physical address. The firmware read by the reading unit 1104 may be sent to the first processor 1101, and the first processor 1101 executes the firmware read from the first storage area.
[0090] In at least one embodiment, a mapping table may be stored in the memory management unit 1103, where the mapping table records the mapping relationship between the virtual address and the physical address, so that the memory management unit 1103 can convert the virtual address (VA) in the first information into a physical address.
[0091] Figure 4 It is a schematic diagram of the architecture of a system for implementing multi-core drive in an embodiment of the first aspect of the present application. Figure 4 and Figure 1 Corresponding. Figure 4 In the example shown, the data processing unit 11 is, for example, a graphics processing unit (GPU), and the data processing unit 11 has 8 cores, which are divided into 8 groups, each group having 1 core. It should be noted that the present application is not limited thereto. Figure 4 The description may also apply to cases where the number of cores in the data processing unit, and / or the number of groups, and / or the number of cores in each group are other values.
[0092] like Figure 4 As shown, User represents the user layer 121 of the host 12, Kernel represents the kernel driver layer 122 of the host 12, FW represents the firmware running in each group of cores, and the data processing unit 11 has a memory 112, and the memory 112 is, for example, the onboard physical memory of the GPU, such as VRAM.
[0093] like Figure 4 As shown, the GPU is configured in a mode of 8 MC1s (i.e., each group of cores has 1 core), and the 8 MC1s (e.g., labeled MC1-0, MC1-1, ..., MC1-7) can run independently without affecting each other, and each MC1 executes an independent FW (e.g., each FW is respectively labeled FW1, FW2, ..., FW7).
[0094] The memory 112 (eg, VRAM) is divided into 8 parts (ie, 8 storage areas, represented as MemRegion0, MemRegion1, ..., MemRegion7), and each MC1 corresponds to a storage area of the memory 112, thereby achieving physical isolation.
[0095] A driver device unit (eg, DRMDevice) is registered in the kernel driver layer 122 (eg, in the Linux Kernel state), and corresponds to a set of device files in the user layer 121 (eg, in the user state), for example, the set of device files is / dev / dri / card0 and / or / dev / dri / render128.
[0096] Eight data processing unit device nodes (eg, GPUDeviceNode0 - GPUDeviceNode7 ) are registered in the kernel driver layer 122 (eg, in the Linux Kernel state).
[0097] In addition, the kernel driver layer 122 can also register 8 display device units (eg, DisplayDevice0 to DisplayDevice7), each display device unit having a mapping relationship with each driver device unit (eg, DRMDevice0 to DRMDevice7). Each display device unit is used to drive a corresponding display device.
[0098] The following takes the first core group MC1-0 as an example to explain the process of task processing by the data processing unit 11. The description is also applicable to other core groups.
[0099] For example:
[0100] like Figure 4 As shown in ①, a set of device files / dev / dri / card0 and / or / dev / dri / render128 of the user layer 121 (for example, in user mode) interact with an application (APP) running in the host 12 to obtain a task sent by the application;
[0101] like Figure 4 As shown in ②, the group of device files sends the task to the driver device unit (e.g., DRMDevice), and the driver device unit (e.g., the kernel mode driver module of the driver device unit) selects the data processing unit device node according to the process of the task, for example, the selected data processing unit device node is GPUDeviceNode0), and sends the received task to the selected data processing unit device node (e.g., GPUDeviceNode0);
[0102] like Figure 4 As shown in ③, the data processing unit device node (for example, GPUDeviceNode0) stores the task in the second storage area (for example, KCCB) in the corresponding storage area MemRegion0 according to the stored second information;
[0103] A memory management unit (MMU) in the first group of cores MC1-0 converts a virtual address (VA) in the first information stored in the data processing unit device node (e.g., GPUDeviceNode0) into a physical address, and a reading unit in the first group of cores MC1-0 reads firmware FW0 from a first storage area in the storage area MemRegion0 according to the physical address, and sends the read FW1 to the first processor in the first group of cores MC1-0, wherein a dash-dot line with an arrow pointing from FW0 to the first storage area in MemRegion0 indicates that FW0 is stored in the first storage area in MemRegion0;
[0104] like Figure 4 As shown in ④, the first processor in the first group of cores MC1-0 executes FW0 to read the task from the second storage area (eg, KCCB) in the storage area MemRegion0;
[0105] like Figure 4 As shown in ⑤, the first processor in the first group of cores MC1-0 sends the task to the first executor in the first group of cores MC1-0, and the first executor in the first group of cores MC1-0 executes the task.
[0106] According to the above description, / dev / dri / card0 and / or / dev / dri / render128 of the user layer are mapped to DRMDevice0, DRMDevice0 is bound to GPUDeviceNode0, and GPUDeviceNode0 is mapped to MemRegion0. Other device nodes also have similar mapping relationships. As a result, the isolation of data processing units (e.g., GPUs) and internal storage resources can be achieved for users from software to hardware. In addition, for the user, the user interacts with the user layer 121, and what is presented to the user is a device (i.e., a set of device files), so the user perceives a device, and a driver device unit of the host selects a data processing unit device node for each process of the host to map to a corresponding set of cores, so that the isolation of each group of cores can be achieved without the user being aware of it.
[0107] Embodiments of the second aspect
[0108] The embodiment of the second aspect of the present application provides a method for implementing multi-core driving, which is applied to the system described in the embodiment of the first aspect.
[0109] Figure 5 FIG. 1 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the second aspect. Figure 5 As shown, the method for implementing multi-core driving includes:
[0110] Operation 501: divide multiple cores of a data processing unit into at least two groups (MCn, n represents the number of cores in a group of cores), and each group of cores respectively executes independent firmware;
[0111] Operation 502: Divide the memory of the data processing unit into at least two storage regions (MemRegion-N), and each of the storage regions is mapped to a group of cores respectively; and
[0112] Operation 503: configure a driver device unit (DRMDevice) and at least two data processing unit device nodes (GPUDeviceNode) for a host (host) communicating with the data processing unit, wherein each of the data processing unit device nodes is mapped to one of the storage areas.
[0113] like Figure 5 As shown, in at least one embodiment, the method further includes:
[0114] Operation 504: configure at least one device unit for the host, wherein the device unit has a set of device files, and the device files are mapped to the driver device unit (DRMDevice).
[0115] like Figure 5 As shown, in at least one embodiment, the method further includes:
[0116] Operation 505: The kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for the process running on the host.
[0117] In at least one embodiment, in operation 505, the kernel mode driver (KMD) selects a data processing unit device node (GPUDeviceNode) for the process, including:
[0118] If the connection (Connection) corresponding to the process has been created, select the data processing unit device node (GPUDeviceNode) selected when the connection (Connection) was last created; and / or
[0119] If the connection (Connection) corresponding to the process has not been created, the data processing unit device node (GPUDeviceNode) is selected for the process according to the following conditions:
[0120] Selecting the GPUDeviceNode with the least number of bound connections; and / or
[0121] Selecting the GPUDeviceNode with the least number of ioctl usages; and / or
[0122] Select the data processing unit device node (GPUDeviceNode) with the shortest cumulative usage time of input and output control (ioctl); and / or
[0123] According to a predetermined ordering of the data processing unit device nodes (GPUDeviceNode), the data processing unit device node (GPUDeviceNode) is selected.
[0124] In at least one embodiment, each of the data processing unit device nodes (GPUDeviceNode) stores:
[0125] Address information of the storage area corresponding to the data processing unit device node.
[0126] In at least one embodiment, each of the data processing unit device nodes (GPUDeviceNode) further stores:
[0127] The storage area includes first information of a first storage area for storing the firmware; and the storage area includes second information of a second storage area for storing tasks.
[0128] like Figure 5 As shown, in at least one embodiment, the method further includes:
[0129] Operation 506: A memory management unit (MMU) of the group of cores converts a virtual address (VA) in the first information into a physical address; and
[0130] Operation 507 : The reading unit of the group of cores reads the firmware from the first storage area according to the physical address.
[0131] like Figure 5 As shown, in at least one embodiment, the method further includes:
[0132] Operation 508: The first processor of the group of cores executes the firmware read from the first storage area to read the task from the second storage area and send the task to the first executor of the group of cores; and
[0133] Operation 509: The first executor executes the received task.
[0134] In addition, for the contents of the embodiment of the second aspect corresponding to the embodiments of the first aspect, reference can be made to the description of the embodiments of the first aspect, which will not be repeated here.
[0135] Embodiments of the third aspect
[0136] The third aspect of the present application provides a method for implementing multi-core driving. The method can be applied to Figure 1 Host 12 in.
[0137] Figure 6 FIG. 1 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the third aspect. Figure 6 As shown, the method for implementing multi-core driving includes:
[0138] Operation 601: configuring a driver device unit (DRMDevice) and at least two data processing unit device nodes (GPUDeviceNode) for the host, wherein each of the data processing unit device nodes is mapped to one of the storage areas;
[0139] Operation 602: configure at least one device unit for the host, each of the device units having a set of device files, and the device files are mapped to the driver device unit (DRMDevice); and
[0140] Operation 603: The kernel mode driver (KMD) module of the driver device unit (DRMDevice) selects one of the at least two data processing unit device nodes (GPUDeviceNode) for the process running on the host.
[0141] In at least one embodiment, the kernel mode driver (KMD) selects a data processing unit device node (GPUDeviceNode) for the process, including:
[0142] If the connection (Connection) corresponding to the process has been created, select the data processing unit device node (GPUDeviceNode) selected when the connection (Connection) was last created; and / or
[0143] If the connection (Connection) corresponding to the process has not been created, the data processing unit device node (GPUDeviceNode) is selected for the process according to the following conditions:
[0144] Selecting the GPUDeviceNode with the least number of bound connections; and / or
[0145] Selecting the GPUDeviceNode with the least number of ioctl usages; and / or
[0146] Select the data processing unit device node (GPUDeviceNode) with the shortest cumulative usage time of input and output control (ioctl); and / or
[0147] According to a predetermined ordering of the data processing unit device nodes (GPUDeviceNode), the data processing unit device node (GPUDeviceNode) is selected.
[0148] In at least one embodiment, each of the data processing unit device nodes (GPUDeviceNode) stores:
[0149] Address information of the storage area corresponding to the data processing unit device node.
[0150] In at least one embodiment, each of the data processing unit device nodes (GPUDeviceNode) further stores:
[0151] The first information of the first storage area for storing the firmware in the storage area; and
[0152] Second information of a second storage area in the storage area for storing tasks.
[0153] In addition, for the contents of the embodiment of the third aspect corresponding to those of the embodiment of the first aspect, reference can be made to the description of the embodiment of the first aspect and will not be repeated here.
[0154] Embodiments of the fourth aspect
[0155] The fourth aspect of the present application provides a method for implementing multi-core driving. The method for implementing multi-core driving can be applied to Figure 1 The data processing unit 11 in.
[0156] Figure 7 FIG. 4 is a schematic diagram of a method for implementing multi-core driving in an embodiment of the fourth aspect. Figure 7 As shown, the method for implementing multi-core driving includes:
[0157] Operation 701: divide multiple cores of the data processing unit into at least two groups (MCn, n represents the number of cores in a group of cores), and each group of cores respectively executes independent firmware; and
[0158] Operation 702: Divide the memory of the data processing unit into at least two storage regions (MemRegion-N), and each of the storage regions is mapped to a group of cores respectively.
[0159] The host is configured with a driver device unit (DRMDevice) and at least two data processing unit device nodes (GPUDeviceNode), and each of the data processing unit device nodes is mapped to one of the storage areas.
[0160] like Figure 7 As shown, the method also includes:
[0161] Operation 703: The memory management unit of the group of cores converts a virtual address (VA) in the first information indicating a storage location of the firmware into a physical address; and
[0162] Operation 704: The reading unit of the group of cores reads the firmware from a first storage area of the storage area mapped to the group of cores according to the physical address.
[0163] like Figure 7 As shown, the method also includes:
[0164] Operation 705: The first processor of the group of cores executes the firmware to read a task from a second storage area of the storage area mapped with the group of cores, and sends the task to the first executor; and
[0165] Operation 706: The first executor of the group of cores executes the received task.
[0166] In addition, for the contents of the embodiment of the fourth aspect corresponding to those of the embodiment of the first aspect, reference can be made to the description of the embodiment of the first aspect and will not be repeated here.
[0167] Embodiments of the fifth aspect
[0168] The embodiment of the fifth aspect provides an electronic device. The electronic device is used to implement Figure 1 The function of the host 12.
[0169] In the present application, the host 12 may be, for example, a computer, a server, a workstation, a laptop computer, a smart phone, etc.; however, the embodiments of the present application are not limited thereto.
[0170] Figure 8It is a schematic diagram of an electronic device used to implement the host function. Figure 8 As shown, the electronic device 800 may include: a processor (e.g., a central processing unit CPU) 810 and a memory 820; the memory 820 is coupled to the central processing unit 810. The memory 820 may store various data; in addition, it may store a program 821 for information processing, and execute the program 821 under the control of the processor 810.
[0171] In addition, if Figure 8 As shown, the host 800 may also include: an input / output (I / O) device 830 and a display 840, etc.; wherein the functions of the above components are similar to those of the prior art and are not described in detail here. It is worth noting that the electronic device 800 does not necessarily have to include Figure 8 In addition, the electronic device 800 may also include Figure 8 For components not shown in the figure, reference may be made to the related art.
[0172] An embodiment of the present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any one of the methods in the embodiments of the second to fourth aspects is implemented.
[0173] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any method in the embodiments of the second to fourth aspects is implemented.
[0174] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, any method in the embodiments of the second to fourth aspects is implemented.
[0175] The acquisition, storage, use, and processing of data in the technical solutions of each embodiment of the present application comply with the relevant provisions of national laws and regulations.
[0176] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0177] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0178] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0179] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0180] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A system for implementing multi-core driving, the system comprising a host and at least one data processing unit communicating with the host, Each of the data processing units comprises: A plurality of cores, wherein the plurality of cores are divided into at least two groups, and each group of cores executes firmware independent of each other; as well as A memory having at least two storage areas, each of which is mapped to a corresponding set of cores, The host comprises: A driver device unit, and at least two data processing unit device nodes, in, Each of the data processing unit device nodes is mapped to one of the storage areas, The kernel mode driver module of the driver device unit selects one of the at least two data processing unit device nodes for a process running on the host.
2. The system of claim 1, wherein: The host also includes: at least one device unit, wherein the device unit has a set of device files, The device file is mapped to the driver device unit.
3. The system of claim 1, wherein: The kernel mode driver selects a data processing unit device node for the process, including: If the connection corresponding to the process has been created, selecting the data processing unit device node selected when the connection was last created; and / or If the connection corresponding to the process has not been created, the data processing unit device node is selected for the process according to the following conditions: Select the data processing unit device node with the least number of bound connections; and / or Select the data processing unit device node with the least number of input and output control usages; and / or Select the data processing unit device node with the shortest cumulative usage time of input and output control; and / or The data processing unit device node is selected according to a predetermined ordering of the data processing unit device nodes.
4. The system of claim 1, wherein: Each of the data processing unit device nodes stores: Address information of the storage area corresponding to the data processing unit device node.
5. The system of claim 4, wherein: Each of the data processing unit device nodes also stores: first information of a first storage area in the storage area for storing the firmware; and Second information of a second storage area in the storage area for storing tasks.
6. The system of claim 5, wherein: The set of cores includes a first processor and a first executor, The first processor executes the firmware to read a task from the second storage area and send the task to the first executor, The first executor executes the received task.
7. The system of claim 6, wherein: The set of cores also includes a memory management unit and a read unit, The memory management unit converts the virtual address in the first information into a physical address, The reading unit reads the firmware from the first storage area according to the physical address, The first processor executes the firmware read from the first storage area.
8. A method for implementing multi-core driving, the method comprising: Dividing the plurality of cores of the data processing unit into at least two groups, wherein each group of cores executes firmware independent of each other; Dividing the memory of the data processing unit into at least two storage areas, each of the storage areas is respectively mapped to a group of cores; as well as A driver device unit and at least two data processing unit device nodes are configured for a host communicating with the data processing unit. Each of the data processing unit device nodes is mapped to one of the storage areas. The method further comprises: The kernel mode driver module of the driver device unit selects one of the at least two data processing unit device nodes for a process running on the host.
9. The method of claim 8, wherein: The method further comprises: A device unit is configured for the host, wherein the device unit has a set of device files. The device file is mapped to the driver device unit.
10. The method of claim 8, wherein: The kernel mode driver selects a data processing unit device node for the process, including: If the connection corresponding to the process has been created, selecting the data processing unit device node selected when the connection was last created; and / or If the connection corresponding to the process has not been created, the data processing unit device node is selected for the process according to the following conditions: Select the data processing unit device node with the least number of bound connections; and / or Select the data processing unit device node with the least number of input and output control usages; and / or Select the data processing unit device node with the shortest cumulative usage time of input and output control; and / or The data processing unit device node is selected according to a predetermined ordering of the data processing unit device nodes.
11. The method of claim 8, wherein: Each of the data processing unit device nodes stores: Address information of the storage area corresponding to the data processing unit device node.
12. The method of claim 11, wherein: Each of the data processing unit device nodes also stores: first information of a first storage area in the storage area for storing the firmware; and Second information of a second storage area in the storage area for storing tasks.
13. The method of claim 12, wherein: The method further comprises: The first processor of the group of cores executes the firmware to read a task from the second storage area and send the task to the first executor of the group of cores; and The first executor executes the received task.
14. The method of claim 13, wherein: The method further comprises: The memory management unit of the group of cores converts the virtual address in the first information into a physical address; and The reading unit of the group of cores reads the firmware from the first storage area according to the physical address, The first processor executes the firmware read from the first storage area.
15. A host implementing a multi-core drive, the host communicating with at least one data processing unit, Each of the data processing units comprises: A plurality of cores, wherein the plurality of cores are divided into at least two groups, and each group of cores executes firmware independent of each other; as well as A memory having at least two storage areas, each of which is mapped to a group of cores, The host comprises: A driver device unit, and at least two data processing unit device nodes, Each of the data processing unit device nodes is mapped to a storage area of the data processing unit. The kernel mode driver module of the driver device unit selects one of the at least two data processing unit device nodes for a process running on the host.
16. The host according to claim 15, wherein: The host also includes: at least one device unit, wherein the device unit has a set of device files, The device file is mapped to the driver device unit.
17. The host according to claim 15, wherein: The kernel mode driver selects a data processing unit device node for the process, including: If the connection corresponding to the process has been created, selecting the data processing unit device node selected when the connection was last created; and / or If the connection corresponding to the process has not been created, the data processing unit device node is selected for the process according to the following conditions: Select the data processing unit device node with the least number of bound connections; and / or Select the data processing unit device node with the least number of input and output control usages; and / or Select the data processing unit device node with the shortest cumulative usage time of input and output control; and / or The data processing unit device node is selected according to a predetermined ordering of the data processing unit device nodes.
18. The host according to claim 15, wherein: Each of the data processing unit device nodes stores: Address information of the storage area corresponding to the data processing unit device node.
19. The host according to claim 18, wherein: Each of the data processing unit device nodes also stores: first information of a first storage area in the storage area for storing the firmware; and Second information of a second storage area in the storage area for storing tasks.
20. A method for implementing a multi-core driver, applied to a host, the host communicating with at least one data processing unit, Each of the data processing units comprises: A plurality of cores, wherein the plurality of cores are divided into at least two groups, and each group of cores executes firmware independent of each other; as well as A memory having at least two storage areas, each of which is mapped to a group of cores, Wherein, the method comprises: A drive device unit and at least two data processing unit device nodes are configured for the host. Each of the data processing unit device nodes is mapped to one of the storage areas. The method further comprises: The kernel mode driver module of the driver device unit selects one of the at least two data processing unit device nodes for a process running on the host.
21. The method of claim 20, wherein: The method further comprises: configuring at least one device unit for the host, wherein the device unit has a set of device files, The device file is mapped to the driver device unit.
22. The method of claim 20, wherein: The kernel mode driver selects a data processing unit device node for the process, including: If the connection corresponding to the process has been created, selecting the data processing unit device node selected when the connection was last created; and / or If the connection corresponding to the process has not been created, the data processing unit device node is selected for the process according to the following conditions: Select the data processing unit device node with the least number of bound connections; and / or Select the data processing unit device node with the least number of input and output control usages; and / or Select the data processing unit device node with the shortest cumulative usage time of input and output control; and / or The data processing unit device node is selected according to a predetermined ordering of the data processing unit device nodes.
23. The method of claim 20, wherein: Each of the data processing unit device nodes stores: Address information of the storage area corresponding to the data processing unit device node.
24. The method of claim 23, wherein: Each of the data processing unit device nodes also stores: The first information in the storage area for storing the first storage area of the firmware; and Second information of a second storage area in the storage area for storing tasks.
25. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 8 to 14 and 20 to 24 is implemented.
26. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 8 to 14 and 20 to 24 is implemented.
Citation Information
Patent Citations
Apparatus and method for accelerating operations in processor which uses shared virtual memory
CN104204990A
GPU debugging device, GPU and debugging system
CN115309602A
Systems and methods for efficient multi-GPU execution of core through
CN118796493A
GPU-native packet I / O method and apparatus for GPU application on commodity ethernet
US20220272052A1