Computing system, computing management method, and electronic device
By introducing first and second switching modules into the graphics processor cluster, the interconnection between the graphics processor and the storage module and the permission configuration of the storage unit are realized, which solves the problem of low storage resource utilization and improves computing efficiency and storage space utilization flexibility.
Patent Information
- Application Number
- CN202511892950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-12-16
AI Technical Summary
Insufficient video memory resources in graphics processor clusters lead to low storage resource utilization, which affects computing efficiency.
The first switching module enables interconnection between graphics processors, and the second switching module connects the graphics processors to the storage module. The storage module serves as storage space and divides storage units, allowing storage units to be used by one graphics processor or shared by multiple processors. Usage permissions are configured to reduce idle resources.
It improved the reuse rate of storage resources, optimized hardware resource configuration, enhanced the accuracy and efficiency of storage access, and improved computing performance.
Smart Images

Figure CN121349379B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a computing system, a computing management method and an electronic device. BACKGROUND
[0002] In large-scale computing model training and actual inference process, usually need to use multiple graphics processors to cooperate to realize, graphics processor cluster exists because its memory storage resource is insufficient and affects the overall computing power utilization. For this reason, there appears an extension of other storage modules directly as graphics processor expansion storage resource mode, but due to the lack of reasonable storage management leads to unreasonable storage resource utilization, and there is a risk of further affecting the computing efficiency of graphics processor. SUMMARY
[0003] The present application provides a computing system, a computing management method and an electronic device to at least solve the problem of low storage resource utilization affecting computing efficiency in related technologies.
[0004] The present application provides a computing system, the computing system comprising: a first switching module, a plurality of graphics processors, a second switching module and a plurality of storage modules; the plurality of graphics processors are connected with the first switching module to connect with each other for realizing computing tasks; the second switching module is connected with the graphics processors respectively; the plurality of storage modules are connected with the second switching module respectively, the plurality of storage modules are used as storage space, the storage space comprises a plurality of storage units, the storage unit is associated with one or more graphics processors having its use permission; wherein, the second switching module is used for responding to the storage access request of the graphics processor, and routing the storage access request to the storage unit corresponding to the target physical address carried by the graphics processor.
[0005] The present application also provides a computing management method, the computing management method is applied to the computing system as described above, the computing management method comprising: the graphics processor generates a storage access request and transmits it to the second switching module; the second switching module acquires the storage access request, analyzes the target physical address carried by the storage access request, and queries the storage space formed by the storage module; wherein, the storage space comprises a plurality of storage units, the storage unit is associated with one or more graphics processors having its use permission; the second switching module routes the storage access request to the storage unit corresponding to the target physical address.
[0006] The present application also provides an electronic device, the electronic device comprising: a memory and a processor, the memory is used for storing computer programs; the processor is used for executing computer programs to realize the steps of any one of the above computing management methods.
[0007] Through this application, since the graphics processors (GPUs) are interconnected via a first switching module to achieve collaborative computing, and the storage modules are connected to the GPUs via a second switching module, multiple storage modules jointly serve as storage space. Storage units are then divided based on this storage space to redistribute storage scheduling units. Furthermore, storage units can be configured to be used by a single GPU or shared by multiple GPUs. This reduces the idle storage resources that exist when storage modules are bound to GPUs, thereby solving the problem of low storage resource utilization. This improves storage resource reuse, optimizes hardware resource configuration, and enhances the accuracy and efficiency of storage access. Therefore, it solves the technical problem of low storage resource utilization, achieving the technical effect of improving the flexibility of storage space usage and thus improving computing performance. Attached Figure Description
[0008] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram of the structure of an embodiment of the computing system of this application;
[0010] Figure 2 This is a schematic diagram of the structure of a graphics processor according to an embodiment of the present application;
[0011] Figure 3 This is a schematic diagram of the structure of an embodiment of the second switching module and storage module of this application;
[0012] Figure 4 A flowchart illustrating an embodiment of storage space partitioning;
[0013] Figure 5 This is a flowchart illustrating an embodiment of the computational management method of this application;
[0014] Figure 6 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0016] It should be noted that in the description of the present application, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent in such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, not to describe a specific order or sequence.
[0017] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0018] To solve the technical problems in the related art, the present application provides a computing system, a computing management method and an electronic device. The computing system comprises a first switching module, a plurality of graphic processors, a second switching module and a plurality of storage modules; the plurality of graphic processors are connected with the first switching module to be connected with each other for implementing a computing task; the second switching module is connected with the graphic processors respectively; the plurality of storage modules are connected with the second switching module respectively, the plurality of storage modules serve as storage spaces, the storage spaces comprise a plurality of storage units, and the storage units are associated with one or more graphic processors having usage permissions thereof; wherein the second switching module is configured to, in response to obtaining a storage access request of a graphic processor, route the storage access request to a storage unit corresponding to a target physical address carried by the graphic processor. The detailed working principle of the present application is described by examples.
[0019] In combination with the specific application environment architecture or the specific hardware architecture on which the execution of the computing management method depends, the specific application environment architecture or the specific hardware architecture is described herein.
[0020] That is, the embodiment of the present application provides a computing system, and the computing system is described in detail in combination with the specific architecture and working principle of the computing system.
[0021] Please refer to Figure 1 , Figure 1 The structural schematic diagram of an embodiment of the computing system of the present application is shown in the figure.
[0022] In an embodiment, the computing system can comprise a first switching module 10, a plurality of graphic processors 20, a second switching module 30 and a plurality of storage modules 40.
[0023] The plurality of graphic processors 20 are connected with the first switching module 10 to be connected with each other for implementing a computing task.
[0024] The second switching module 30 is connected with the graphic processors 20 respectively.
[0025] A plurality of storage modules 40 are respectively connected with the second switching module 30. The plurality of storage modules 40 serve as storage spaces, and the storage spaces include a plurality of storage units. The storage units are associated with one or more graphics processors 20 having usage permissions thereof.
[0026] The second switching module 30 is configured to, in response to obtaining a storage access request of the graphics processor 20, route the storage access request to a storage unit corresponding to a target physical address carried by the storage access request.
[0027] That is, in the computing system of the embodiment, the first switching module 10 is used to implement the interconnection between the graphics processors 20, and the second switching module 30 is used to implement the connection between the graphics processors 20 and the storage spaces. The second switching module 30 can be regarded as an intermediate node for the graphics processors 20 to access the storage spaces. The second switching module 30 is connected with the plurality of storage modules 40. In addition, the storage modules 40 constitute complete storage spaces. The storage spaces are re-divided to form storage units, and the storage units are controlled to be associated with one or more graphics processors 20 having usage permissions in advance. When the graphics processor 20 initiates a storage access request, the second switching module 30 can receive the storage access request and accurately route it to the corresponding storage unit according to the target physical address carried therein, thereby completing data access.
[0028] In this way, the graphics processors 20 can be interconnected by the first switching module 10 to implement collaborative computing. The storage modules 40 are connected with the graphics processors 20 by the second switching module 30. The plurality of storage modules 40 collectively serve as storage spaces to divide storage units on the basis of the storage spaces. The storage units can be configured to be used by one graphics processor 20 or shared by a plurality of graphics processors 20. This can reduce the idle storage resources when the storage modules 40 are bound with the graphics processors 20, thereby solving the problem of low storage resource utilization, improving the storage resource reuse rate, optimizing the hardware resource configuration, and improving the accuracy and efficiency of storage access. Therefore, the technical problem of low storage resource utilization can be solved, and the technical effect of improving the flexibility of storage space usage and improving the computing performance can be achieved.
[0029] Optionally, the first switching module 10 can be an NVLink Switch (high-speed interconnection switch).
[0030] The second switching module 30 can be based on the CXL (Compute Express Link) protocol. For example, the second switching module can be a CXL switch.
[0031] The storage module 40 can be a memory node of a central processing unit.
[0032] Further, the usage right includes a modification right and an access right.
[0033] When the storage unit is associated with multiple graphic processors 20 having usage rights thereof, one of the graphic processors 20 has a modification right and an access right, and the rest of the graphic processors 20 have an access right.
[0034] In the embodiment, the usage right is divided into a modification right and an access right. When the storage unit is associated with multiple graphic processors 20 having usage rights thereof, the multiple graphic processors 20 associated are assigned with differentiated rights. That is, one of the graphic processors 20 can have both a modification right and an access right, which can be considered as the master graphic processor of the storage unit. The master graphic processor writes, modifies, and the like, data in the storage unit, and shares the data in the storage unit to other graphic processors 20 associated with the storage unit, that is, the other graphic processors 20 associated with the storage unit only have an access right. The graphic processors 20 can perform operations on the storage unit based on the rights obtained by themselves. The graphic processor 20 having a modification right can read and write and modify the data in the storage unit, and the graphic processor 20 having only an access right can only read the data in the storage unit and cannot perform a modification operation, thereby being beneficial to maintaining the reliability of the data stored in the storage unit, reducing the risk of tampering with the data in the storage unit, and further being beneficial to improving the computing reliability to optimize the computing performance.
[0035] In other words, by subdividing the usage right into a modification right and an access right and specifying the right assignment rule when multiple graphic processors 20 are associated with a storage unit, the storage resource can be shared and accessed by multiple graphic processors 20, and the utilization rate of the storage resource can be improved. Moreover, by configuring a single graphic processor 20 with a modification right, the risk of conflict caused by multiple graphic processors 20 modifying data at the same time can be reduced, and the consistency and integrity of the storage data can be improved, thereby further improving the stability and reliability of the system storage access to optimize the computing performance.
[0036] For example, the computing task can include pipeline parallelism, which means that the output of a first processor is used as the input of a second processor, and the activation value information is transmitted between the two. The first processor means the graphic processor 20 on the upstream side of the pipeline parallelism, and the second processor means the graphic processor 20 on the downstream side of the pipeline parallelism.
[0037] The storage unit can include an activation value unit, and the activation value unit is associated with the first processor and the second processor. The first processor writes the activation value information into the activation value unit by using the modification right thereof, and the second processor accesses the activation value unit to obtain the activation value information, so as to realize the forward propagation of the pipeline parallelism.
[0038] In the training mode of pipeline parallelism, the graphics processor 20 on the upstream side as the first processor can output the activation value information. The graphics processor 20 on the downstream side as the second processor can take the activation value information output by the first processor as input. When the activation value unit in the storage unit is associated with the first processor and the second processor, the first processor has the modification permission and the access permission, and the second processor only has the access permission. In this way, when the forward propagation of pipeline parallelism is performed, the first processor can write the activation value information into the activation value unit by using the modification permission, and the second processor can read the activation value information in the unit by using the access permission, so as to complete the data transmission between the upstream and downstream processors and realize the pipeline parallelism.
[0039] The storage unit storing the activation value information is used as the activation value unit, the activation value unit can be associated with the upstream and downstream graphics processors 20 in pipeline parallelism, and the use permission of the first processor and the second processor can be differentiated and allocated in this way, so as to facilitate the directional writing and accurate reading of the activation value information, thereby facilitating the reliable forward propagation of pipeline parallelism. Moreover, the repeated storage of the shared activation value information can be reduced by sharing the activation value unit to transmit data, and the storage resource utilization rate can be further improved; the data writing conflict can be reduced by configuring the modification permission for the first processor, which is beneficial to guarantee the reliability of the storage of the activation value information, so as to facilitate the reliable acquisition of the activation value information by the second processor with the access permission, and to improve the calculation efficiency of the pipeline parallelism while guaranteeing the data consistency of the activation value information.
[0040] In general, in pipeline parallelism, the activation value for forward propagation and the gradient information for backward propagation need to be transmitted between the graphics processors 20 (GPU, Graphics Processing Unit) and the graphics processors 20. Generally, it can be considered that the amount of data to be transmitted is not large, and generally includes the activation value information of a single layer. Therefore, a region in a certain memory pool can be set as a shared region between two adjacent graphics processors 20, the activation value transmitted by the previous graphics processor 20 is stored in the shared region, and the next graphics processor 20 is reminded to read, and the next graphics processor 20 can read the activation value from the memory address, thereby avoiding a large amount of data transmission.
[0041] Further, the load / store (read / write) semantic access provided by CXL can also be directly read from the memory pool to the cache (cache unit) of the graphics processor 20, reducing the data access path and improving the access efficiency.
[0042] In an alternative embodiment, at least part of the graphics processors 20 associated with the storage unit can also have the modification right, which is not limited herein. For example, two, three, or four graphics processors 20 can have the modification right.
[0043] That is, please refer to Figure 1 and Figure 2 , Figure 2 is a structural schematic diagram of an embodiment of the graphics processor of the present application.
[0044] In an embodiment, the graphics processor 20 includes a cache unit. The second exchange module 30 provides access semantic information. The access semantic information includes read semantic information and / or write semantic information. The access semantic information is used to enable the graphics processor 20 to activate the value unit to read the activation value information and write the activation value information to the cache unit.
[0045] That is, the cache unit is arranged in the graphics processor 20, the second exchange module 30 can provide the graphics processor 20 with access semantic information including read semantic information and / or write semantic information, and the graphics processor 20 can directly write the activation value information to its own cache unit. The graphics processor 20 can directly perform the read and / or write operation of the activation value information based on the access semantic information, and directly write the obtained activation value information to its own cache unit during the operation. In this way, the additional data transfer link can be reduced, the data access path can be shortened, and the access and cache efficiency of the activation value information can be significantly improved, thereby realizing efficient cache storage and calling of the activation value information. At the same time, the direct reuse of the cache unit can reduce the data repeated transmission and storage overhead, thereby being beneficial to reducing the resource occupation of the computing system, further optimizing the execution process of the computing task to improve the overall computing performance of the computing system.
[0046] Optionally, the storage unit can also include a gradient unit for storing gradient information propagated backward.
[0047] The second processor has the modification right of the gradient unit, and is used to write the gradient information in the pipeline parallelism to the gradient unit, so that the gradient unit stores the gradient information of the second processor in the pipeline parallelism. The first processor accesses the gradient unit to read and obtain the gradient information, and realizes the backward propagation of the pipeline parallelism.
[0048] In the training task of pipeline parallelism, the gradient unit can also be associated with the first processor and the second processor, and the second processor on the downstream side has the modification authority of the gradient unit. After the second processor generates gradient information in the pipeline parallelism process, the second processor writes the gradient information into the gradient unit by using the modification authority, so that the gradient unit stores the gradient information generated by the second processor. The first processor has the access authority of the gradient unit, and the first processor reads and obtains the gradient information by accessing the gradient unit, and completes the backward propagation process of pipeline parallelism based on the obtained gradient information.
[0049] By the modification authority of the gradient unit of the second processor and the access authority of the first processor, reliable writing and accurate reading of the gradient information can be realized, which can adapt to the data flow demand of pipeline parallel backward propagation. At the same time, the gradient unit stores the corresponding gradient information, which can be isolated from the activation value information stored in the activation value unit, thereby reducing the risk of data confusion, thereby ensuring the accuracy of the backward propagation data. Similarly, combined with the access semantic information, the additional data transfer link in the pipeline parallel backward propagation process can be reduced, thereby shortening the gradient information transmission path and improving the backward propagation efficiency. The activation value transmission of the forward propagation is matched, so that the data interaction architecture of pipeline parallelism can be improved, and the parallel processing performance of the computing system can be further optimized.
[0050] Please continue to refer to Figures 1 to 2 In an embodiment, the graphics processor 20 includes a stream multiprocessor 21, a communication bus 22, a root unit 23, a terminal node 24, and a display memory control unit 25.
[0051] The stream multiprocessor 21 is configured to execute a computing task.
[0052] The root unit 23 is connected to the stream multiprocessor 21 through the communication bus 22, and is configured to scan and identify a storage unit, and construct a global address mapping table of the storage unit.
[0053] The terminal node 24 is connected to the stream multiprocessor 21 through the communication bus 22, and is further configured to be connected to a host computer.
[0054] The display memory control unit 25 is connected to the stream multiprocessor 21 through the communication bus 22, and is further configured to be connected to a display memory.
[0055] The graphics processor 20 includes the stream multiprocessor 21, the communication bus 22, the root unit 23, the terminal node 24, and the display memory control unit 25. That is, the stream multiprocessor 21, the communication bus 22, the root unit 23, the terminal node 24, and the display memory control unit 25 are respectively used as functional units, and the functional units can be internally connected through the communication bus 22.
[0056] The stream multiprocessor 21 can calculate components to perform a calculation task; the root unit 23 is connected with the stream multiprocessor 21 through the communication bus 22, is responsible for scanning and identifying a storage unit, and further constructs a global address mapping table of the storage unit to provide address information basis for data access. The terminal node 24 can be connected with the stream multiprocessor 21 through the communication bus 22, and simultaneously bears a connection function with a host to realize interaction between the graphics processor 20 and the host. The display memory control unit 25 is connected with the stream multiprocessor 21 through the communication bus 22, and is responsible for connection with a display memory to guarantee data transmission between the graphics processor 20 and the display memory. The functional units are connected through the communication bus 22 to form a structured interconnection, realize division of labor and cooperation and efficient data flow, and improve internal collaborative efficiency of the graphics processor 20; the global address mapping table constructed by the root unit 23 enables the stream multiprocessor 21 to quickly locate the storage unit, and shortens data access addressing time; the terminal node 24 and the display memory control unit 25 respectively realize precise docking with the host and the display memory, perfect external interaction and storage expansion capability of the graphics processor 20, and provide sufficient resource support for a calculation task; the overall architecture layout is reasonable, the functions of the components are clear, and the calculation execution efficiency and system compatibility of the graphics processor 20 can be effectively improved.
[0057] Further, the root unit 23 includes a root complex 231 and a root port 232 connected with the root complex 231.
[0058] The root complex 231 is connected with the stream multiprocessor 21 through the communication bus 22.
[0059] The root complex 231 is connected with the second switch module 30 through the root port 232, and is used for scanning a switch network topology of the second switch module 30 to identify a storage unit.
[0060] The root complex 231 is provided with a decoder 2311, the decoder 2311 is used for mapping a physical address of the storage unit to a system address space of the graphics processor 20, forming a virtual address of the storage unit, constructing a global address mapping table and recording physical address distribution of the storage unit.
[0061] That is, the detailed structure and working principle of the root unit 23 are further refined in this embodiment. In this embodiment, the root unit 23 is composed of a root complex 231 and a root port 232 connected thereto, and can perform address management of the graphics processor 20 and the external storage space. The root complex 231 can establish a fixed connection with the stream multi-processor 21 of the graphics processor 20 through the communication bus 22, and at the same time, can realize physical docking with the second switching module 30 through the root port 232, which can facilitate the smooth communication between the root complex 231 and the second switching module 30. Moreover, in the address mapping construction phase, the root complex 231 can exchange the scanning process of the network topology with the second switching module 30, by traversing the connection nodes, port links and storage module 40 access paths of the switching network, so as to facilitate accurate identification of all associated storage units and clear physical location of each storage unit in the computing system. The decoder 2311 built-in the root complex 231 can perform address conversion function, after the decoder 2311 receives the physical addresses of each storage unit scanned and acquired, the decoder 2311 can map these physical addresses one by one to the system address space of the graphics processor 20 according to the preset address mapping rule, and convert them into virtual addresses recognizable by the graphics processor 20. The root complex 231 can record the physical address distribution information of each storage unit, such as the physical address range, the corresponding switching network port and the storage module 40 number, etc., to integrate the virtual address, the physical address and the distribution information, etc. to construct the global address mapping table of the storage unit. When the stream multi-processor 21 needs to access the storage unit, it can query the global address mapping table from the root complex 231 through the communication bus 22, quickly obtain the correspondence between the virtual address and the physical address of the target storage unit, and provide address support for subsequent data access.
[0062] This can significantly improve the interaction efficiency of the graphics processor 20 and the storage system. By using the root complex 231 to realize automatic identification of the storage unit by scanning the switching network topology of the second switching module 30, the address mapping relationship can be configured without manual intervention, which reduces the complexity of system deployment and maintenance, and ensures that all connected storage units can be accurately identified, avoiding omission or misjudgment. And through the decoder 2311, the standardized mapping of the physical address to the system address space of the graphics processor 20 is realized, which adapts the dispersed physical addresses of the storage units to the virtual address management of the graphics processor 20, so that the stream multi-processor 21 can access different storage units through a unified virtual address, which can improve the flexibility and consistency of address access.
[0063] The global address mapping table records the physical address distribution information of the storage unit completely, and the stream multi-processor 21 does not need to jump and query one by one when accessing, but can directly locate the physical position of the target storage unit through the mapping table, greatly shortens the addressing time, can reduce the data access delay, and is beneficial to adapt to the application scenarios such as large model training which have certain requirements on the storage access speed. Moreover, the global address mapping table formed by the address mapping can be beneficial to enhance the compatibility and expansibility of the computing system. When the storage module 40 is added or the switching network topology is adjusted, the root complex 231 can rescan and update the global address mapping table, so that the address matching logic of the graphics processor 20 can be avoided to be modified, thereby reducing the routing cost of the computing system, and being beneficial to further guarantee the computing efficiency of the graphics processor 20, thereby guaranteeing the computing performance of the computing system.
[0064] That is, the stream multi-processor 21 generates a storage access request carrying a target virtual address.
[0065] The decoder 2311 obtains the storage access request to parse the target virtual address, queries the physical address corresponding to the target virtual address in the global address mapping table as the target physical address, converts the target virtual address in the storage access request into the target physical address, and transmits the updated storage access request to the second switching module 30 through the root port 232.
[0066] In the process of executing the computing task by the stream multi-processor 21, when the storage unit needs to be accessed, the stream multi-processor 21 can generate a storage access request carrying a target virtual address. The storage access request is obtained by the decoder 2311 to parse the target virtual address contained in the request. The decoder 2311 can query the pre-constructed global address mapping table to locate the physical address corresponding to the target virtual address in the global address mapping table, and take the physical address corresponding to the target virtual address as the target physical address. The decoder 2311 can replace the original target virtual address in the storage access request with the target physical address. After updating the request, the decoder 2311 transmits the updated storage access request to the second switching module 30 through the root port 232, thereby laying a foundation for subsequent accurate access to the storage unit.
[0067] That is, in the embodiment, the graphics processor 20 can undertake the functions of resolving the target virtual address, address translation, and request transmission guidance by the decoder 2311, and the stream multi-processor 21 does not need to participate in the address mapping related operations, so as to reduce the dispersion of computing resources, and the computing resources in the execution process of the computing task can be guaranteed, so as to improve the utilization rate and computing efficiency of the computing resources of the stream multi-processor 21. In addition, the decoder 2311 can accurately convert the target virtual address to the target physical address based on the global address mapping table, effectively reduce the risk of storage access failure caused by address mapping errors, guarantee the accuracy and reliability of storage access, and realize the automatic processing of the storage access request of the storage unit, which can shorten the request processing period, improve the response speed of storage access, and further enhance the overall stability and computing performance of the computing system.
[0068] The terminal node 24 is used for the host to perform control communication and / or event notification. The control communication includes transmission of training data of the computing model provided in the graphics processor 20, and the event notification represents asynchronous event notification of the terminal node 24 through the interrupt mechanism. And / or,
[0069] The display memory control unit 25 is used for display memory management of the display memory connected to the graphics processor 20. The display memory management includes interleaved access.
[0070] Please refer to Figures 1 to 3 , Figure 3 FIG. 2 is a structural schematic diagram of an embodiment of the second exchange module 30 and the storage module.
[0071] In an embodiment, the second exchange module 30 includes a management unit 31, a communication unit 32, and a maintenance unit 33.
[0072] The management unit 31 is used for managing whether the storage unit is associated with the graphics processor 20 based on the storage access request.
[0073] The communication unit 32 performs control plane communication with the graphics processor 20, and is used for obtaining the control instruction of the graphics processor 20, analyzing the control instruction, and feeding back to the management unit 31.
[0074] The maintenance unit 33 is connected with the communication unit 32, and is used for maintaining a usage relationship file. The usage relationship file includes the association relationship between the storage unit and the graphics processor 20 having the usage permission thereof.
[0075] The second exchange module 30 can reliably manage the use permissions and the association relationship of the storage units through the cooperative operation of the management unit 31, the communication unit 32, and the maintenance unit 33. The communication unit 32 can perform control plane communication with the graphics processor 20, obtain the control instruction sent by the graphics processor 20, and feed back the analysis result of the control instruction to the management unit 31 after the analysis of the control instruction by the communication unit 32. The management unit 31 receives the analyzed control instruction and judges and manages whether the storage unit establishes an association relationship with the graphics processor 20 in combination with the storage access request. The maintenance unit 33 is connected with the communication unit 32 and can maintain the use relationship file, that is, record the association relationship between the storage unit and the graphics processor 20 having the use permission of the storage unit in the use relationship file, to ensure the real-time and accuracy of the association information.
[0076] That is, the management unit 31, the communication unit 32, and the maintenance unit 33 can be beneficial to guarantee the smoothness of the control interaction between the graphics processor 20 and the second exchange module 30. Specifically, the management unit 31 can dynamically manage the association relationship of the storage unit to facilitate the flexible configuration and adjustment of the permissions. And maintain the corresponding use relationship file in the maintenance unit 33 to facilitate the consistency between the shared identification item and the real use scenario. And through the logical separation of the control plane and the data plane, the independence and security of the permission management can be guaranteed while the data transmission efficiency of the storage access is guaranteed. Maintaining the use relationship file can be beneficial to the visualization of the association relationship between the storage unit and the graphics processor 20, so as to facilitate the troubleshooting of the related use permissions of the computing system when the computing system is abnormal, thereby reducing the maintenance cost.
[0077] Optionally, the use relationship file includes a storage identification item, an initial association item, a current association item, and a shared marker item.
[0078] The storage identification item is used to record the unit identification of the storage unit. That is, the unit identification of each storage unit can be recorded through the storage identification item to realize the differentiation of the storage units.
[0079] The initial association item is used to record the processor identification of the initially associated virtual processor of the storage unit. The virtual processor is a default virtual graphics processor 20, which is not included in the plurality of graphics processors 20. In other words, the initial association item is used to record the processor identification of the initially associated virtual processor of the storage unit, which is a default setting and is not included in the plurality of graphics processors 20.
[0080] The current association item is used to record the processor identifier of the currently associated graphics processor 20 of the storage unit. When the storage unit is not currently associated with a graphics processor 20, the current association item records the processor identifier of the virtual processor. The current association item can record the processor identifier of the graphics processor 20 that the storage unit is currently actually associated with. If the storage unit is not currently associated with any graphics processor 20, the current association item automatically records the processor identifier of the virtual processor.
[0081] The sharing marker item is used to store whether the storage unit is shared by multiple graphics processors 20. When the storage unit is shared by multiple graphics processors 20, the sharing marker item records a first sharing marker. When the storage unit is not shared by multiple graphics processors 20, the sharing marker item records a second sharing marker. In general, the sharing marker item can explicitly store whether the storage unit is shared by multiple graphics processors 20. When the storage unit supports multiple graphics processors 20 sharing, the sharing marker item records the first sharing marker. When the storage unit does not support sharing, the sharing marker item records the second sharing marker.
[0082] The use relationship file can completely record the permission association information of the storage unit and the graphics processor 20 by the structured design of the storage identifier item, the initial association item, the current association item, and the sharing marker item. In this embodiment, the identification of the storage unit, the initial association state, the current association state, and the sharing attribute can be classified and recorded by the storage identifier item, the initial association item, the current association item, and the sharing marker item, which is beneficial to improving the clarity of the association relationship between the storage unit and the graphics processor 20 and can be beneficial to the reliability and management efficiency of the permission management. The default virtual processor can be beneficial to uniformly identifying and managing the storage unit that is not associated with the graphics processor 20, can be beneficial to reducing the situation of association state confusion or being unable to confirm, and is beneficial to confirming the storage unit with abnormal association state. Moreover, the sharing marker item can improve the efficiency of quickly distinguishing the sharing attribute of the storage unit, which is beneficial to enabling the management unit 31 to quickly judge the permission allocation of the storage unit, thereby improving the efficiency of permission configuration and adjustment and further enhancing the stability, maintainability, and resource utilization rate of the computing system.
[0083] Further, when the storage unit is shared by multiple graphics processors 20, the current association item records the processor identifier of one or more associated graphics processors 20, thereby reducing the recording complexity, improving the recording efficiency, reducing the maintenance burden of the use relationship file, and further being beneficial to guaranteeing the computing performance.
[0084] Please refer to Figures 1 to 4 , Figure 4 The flowchart of an embodiment is divided for the storage space.
[0085] In an embodiment, the management unit 31 is configured to evaluate the unit size of the storage unit by using a management parameter. The management parameter includes at least one of resource information of the second switching module 30 and storage unit management efficiency. The plurality of storage modules 40 are spatially divided according to the unit size to form the storage unit.
[0086] As Figure 4 As exemplarily shown in the figure, the storage units S1-S3 are accessible by the GPU 20GPU2, and the storage units S10-S13 are accessible by the GPU 20GPU1.
[0087] Please continue to refer to Figures 1 to 3 In an embodiment, the second switching module 30 further includes a plurality of virtual switching units 34, and each virtual switching unit 34 is provided with a plurality of virtual bridges 341.
[0088] The upstream virtual bridge 341 in the virtual switching unit 34 is connected with the storage unit, and the downstream virtual bridge 341 is connected with the GPU 20, and is configured to realize the cross-domain access of the GPU 20 to the storage unit.
[0089] In a word, in the embodiment, a multi-GPU shared memory pooling architecture based on a CXL Switch is provided. A first number of GPUs are connected to a second number of memory expansion devices, i.e., storage modules 40, through a centralized CXL Switch. The first number can be the same or different, which is not limited herein.
[0090] The memory expansion devices can be directly accessed by the GPUs in a load / store semantic manner. Each GPU is integrated with an interface conforming to the CXL standard and is configured in an RP (Root Port) mode, and is responsible for communication with the uplink port of the CXL Switch. The CXL Switch can route and forward CXL messages, and its downlink port is connected to each memory expansion device, and the memory expansion devices form a large-capacity storage space that can be uniformly addressed by the GPUs.
[0091] In addition, the GPUs are high-speed interconnected through an NVSwitch, and in a scenario requiring extremely high communication bandwidth, the computing system can use the NVLink technology to meet the demand for high-performance data transmission.
[0092] In this way, in the embodiment, the memory capacity accessible by the GPU can be increased, thereby overcoming the limitation of insufficient local GPU memory capacity, and avoiding the use of various complex distributed parallel training strategies due to the GPU memory bottleneck. Compared with the traditional single-device memory management, the memory management in the embodiment can be enriched, such as the control management of multi-GPU access, the granularity of memory page management, etc.
[0093] At the same time, the architecture of the computing system in this embodiment adopts a single CXL Switch, which supports a maximum of 4096 connected devices according to, for example, the CXL3.0 protocol specification. Therefore, the computing system in this embodiment can theoretically support the construction of a large-scale cluster required for training a large model.
[0094] The graphics processor 20 can uniformly manage the storage module 40, and therefore, the CXL root complex 231 structure is connected in the graphics processor 20 architecture through a system bus. The CXL root complex 231 can serve as the core control unit of the graphics processor 20 to manage the storage module 40, perform CXL device enumeration, scan the switching network topology, identify the connected memory devices, perform link training to negotiate the link rate, build a global address mapping table, and record the physical address distribution of the memory pool. At the same time, the HDM (Host-managed Device Memory) decoder 2311 can be integrated in the root complex 231, which is responsible for mapping the corresponding CXL memory to the system address space of the graphics processor 20. The RP port connected to the root complex 231 is responsible for processing CXL protocol data packets and can support multiple RP ports to improve data transmission capability.
[0095] Specifically, the graphics processor 20 stream multiprocessor 21 initiates a load / store request, the HDM decoder intercepts the request and parses the target virtual address, converts it to the physical address of the CXL memory device according to the address remapping table established in the initialization process, and transmits it to the uplink port of the second switching module 30 through the protocol stack encapsulation processing of the RP interface. The second switching module 30 forwards it to the corresponding memory device according to the port routing ID carried in the Flit file header to realize the access of the graphics processor 20 to the storage space. Flit is the data transmission format of the CXL protocol.
[0096] In addition, the system bus connected to the SM (Streaming Multiprocessor, stream multiprocessor 21) in the GPU also mounts a PCIe EP (PCIe Endpoint, high-speed serial computer expansion bus standard terminal device) interface, which realizes the control plane communication with the host to realize, for example, the transmission of model training data or asynchronous event notification through the interrupt mechanism. An independent video memory controller can be responsible for managing the local video memory on the GPU, such as interleaved access. The PCIe EP, CXL, and video memory controller can be controlled through the on-chip network on the communication bus 22 to isolate the transmission traffic between them.
[0097] Since multiple GPUs can access the pooled memory, the access of multiple GPUs needs to be managed and controlled. A new CXL switch architecture is designed in this paper, and a dynamic pooled memory management engine is integrated into the CXL switch. The engine receives data from the uplink port of the CXL switch, processes the data, and then forwards the data to the corresponding downlink port. The CXL pooled memory is logically divided into storage units with a predefined size M. The management unit 31 allocates or releases these storage units according to the memory access requests of the GPUs. These requests can be collected in the Mailbox (i.e., communication unit 32), which can communicate with the GPUs through the CXL.io sub-protocol. CXL.io is a sub-protocol of the CXL protocol. The GPU sends commands to the CXL switch, and the Mailbox can identify and analyze these commands for use by the memory management unit 31 (MMU). The maintenance unit 33 can record the status of the storage units, such as allocation, release, etc., and also record the associated GPUs. The storage units can be shared by multiple GPUs, so the shared record of the storage units can also be recorded. The management unit 31 can update and maintain the global address mapping table.
[0098] The data in the global address mapping table can initially set the owner of each storage unit, i.e., the initial associated item, to 0h (hexadecimal), which means that 0h indicates that it has not been accessed by any GPU. When the GPU starts accessing, the required number of storage units is calculated according to the size of the access request, and the corresponding GPU is allocated, and the GPU number in hexadecimal is recorded in the owner field. Moreover, if a storage unit is released, the current associated item in the owner field will be reset to 0h. If a storage unit is owned by multiple GPUs, the shared flag in the storage unit record table will be set to 1b (binary). Since the storage unit sharing is only for read permission, the shared GPU 20 does not have the right to modify the data in the storage unit, so the specific shared GPU 20 does not need to be recorded, and the main GPU, i.e., the GPU with the right to modify, is responsible for the storage unit, including opening and closing the shared permission, i.e., setting the shared flag to 0b or 1b. The global address mapping table in this embodiment is illustrated by the following table.
[0099] Table 1 Global address mapping table
[0100]
[0101] The specific application principle of the storage space is illustrated as follows.
[0102] In an embodiment, the GPU is used to store a copy of a computational model.
[0103] When the graphics processor is training a computational model and its video memory is occupied, it acquires training data and performs forward propagation to generate intermediate variable activation values. The graphics processor then unloads these intermediate variable activation values to the memory unit connected to it.
[0104] During training backpropagation, the graphics processor retrieves the activation values of intermediate variables carried by the memory unit, calculates the propagation gradient using the activation values of the intermediate variables, and updates the computation model using the propagation gradient.
[0105] In layman's terms, in data parallelism, each graphics processor (GPU) typically stores a complete copy of the model. The GPU memory usage of intermediate activation values accounts for a large portion of the storage space, and this usage increases with the amount of input training data. While recalculating activation values can reduce some memory usage, re-performing forward computation impacts computational efficiency. Therefore, in this embodiment, intermediate activation values can be offloaded to storage spaces such as the CXL memory pool to avoid the high latency caused by recalculation. Specifically, each GPU initiates a memory access to the storage space to obtain training data. During the forward propagation phase, each GPU temporarily stores the intermediate activation values of its corresponding training data in the storage space. During the backpropagation phase, the intermediate activation values are retrieved from the storage space to calculate gradients, and the model parameters are updated synchronously using ensemble communication.
[0106] The graphics processing unit (GPU) offloads intermediate variable activation values generated during training to its connected storage unit, effectively freeing up GPU memory and reducing training interruptions or memory overflows caused by excessive memory usage. This facilitates the GPU's support for storing and executing training tasks on larger-scale computational models. Furthermore, the GPU efficiently retrieves intermediate variable activation values from the storage unit when needed, ensuring smooth gradient calculation during backpropagation, maintaining the continuity and integrity of the training process, and improving the stability of the computational model training. This solution, through the collaborative work of the storage unit and the GPU, achieves flexible storage and reuse of intermediate variable activation values, eliminating the need to repeatedly generate intermediate data, reducing computational resource waste, and providing data support for multiple training iterations. This improves the training efficiency and convergence speed of the computational model, flexibly addressing the training needs of computational models of different scales, and enhancing the practicality and scalability of the entire training system.
[0107] In one embodiment, the computing system includes a computing model for performing computing tasks. The computing model includes multiple tensors, at least some of which are stored in storage space.
[0108] When the graphics processor performs a computational task, it reads the tensor required to perform the computational task from its own storage space as the target tensor, and then uses the target tensor to perform the computational task.
[0109] In general, in tensor parallelism, a few layers or even a single layer of the model cannot be placed in a graphics processor, so the tensor data is usually split into multiple parts, and each graphics processor is responsible for calculating a part of the tensor, but a large amount of communication between graphics processors is needed for synchronization. In the computing system of the embodiment, since the activation values are placed in the storage space, in some size of the model, tensor parallelism can be performed without tensor splitting, thereby reducing the risk of communication overhead caused by a large amount of synchronization. For a very large model, a single graphics processor can still not be able to place a single layer, and part of the tensor can be placed in the storage space, and the graphics processor reads from the memory pool when needed to perform operation, so as to reduce the memory occupation of the graphics processor.
[0110] In other words, the computing system can store at least part of the tensor of the computing model in the storage space, so that for a computing model of a certain size, the graphics processor can directly read the target tensor required for performing the computing task from the storage space, so as to be able to perform tensor parallelism without splitting the tensor of the computing model, thereby reducing the risk of communication overhead caused by a large amount of synchronization operation required by tensor parallelism. Moreover, for a computing model of a very large size, when a single graphics processor cannot accommodate a single layer of the computing model, the computing system can store part of the tensor of the computing model in the storage space, and the graphics processor can read the target tensor required for performing the computing task from the storage space for operation, thereby effectively reducing the memory occupation of the graphics processor and breaking through the memory capacity limit of the graphics processor. By storing at least part of the tensor of the computing model in the storage space and enabling the graphics processor to read the target tensor on demand to perform the computing task, the scheme can take into account the computing requirements of computing models of different sizes and improve the running efficiency of the computing system, thereby further enhancing the practicability and expansibility of the computing system and improving the computing performance of the computing system.
[0111] In an embodiment, the computing system further includes a control module and a computing model, and the computing model includes one or more domain models.
[0112] The graphics processor is configured to locally deploy at least part of the domain model.
[0113] The storage space is configured to store and maintain a model copy of all the domain models.
[0114] The control module is connected with the graphics processor and is configured to acquire a number of word pieces processed by the graphics processor using the domain model, and in response to the number of word pieces exceeding a number threshold, the domain model is taken as a to-be-scheduled model. An idle domain model in the plurality of graphics processors is queried as a to-be-unloaded model, and the to-be-unloaded model is unloaded and an auxiliary model is replaced. The auxiliary model represents a model constructed using the model copy of the to-be-scheduled model in the storage space.
[0115] Further, the control module is further configured to monitor whether the token is successfully routed to the local domain model of the graphics processor. In response to the token failing to be routed to the local domain model, the control module controls the token to be exchanged between the graphics processors through the first exchange module.
[0116] That is, the control module acquires the number of tokens processed by the graphics processor using the domain model, and when the number of tokens exceeds the number threshold, the corresponding domain model is taken as a model to be scheduled. The control module queries and unloads the idle domain model in the plurality of graphics processors, and replaces the idle domain model with an auxiliary model constructed using the model copy of the model to be scheduled in the storage space. This can dynamically balance the computing load of the domain model between the plurality of graphics processors, and effectively reduce the risk of load imbalance caused by uneven distribution of tokens in the expert parallel strategy. The control module monitors the routing of the token to the local domain model of the graphics processor, and controls the token to be exchanged between the graphics processors through the first exchange module when the routing fails, so as to ensure that the token can be normally matched with the required domain model and the computing task can be continuously executed. The local deployment of the domain model in the graphics processor can improve the token processing efficiency, and the storage space stores and maintains the model copy of all domain models, which provides data support for dynamic replacement and routing compensation. The two are cooperatively adapted to the sparse utilization characteristics of the expert parallel strategy, ensuring the efficiency of local computing while taking into account the flexible scheduling of global models, which can significantly improve the overall utilization rate of the domain model and enhance the adaptation ability and running stability of the computing system to different scale computing tasks.
[0117] In general, the domain model can be considered as an expert model. In model training, the expert parallel strategy is used. Since the utilization rate of the expert model is relatively sparse, the average activation of the expert model per Token (token) is about 1%, therefore, only part of the expert model is stored in each graphics processor in the expert parallel strategy. However, the problem of how many expert models are placed on each graphics processor still exists, and there is a risk of load imbalance. In the embodiment, a certain number of expert models are maintained in the local memory of each graphics processor, and a complete expert model is maintained in the memory pool. When the Token is routed to the expert model that is not in the local graphics processor, All2All communication is performed through the first exchange module. The All2All communication indicates full interconnection data exchange between all graphics processor nodes. When the number of tokens calculated on a certain graphics processor is too large, the complete expert model copy in the storage space can be used to replace some idle expert models on the graphics processor, so as to maintain the dynamic balance of the computing load of the expert model between the graphics processors.
[0118] The application principle of the storage space is clearly shown in the following table:
[0119] Table 2: Mapping relationship table of training mode and memory, storage space storage data and beneficial effects
[0120]
[0121] In the large model inference scenario, the memory occupation of the graphics processor includes model parameters and KV cache (key-value cache), wherein the storage space of the KV cache occupies a large part and can reach several TB (terabytes). The CXL memory pool can store the KV cache using fewer tokens and load the same into the corresponding graphics processor memory when needed, thereby relieving the bottleneck of the graphics processor memory capacity. In the large model inference scenario, each graphics processor stores a part of the model, and the KV cache corresponding to the part of the model is unloaded to the CXL memory pool. The unloading is based on the importance of the attention score of the token. The KV cache of the most important historical token is reserved in the local memory of each graphics processor, and is updated in time according to the importance, that is, the KV cache is replaced with the storage space.
[0122] In the embodiment, the computing system architecture can use the CXL bus level high-performance link to replace part of the network interconnection link. Meanwhile, compared with using the central processor memory, the graphics processor directly connects the CXL storage space, which can reduce the transit overhead on the central processor side. Therefore, the performance in the large model training or inference scenario can be significantly improved.
[0123] Embodiments of the present application provide a computing management method. The computing management method is described in detail in combination with the execution process of the computing management method.
[0124] Please refer to Figure 5 , Figure 5 The flowchart of an embodiment of the computing management method of the present application is shown.
[0125] S101: The graphics processor generates a storage access request and transmits the same to the second switching module.
[0126] In the embodiment, when the graphics processor needs to read or write a storage resource during the execution of a computing task, the graphics processor generates a storage access request carrying a target physical address. The graphics processor transmits the storage access request to the second switching module. The second switching module serves as a routing transit of storage access and can analyze and direct the forwarding of the storage access request.
[0127] S102: The second switching module acquires the storage access request, analyzes the target physical address carried by the storage access request, and queries the storage space formed by the storage module; wherein the storage space includes a plurality of storage units, and the storage unit is associated with one or more graphics processors having the use authority thereof.
[0128] In this embodiment, when the second switching module receives the storage access request transmitted by the graphics processor, the target physical address carried by the storage access request can be parsed to determine the specific physical location of the storage resource accessed by the graphics processor. In this embodiment, the computing system includes a storage space composed of multiple storage modules. The storage space is divided into multiple independent storage units, and each storage unit is associated with one or more graphics processors having usage permission. In this way, the second switching module queries the storage space to determine the storage unit corresponding to the target physical address. Alternatively, the access permission of the graphics processor to the target storage unit can also be determined based on the usage permission.
[0129] S103: The second switching module routes the storage access request to the storage unit corresponding to the target physical address.
[0130] In this embodiment, after the second switching module completes the target physical address parsing and storage space querying, the storage access request can be routed to the storage unit corresponding to the target physical address based on the parsing result and the authorized association information of the storage unit. The data link between the graphics processor and the storage unit is realized based on the switching network topology of the second switching module. Precise routing can reduce the path redundancy of storage access, and can also be combined with the authorized association mechanism of the storage unit to facilitate the access of graphics processors with usage permission to the corresponding storage unit, thereby ensuring the security of storage access and improving the data read-write efficiency. Moreover, the storage unit can be associated with multiple graphics processors with usage permission, which facilitates the sharing of storage resources by multiple graphics processors, reduces the waste of storage resources caused by allocation, and improves the utilization rate of storage resources.
[0131] As can be seen, in this embodiment, the graphics processor is interconnected to realize collaborative computing through the first switching module, the storage module is connected with the graphics processor through the second switching module, multiple storage modules are collectively used as a storage space to divide storage units in the storage space, to redivide the storage scheduling unit, and the storage unit can be used by one graphics processor or shared by multiple graphics processors, which can reduce the idle storage resources when the storage module is bound with the graphics processor, thereby solving the problem of low storage resource utilization, improving the storage resource reuse rate, optimizing the hardware resource configuration, and improving the accuracy and efficiency of storage access, and thus improving the flexibility of storage space usage and the computing performance.
[0132] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment.
[0133] The features of the embodiments of the computing management method can refer to the related descriptions of the embodiments of the computer system, which will not be repeated here.
[0134] That is, the current efficient coherent memory semantics function provided by using the CXL.mem sub-protocol can be used to cooperate with the efficient management of the software side of the memory pooling resource, which is beneficial to realize the dynamic abstraction and fine scheduling of the memory pooling, so as to allow the heterogeneous host to directly access the shared memory block with ultra-low latency on demand. Among them, CXL.mem is a sub-protocol in the CXL protocol. That is, the CXL.mem and the cache coherence controller such as the host interface, the memory pool side device controller need to be integrated on the computing system hardware; the address translation is realized on the software side of the computing system through the operating system or the virtualization layer, for example, the OS (Operating System) function and the resource arbitration strategy can be extended based on the CXL device memory semantics.
[0135] However, the current CXL-based memory pooling is mainly around the host CPU side, and there is no memory pooling around the GPU device side. In the application scenario of AI (Artificial Intelligence) large models, due to the fact that a single GPU memory cannot completely store the intermediate variables generated in the large model training or inference process, such as activation values and KV cache, the existing multi-GPU parallel training or inference scheme is relatively low in efficiency. In the multi-GPU architecture, the number of connected GPUs in the node is limited, usually 8 or 16 cards per machine, and more GPUs in multiple machines are connected through a network, and the efficiency of network connection is relatively low, such as the time delay is usually in the order of microseconds. The inefficiency of such interconnection leads to insufficient utilization of computing power of the existing multi-GPU architecture. The super-node architecture can support the connection of 256 GPU cards and expand the GPU memory by introducing the CPU (Central Processing Unit) side memory through consistent interconnection, but this architecture is currently not widely popularized due to high cost, and there is no shared memory pool between multiple GPUs, and there is no optimization for the management of large memory, which further aggravates the problem of insufficient computing power utilization. The computing system and the computing management method of the present application can overcome the above problems to some extent.
[0136] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of an embodiment of an electronic device of the present application.
[0137] In an embodiment, the electronic device can include a device body 51 and a computing system 52, wherein the computing system 52 is arranged in the device body 51.
[0138] The computing system 52 can be as described in any of the preceding embodiments. That is, the computing system 52 can include at least a first switch module, a plurality of graphic processors, a second switch module, and a plurality of storage modules. The plurality of graphic processors are respectively connected to the first switch module for performing computing tasks. The second switch module is respectively connected to the graphic processors. The plurality of storage modules are respectively connected to the second switch module. The plurality of storage modules serve as storage spaces, and the storage spaces include a plurality of storage units. The storage units are associated with one or more graphic processors having usage permissions thereof. The second switch module is configured to route a storage access request to a storage unit corresponding to a target physical address carried by the storage access request in response to obtaining the storage access request.
[0139] Embodiments of the present application also provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above computing management method embodiments. The computing management method includes at least that a graphic processor generates a storage access request and transmits the storage access request to a second switch module; the second switch module obtains the storage access request, parses a target physical address carried by the storage access request, and queries a storage space formed by a plurality of storage modules; the storage space includes a plurality of storage units, and the storage units are associated with one or more graphic processors having usage permissions thereof; and the second switch module routes the storage access request to a storage unit corresponding to the target physical address.
[0140] Embodiments of the present application also provide a computer readable storage medium storing a computer program. When the computer program is executed, the steps in any of the above computing management method embodiments are performed.
[0141] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0142] Embodiments of the present application also provide a computer program product including a computer program. When the computer program is executed by a processor, the steps in any of the above computing management method embodiments are performed.
[0143] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium stores a computer program, the computer program is executed by a processor to implement the steps in any of the above computing management method embodiments.
[0144] Those skilled in the art will further appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the above description has generally been stated in terms of the functional components and steps of the examples. Whether such functionality is implemented in hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in varying ways for each particular application, but such implementation should not be interpreted as a departure from the scope of the present application.
[0145] The above provides a computing system, a computing management method and an electronic device. The principles and implementation manners of the present application are described by applying specific examples. The above description of the examples is only used to help understand the method and core idea of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A computing system, characterized in that, The computing system includes: First switching module; Multiple graphics processors are connected to the first switching module to interconnect and are used to perform computing tasks. The second switching module is connected to the graphics processor respectively; Multiple storage modules are connected to the second switching module, and the multiple storage modules serve as storage space. The storage space includes multiple storage units, and the storage units are associated with one or more graphics processors that have access to them. The second switching module is used to route the storage access request to the storage unit corresponding to the target physical address carried by the graphics processor in response to the acquisition of the storage access request of the graphics processor. The usage permissions include modification permissions and access permissions; When the storage unit is associated with multiple graphics processors that have usage rights, one of the graphics processors has the modification right and the access right, and the other graphics processors have the access right. The computing task includes pipelined parallelism, where the output of the first processor is used as the input of the second processor, and activation value information is passed between them; the first processor represents the pipelined parallelism as the upstream graphics processor, and the second processor represents the pipelined parallelism as the downstream graphics processor. The storage unit includes an activation value unit, which is associated with the first processor and the second processor. The first processor uses its modification permission to write the activation value information into the activation value unit, and the second processor accesses the activation value unit to obtain the activation value information, so as to realize the pipelined forward propagation.
2. The computing system according to claim 1, characterized in that, The storage unit includes a gradient unit; The second processor has the modification permission of the gradient unit, and is used to write the gradient information in the pipeline parallelism into the gradient unit, so that the gradient unit stores the gradient information of the second processor in the pipeline parallelism; The first processor accesses the gradient unit to read and acquire the gradient information, thereby realizing the pipelined parallel backpropagation.
3. The computing system according to claim 1, characterized in that, The second switching module includes a management unit, a communication unit, and a maintenance unit; The management unit is used to manage whether the storage unit is associated with the graphics processor based on the storage access request; The communication unit communicates with the graphics processor via the control plane to obtain control commands from the graphics processor, parse the control commands, and feed them back to the management unit. The maintenance unit is connected to the communication unit and is used to maintain the usage relationship file; wherein the usage relationship file includes the association between the storage unit and the graphics processors that have usage rights to it.
4. The computing system according to claim 3, characterized in that, The usage relationship file includes storage identifiers, initial associations, current associations, and shared tag entries; The storage identifier is used to record the unit identifier of the storage unit; The initial association entry is used to record the processor identifier of the virtual processor initially associated with the storage unit; wherein, the virtual processor is a default virtual graphics processor, and the virtual graphics processor is not included in the plurality of graphics processors; The current association entry is used to record the processor identifier of the graphics processor currently associated with the storage unit; wherein, when the storage unit is not currently associated with the graphics processor, its current association entry records the processor identifier of the virtual processor; The shared tag is used to determine whether the storage unit is shared with multiple graphics processors; when the storage unit is shared with multiple graphics processors, the shared tag records a first shared tag; when the storage unit is not shared with multiple graphics processors, the shared tag records a second shared tag.
5. The computing system according to claim 1, characterized in that, The graphics processor includes a streaming multiprocessor, a communication bus, a root unit, a terminal node, and a video memory control unit. The streaming multiprocessor is used to execute the computing task; The root unit is connected to the stream multiprocessor via the communication bus and is used to scan and identify the memory units and construct a global address mapping table for the memory units; The terminal node is connected to the streaming multiprocessor via the communication bus and is also used to connect to the host. The video memory control unit is connected to the streaming multiprocessor via the communication bus and is also used to connect to the video memory.
6. The computing system according to claim 5, characterized in that, The root unit includes a root complex and a root port connected thereto; The root complex is connected to the streaming multiprocessor via the communication bus; The root complex is connected to the second switching module through the root port and is used to scan the switching network topology of the second switching module to identify the storage unit; The root complex is equipped with a decoder, which is used to map the physical address of the storage unit to the system address space of the graphics processor to form the virtual address of the storage unit, so as to construct the global address mapping table and record the physical address distribution of the storage unit.
7. The computing system according to claim 5, characterized in that, The terminal node is used by the host to perform control communication and / or event notification; wherein, the control communication includes the transmission of training data to the computational model contained in the graphics processor, and the event notification indicates that the terminal node performs asynchronous event notification through an interrupt mechanism; and / or, The video memory control unit is used to perform video memory management on the video memory connected to the graphics processor; wherein, the video memory management includes interleaved access.
8. The computing system according to claim 1, characterized in that, The graphics processor is used to store copies of the computational model; When the graphics processor is training the computational model and the video memory is occupied, it acquires training data and performs training forward propagation to generate intermediate variable activation values. The graphics processor then unloads the intermediate variable activation values to the storage unit connected to it. During training backpropagation, the graphics processor retrieves the activation values of the intermediate variables carried by the storage unit, calculates the propagation gradient using the activation values of the intermediate variables, and updates the computational model using the propagation gradient.
9. The computing system according to claim 1, characterized in that, The computing system includes a computing model for performing the computing task; the computing model includes multiple tensors, and at least some of the tensors are stored in the storage space. When the graphics processor executes the computation task, it reads the tensor required to execute the computation task from the storage space as the target tensor, and uses the target tensor to execute the computation task.
10. The computing system according to claim 1, characterized in that, The computing system also includes a control module and a computing model, the computing model including one or more domain models; The graphics processor is used to locally deploy at least a portion of the domain model; The storage space is used to store and maintain model copies of all the domain models; The control module is connected to the graphics processor and is used to obtain the number of lexical units processed by the graphics processor using the domain model. In response to the number of lexical units exceeding the threshold, the domain model is designated as a model to be scheduled. The module also queries the idle domain models among the multiple graphics processors as models to be unloaded, unloads the unloaded models, and replaces the auxiliary models. The auxiliary models are models constructed using model copies of the models to be scheduled in the storage space.
11. The computing system according to claim 10, characterized in that, The control module is also configured to monitor whether the word routing to the local domain model of the graphics processor is successful; in response to the failure of the word routing to the local domain model, the control module is configured to control the word to perform inter-graphics processor interconnection data exchange through the first exchange module.
12. A computational management method, characterized in that, The computing management method is applied to the computing system as described in any one of claims 1 to 11, the computing management method comprising: The graphics processor generates a storage access request and transmits it to the second switching module; The second switching module obtains the storage access request, parses the target physical address carried in the storage access request, and queries the storage space formed by the storage module; wherein, the storage space includes multiple storage units, and the storage unit is associated with one or more graphics processors that have access to it; The second switching module routes the storage access request to the storage unit corresponding to the target physical address.
13. An electronic device, characterized in that, The electronic device includes: Equipment body; The computing system according to any one of claims 1 to 11, wherein the computing system is disposed on the device body.
Citation Information
Patent Citations
Data transmission method and device based on high-speed signal switching chip and medium
CN110401466A
GPU server and data transmission method
CN111782565A