Computer system and control method of computer system
By using the redundant design of multiple timers in AI computer systems, the high cost problems caused by different GPU types are solved, compatibility with different types of image processors is achieved, and the production and maintenance costs of computer systems are reduced.
Patent Information
- Application Number
- CN202510398395.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Due to the different types of GPUs in different computing nodes, existing AI computer systems need to design different computer system boards, resulting in higher costs.
Multiple retimers are used to divide into the first group and the second group. The number of data input channels of the first group of retimers is equal to the number of data output channels of the image processor, and the number of data input channels of the second group of retimers is greater than the number of data output channels of the image processor. Through redundant design, the computer system is compatible with different types of image processors.
This reduces the cost of computer systems, avoids the need to design separate boards for different GPU specifications, and achieves compatibility with different types of image processors.
Smart Images

Figure CN120336246A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to computer systems and control methods for computer systems. Background Art
[0002] With the development of artificial intelligence technologies, large AI (Artificial Intelligence) models have rapidly emerged, and the demand for computing resources has shown an exponential growth trend. The data scale required for the operation and training of large AI models is huge. For example, for large language models with 1.8T parameters, a vast amount of text data needs to be processed during training and inference, posing extremely high requirements for computing power. Each server manufacturer is also launching its own AI computer system designs. However, the product specifications and interconnection architectures of existing AI computer systems are relatively single and do not consider compatibility with other products, resulting in the problem that the compatibility and adaptation of multiple image processors cannot be achieved. Summary of the Invention
[0003] This application provides a computer system and a control method for a computer system to at least solve the problem in related technologies that different types of GPUs in different computing nodes require the design of different computer system boards, resulting in a relatively high cost of the computer system.
[0004] This application provides a computer system, including: computing nodes, where the computing nodes can deploy at least two types of multiple image processors; multiple retimers, where the multiple retimers are connected to the multiple image processors, the multiple retimers include a first group of retimers and a second group of retimers, the total number of data input channels of the presence signal inputs of the first group of retimers is equal to the total number of data output channels of the image processors connected to the first group of retimers, and the total number of data input channels of the presence signal inputs of the second group of retimers is greater than the total number of data output channels of the image processors connected to the first group of retimers; a switching node, where the switching node is connected to the output ends of the multiple retimers, and interconnection between the computing nodes and other computing nodes is achieved through the switching node.
[0005] This application also provides a control method for a computer system, including: obtaining first data transmission parameters of physical ports of multiple image processors in the computer system and second data transmission parameters corresponding to the switching node of the computer system; adjusting the first data transmission parameters through the retimers in the computer system to adapt to the second data transmission parameters.
[0006] The present application also provides a control device for a computer system, including: an acquisition unit, configured to acquire first data transmission parameters of physical ports of multiple image processors of the computer system and second data transmission parameters corresponding to a switching node of the computer system; and an adjustment unit, configured to adjust the first data transmission parameters through a retimer in the computer system to adapt to the second data transmission parameters.
[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; and a processor, configured to implement the steps of any one of the above computer system control methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any one of the above computer system control methods when executed by a processor.
[0009] The present application also provides a computer program product including a computer program, and the computer program implements the steps of any one of the above computer system control methods when executed by a processor.
[0010] Through the present application, for compatibility with different types of image processors, multiple retimers are divided into a first group of retimers and a second group of retimers. The total number of all data input channels corresponding to the first group of retimers is equal to the total number of first data output channels of multiple image processors, while the total number of all data input channels of the second group of retimers is greater than or equal to the total number of second data output channels of multiple image processors. Through the redundant setting of the second group of retimers, the computing nodes in the computer system provided by the embodiments of the present application can be compatible with different types of image processors, without the need to design different connection relationships according to different types of image processors. Therefore, the problem in the related art that different computer system boards need to be designed due to different types of GPUs in different computing nodes, resulting in a relatively high cost of the computer system, is solved. Through the redundant design of the retimer, the need for separate board design for different GPU specifications is eliminated, thereby achieving the technical effect of reducing the cost of the computer system. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 Schematic diagram of a computer system provided by an embodiment of the present application Figure 1 ;
[0013] Figure 2 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 1 ;
[0014] Figure 3 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 2 ;
[0015] Figure 4 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 3 ;
[0016] Figure 5 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 4 ;
[0017] Figure 6 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 5 ;
[0018] Figure 7 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 6 ;
[0019] Figure 8 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 7 ;
[0020] Figure 9 A schematic diagram of a computer system provided by an embodiment of the present application Figure 2 ;
[0021] Figure 10 A schematic diagram of a retimer connection provided by an embodiment of the present application Figure 8 ;
[0022] Figure 11 A flowchart of a control method for a computer system provided by an embodiment of the present application;
[0023] Figure 12 A schematic diagram of a control device for a computer system provided by an embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0025] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0026] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0027] An embodiment of this application provides a computer system, and the computer system will be described in detail in combination with the components of the computer system. As Figure 1 shown, the computer system provided by the embodiment of this application includes: a computing node 10, a plurality of retimers 11, and a switching node 12.
[0028] The computing node 10, wherein the computing node can deploy at least two types of a plurality of image processors;
[0029] A plurality of retimers 11, wherein the plurality of retimers are connected to the plurality of image processors, the plurality of retimers include a first group of retimers and a second group of retimers, the total number of data input channels where the presence signals of the first group of retimers are input is equal to the total number of the first data output channels of the plurality of image processors, and the total number of data input channels where the presence signals of the second group of retimers are input is greater than or equal to the total number of the second data output channels of the plurality of image processors;
[0030] The switching node 12, wherein the switching node is connected to the output ends of the plurality of retimers, and the interconnection between the computing node and other computing nodes is realized through the switching node.
[0031] Optionally, as Figure 1 shown, the computer system provided by the embodiment of this application includes: a computing node 10, a plurality of retimers 11, and a switching node 12, wherein the computing node 10 is connected to the plurality of retimers 11, and the plurality of retimers 11 and the switching node 12 are connected. It should be noted that in the computer system provided by the embodiment of this application, the number of computing nodes can be multiple, and the interconnection between the computing nodes can be realized through the switching node 12.
[0032] In the computing node 10, at least two types of multiple image processors can be deployed. In an optional embodiment, the number of multiple image processors can be at least two. The image processor can be an OAM GPU (Open Accelerator Module Graphics Processing Unit). The OAM GPU is a graphics processor module designed based on the OAM (Open Accelerator Module) standard.
[0033] It should be noted that there are differences in the data transfer rate of the physical ports and the number of data transfer channels of different types of image processors. For example, the serdes rate of the first type of image processor is 56 Gbps, and there are a total of 8 physical ports (where serdes1 is divided into serdes1L and serdes 1H), and each physical port is 8 lanes. The serdes rate of the second type of image processor is 112 Gbps, and there are a total of 8 physical ports (serdes1 is divided into serdes1L and serdes 1H), and each physical port is 6 lanes. The serdes rate of the third type of image processor is 224 Gbps, and there are a total of 8 physical ports (where serdes1 is divided into serdes1L and serdes 1H), and each physical port is 8 lanes.
[0034] Serdes, Serializer / Deserializer, is a technology for high-speed data communication, which is used to convert parallel data into serial data for transmission (called serializer at the sending end), and then convert the serial data back into parallel data at the receiving end (called deserializer).
[0035] It should be noted that, in order to be compatible with different types of image processors, multiple re-timers can be divided into a first group of re-timers and a second group of re-timers. Among them, the total number of all data input channels corresponding to the first group of re-timers is equal to the total number of the first data output channels of multiple image processors, while the total number of all data input channels of the second group of re-timers is greater than or equal to the total number of the second data output channels of multiple image processors. That is to say, there is a redundant design for the second group of re-timers. Through the redundant setting of the second group of re-timers, the computing nodes in the computer system provided by the embodiments of the present application can be compatible with different types of image processors, without the need to design different connection relationships according to different types of image processors. Therefore, the problem in the related art that different computer system boards need to be designed due to different types of GPUs in different computing nodes, resulting in a relatively high cost of the computer system, is solved. Through the redundant design of the re-timers, the need for separate board design for different GPU specifications is eliminated, and thus the technical effect of reducing the cost of the computer system is achieved.
[0036] Optionally, in the computer system provided by the embodiments of the present application, the first group of re-timers is connected to multiple first data output channels of physical ports in multiple image processors, and the second group of re-timers is connected to multiple second data output channels of physical ports in multiple image processors. Among them, based on the number of data output channels corresponding to one logical port of the image processor, the data output channels of the physical port in one image processor are divided into multiple first data output channels and multiple second data output channels.
[0037] In an alternative embodiment, the first group of re-timers is connected to multiple first data output channels of physical ports in multiple image processors, and the second group of re-timers is connected to multiple second data output channels of physical ports in multiple image processors.
[0038] It should be noted that one re-timer is correspondingly connected to 2 image processors. For example, multiple image processors are GPU0 and GPU1. One re-timer in the corresponding first group of re-timers is correspondingly connected to some data output channels of GPU0 and GPU1. When a certain re-timer is damaged, it will only affect the data transmission of some data output channels of the image processors, and will not affect the data transmission of other data output channels. Therefore, by connecting one re-timer to 2 image processors, when the re-timer fails, the impact on the image processors can be reduced.
[0039] It should be noted that based on the number of data output channels corresponding to a logical port of the image processor, the number of data output channels corresponding to an image processor is divided into a first data output channel and a second data output channel. For example, the number of data output channels corresponding to a logical port of the first type of image processor is 4, the corresponding first data output channels of the image processor are Lane0-3, and the corresponding second data output channels are Lane4-7. If the number of data output channels corresponding to a logical port of the second type of image processor is 2, the corresponding first data output channels of the image processor are Lane0-1, and the corresponding second data output channels are Lane2-3. In the application, in order to achieve compatibility of different types of image processors, the first data output channels corresponding to the image processor are set as Lane0-3, and the corresponding second data output channels are set as Lane4-7.
[0040] According to the connection of one retimer corresponding to 2 image processors and the division of data output channels as described above, in order to be compatible with different types of image processors, the first group of retimers is connected to Lane0-3 of the 2 physical ports of the 2 image processors, and the second group of retimers is connected to Lane4-7 of the 2 physical ports of the 2 image processors.
[0041] Through the intelligent division of data output channels and the precise connection strategy of retimers, compatibility of different types of image processors is achieved.
[0042] Optionally, in the computer system provided in the embodiment of the present application, for the connection between the first group of retimers and multiple first data output channels of physical ports in multiple image processors: when the multiple image processors are the first type of image processors, all data input channels at the input end of the first group of retimers are connected to the multiple first data output channels, and the output signal of the first group of retimers is output from the target data output channel at the output end of the first group of retimers. Among them, the output data channels of the first type of image processors are merged through the first group of retimers to adapt to the switching node; when the multiple image processors are the second type of image processors, all data input channels at the input end of the first group of retimers are connected to the multiple first data output channels, and the output signal of the first group of retimers is output from all data output channels at the output end of the first group of retimers.
[0043] In an optional embodiment, taking the compatibility of the first type of image processor and the second type of image processor as an example to illustrate the specific connection relationship of the first group of retimers, the comparison of transmission parameters of the first type of image processor and the second type of image processor is shown in Table 1.
[0044] Table 1
[0045]
[0046] As described in Table 1, the serdes rate of the first type of image processor is 56 Gbps. An OAM GPU has a total of 8 physical ports (where serdes1 is divided into serdes1L and serdes 1H), and each physical port has 8 lanes. From the perspective of the Ethernet interface, every 4 lanes form a port (logical port), that is, a total of 16 ports. The serdes rate of the first type of image processor is 112 Gbps. An OAM GPU is also divided into 8 physical ports, but each physical port has only 6 lanes. From the perspective of the Ethernet interface, every 2 lanes form a port (MAC), and each GPU has a total of 24 ports. The full interconnection of the OAM GPU depends on the Ethernet switch. As mentioned above, the OAM GPUs from different manufacturers have different rates, and the single-lane rate adopted by the switching node is 112 Gbps. To achieve the interconnection of the first type of image processor, the second type of image processor, and the switching node, multiple of the above-mentioned retimers need to be set between the OAM GPU and the switching node to achieve rate matching. The specific connection relationship is as follows:
[0047] When multiple image processors are of the first type, all data input channels at the input end of the first group of retimers are connected to multiple first data output channels, and the output signal of the first group of retimers is output from the target data output channels at the output end of the first group of retimers.
[0048] For example, for a retimer in the first group of retimers, all data input channels (Lane0 - 15) of this retimer are connected to Lane0 - 3 of the S1 port and Lane0 - 3 of the S2 port of GPU0, as well as Lane0 - 3 of the S1 port and Lane0 - 3 of the S2 port of GPU1. The above-mentioned target data output channels can be Lane0 - 3 and Lane8 - 11 at the output end of the retimer, or Lane0 - 7.
[0049] As Figure 2 shown, the Host side is connected to the OAM GPU (i.e., the above-mentioned image processor), and the Line side is connected to the switching node. The serdes signal rate of the Host side is 56 Gbps, while the serdes signal rate of the Line side is 112 Gbps. In the case where the above-mentioned target data output channels are Lane0 - 3 and Lane8 - 11 of the retimer, the connection relationship is as follows.
[0050] As Figure 2As shown, GPU0_S1_Lane0 and GPU0_S1_Lane1 are merged through Lane0 and Lane1 on the Host side of the re-timer to Lane0 on the Line side of the re-timer to output GPU0_S1_Lane0 / 1; GPU0_S1_Lane2 and GPU0_S1_Lane3 are merged through Lane2 and Lane3 on the Host side of the re-timer to Lane1 on the Line side of the re-timer to output GPU0_S1_Lane2 / 3; GPU0_S2_Lane0 and GPU0_S2_Lane1 are merged through Lane4 and Lane5 on the Host side of the re-timer to Lane2 on the Line side of the re-timer to output GPU0_S2_Lane0 / 1; GPU0_S2_Lane2 and GPU0_S2_Lane3 are merged through Lane6 and Lane7 on the Host side of the re-timer to Lane3 on the Line side of the re-timer to output GPU0_S2_Lane2 / 3.
[0051] GPU1_S1_Lane0 and GPU1_S1_Lane1 are merged through Lane8 and Lane9 on the Host side of the re-timer to Lane8 on the Line side of the re-timer to output GPU1_S1_Lane0 / 1; GPU1_S1_Lane2 and GPU1_S1_Lane3 are merged through Lane10 and Lane11 on the Host side of the re-timer to Lane9 on the Line side of the re-timer to output GPU1_S1_Lane2 / 3; GPU1_S2_Lane0 and GPU1_S2_Lane1 are merged through Lane12 and Lane13 on the Host side of the re-timer to Lane10 on the Line side of the re-timer to output GPU1_S2_Lane0 / 1; GPU1_S2_Lane2 and GPU1_S2_Lane3 are merged through Lane14 and Lane15 on the Host side of the re-timer to Lane11 on the Line side of the re-timer to output GPU1_S2_Lane2 / 3.
[0052] In the case where the above-mentioned target data output channels are Lane0-7 of the re-timer, the connection relationship is as described below.
[0053] As Figure 3As shown, GPU0_S1_Lane0 and GPU0_S1_Lane1 are merged through Lane0 and Lane1 on the Host side of the re-timer to Lane0 on the Line side of the re-timer to output GPU0_S1_Lane0 / 1; GPU0_S1_Lane2 and GPU0_S1_Lane3 are merged through Lane2 and Lane3 on the Host side of the re-timer to Lane1 on the Line side of the re-timer to output GPU0_S1_Lane2 / 3; GPU0_S2_Lane0 and GPU0_S2_Lane1 are merged through Lane4 and Lane5 on the Host side of the re-timer to Lane2 on the Line side of the re-timer to output GPU0_S2_Lane0 / 1; GPU0_S2_Lane2 and GPU0_S2_Lane3 are merged through Lane6 and Lane7 on the Host side of the re-timer to Lane3 on the Line side of the re-timer to output GPU0_S2_Lane2 / 3.
[0054] GPU1_S1_Lane0 and GPU1_S1_Lane1 are merged through Lane8 and Lane9 on the Host side of the re-timer to Lane4 on the Line side of the re-timer to output GPU1_S1_Lane0 / 1; GPU1_S1_Lane2 and GPU1_S1_Lane3 are merged through Lane10 and Lane11 on the Host side of the re-timer to Lane5 on the Line side of the re-timer to output GPU1_S1_Lane2 / 3; GPU1_S2_Lane0 and GPU1_S2_Lane1 are merged through Lane12 and Lane13 on the Host side of the re-timer to Lane6 on the Line side of the re-timer to output GPU1_S2_Lane0 / 1; GPU1_S2_Lane2 and GPU1_S2_Lane3 are merged through Lane14 and Lane15 on the Host side of the re-timer to Lane7 on the Line side of the re-timer to output GPU1_S2_Lane2 / 3.
[0055] From the above analysis, it can be seen that the rate matching between the first type of image processor and the switching node is achieved through the re-timer.
[0056] When multiple image processors are of the second type, since the transmission rate of the second type of image processor is 112 Gbps, there is no need to adjust the number of the first data output channels of the second type of image processor. Therefore, all data input channels at the input end of the first group of retimers are connected to the multiple first data output channels, and the output signals of the first group of retimers are output by all data output channels at the output end of the first group of retimers.
[0057] In an alternative embodiment, when multiple image processors are of the second type, the corresponding connection relationship is as Figure 4 shown. GPU0_S1_Lane0 and GPU0_S1_Lane1 are mapped to Lane0 and Lane1 on the Line side of the re-timer through Lane0 and Lane1 on the Host side of the re-timer to output GPU0_S1_Lane0 and GPU0_S1_Lane1; GPU0_S1_Lane2 and GPU0_S1_Lane3 are mapped to Lane2 and Lane3 on the Line side of the re-timer through Lane2 and Lane3 on the Host side of the re-timer to output GPU0_S1_Lane2 and GPU0_S1_Lane3; GPU0_S2_Lane0 and GPU0_S2_Lane1 are mapped to Lane4 and Lane5 on the Line side of the re-timer through Lane4 and Lane5 on the Host side of the re-timer to output GPU0_S2_Lane0 and GPU0_S2_Lane1. GPU0_S2_Lane2 and GPU0_S2_Lane3 are mapped to Lane6 and Lane7 on the Line side of the re-timer through Lane6 and Lane7 on the Host side of the re-timer to output GPU0_S2_Lane2 and GPU0_S2_Lane3.
[0058] GPU1_S1_Lane0 and GPU1_S1_Lane1 are mapped to Lane8 and Lane9 on the Line side of the re-timer through Lane8 and Lane9 on the Host side of the re-timer to output GPU1_S1_Lane0 and GPU1_S1_Lane1; GPU1_S1_Lane2 and GPU1_S1_Lane3 are mapped to Lane10 and Lane11 on the Line side of the re-timer through Lane10 and Lane11 on the Host side of the re-timer to output GPU1_S1_Lane2 and GPU1_S1_Lane3; GPU1_S2_Lane0 and GPU1_S2_Lane1 are mapped to Lane12 and Lane13 on the Line side of the re-timer through Lane12 and Lane13 on the Host side of the re-timer to output GPU1_S2_Lane0 and GPU1_S2_Lane1. GPU1_S2_Lane2 and GPU1_S2_Lane3 are mapped to Lane14 and Lane15 on the Line side of the re-timer through Lane14 and Lane15 on the Host side of the re-timer to output GPU1_S2_Lane2 and GPU1_S2_Lane3.
[0059] From the above analysis, it can be seen that since the transmission rate of the second type of image processor is 112 Gbps, there is no need to adjust the number of the first data output channels of the second type of image processor. Therefore, it can be directly remapped to the switching node through the re-timer.
[0060] Through the above connection relationship, the compatibility of the first data output channels of the first type of image processor and the second type of image processor is achieved. There is no need to change the connection relationship between the image processor and the re-timer according to different types of image processors, avoiding the need to produce different dedicated boards for different image processors in the prior art, and thus achieving the technical effect of reducing the production cost and maintenance cost of the computer system.
[0061] Optionally, in the computer system provided in the embodiments of the present application, for the connection between the second set of retimers and the multiple second data output channels of the physical ports in the multiple image processors: when the multiple image processors are the first type of image processors, all data input channels at the input end of the second set of retimers are connected to the multiple second data output channels, and the output signal of the second set of retimers is output from the target data output channel among the output ends of the second set of retimers. Among them, the data output channels of the first type of image processors are merged through the second set of retimers to adapt to the switching node; when the multiple image processors are the second type of image processors, the target data input channel among the input ends of the second set of retimers is connected to the multiple second data output channels, and the output signal of the second set of retimers is output from the target data output channel among the output ends of the second set of retimers.
[0062] In an optional embodiment, the connection relationship between the second set of retimers and the second data output channels is as described below. Since the transmission rate of the first type of image processors is less than that of the switching node, it is necessary to merge the data output channels. Although it is not necessary to merge the data output channels for the second type of image processors, the second data output channels corresponding to the second type of image processors do not have Lane6 and Lane7. That is, after adjusting the number of transmission channels of the first type of image processors through the retimers, for both the first type of image processors and the second type of image processors, the number of data output channels of the output signal of the retimers is the same, both being 8. Therefore, for the convenience of setting the connection relationship between the subsequent retimers and the high-density connectors, for both the first type of image processors and the second type of image processors, the retimers need to output signals using the same data output channels. That is, for the first type of image processors, all data input channels at the input end of the second set of retimers are connected to the multiple second data output channels, and the output signal of the second set of retimers is output from the target data output channel among the output ends of the second set of retimers. For the second type of image processors, the target data input channel among the input ends of the second set of retimers is connected to the multiple second data output channels, and the output signal of the second set of retimers is output from the target data output channel among the output ends of the second set of retimers.
[0063] It should be noted that the target data output channels can be Lane0-3, Lane8-11 at the output end of the retimer, or Lane0-7. In the first type of image processor, when the target data output channels are Lane0-3, Lane8-11 at the output end of the retimer, the corresponding target data output channels of the second type of image processor also need to be Lane0-3, Lane8-11 at the output end of the retimer, that is, the target data output channels adopted by the first type of image processor and the second type of image processor at the output end of the retimer need to be the same.
[0064] It should be noted that the target data input channels can be Lane0-1, Lane4-5, Lane8-9, Lane12-13 at the input end of the retimer, or Lane0-7. As long as it is ensured that the target data output channels adopted by the first type of image processor and the second type of image processor at the output end of the retimer need to be the same, there is no need to specifically limit which Lanes the target data input channels are.
[0065] Taking one retimer in the second group of retimers as an example to illustrate the above connection relationship. For the first type of image processor, when the target data output channels are Lane0-3, Lane8-11 at the output end of the retimer, as Figure 5 shown, the re-timer doubles the 4lane 56Gbps signal on the Host side to a 2lane 112Gbps signal on the Line side. The re-timer connects all data output channels (Lane0-15) to Lane4-7 of the S1 port and Lane4-7 of the S2 port of GPU0, and also connects to Lane4-7 of the S1 port and Lane4-7 of the S2 port of GPU1.
[0066] As Figure 5, GPU0_S1_Lane4 and GPU0_S1_Lane5 are merged through Lane0 and Lane1 on the Host side of the re-timer to Lane0 on the Line side of the re-timer to output GPU0_S1_Lane4 / 5; GPU0_S1_Lane6 and GPU0_S1_Lane7 are merged through Lane2 and Lane3 on the Host side of the re-timer to Lane1 on the Line side of the re-timer to output GPU0_S1_Lane6 / 7; GPU0_S2_Lane4 and GPU0_S2_Lane5 are merged through Lane4 and Lane5 on the Host side of the re-timer to Lane2 on the Line side of the re-timer to output GPU0_S2_Lane4 / 5; GPU0_S2_Lane6 and GPU0_S2_Lane7 are merged through Lane6 and Lane7 on the Host side of the re-timer to Lane3 on the Line side of the re-timer to output GPU0_S2_Lane6 / 7.
[0067] GPU1_S1_Lane4 and GPU1_S1_Lane5 are merged through Lane8 and Lane9 on the Host side of the re-timer to Lane8 on the Line side of the re-timer to output GPU1_S1_Lane4 / 5; GPU1_S1_Lane6 and GPU1_S1_Lane7 are merged through Lane10 and Lane11 on the Host side of the re-timer to Lane9 on the Line side of the re-timer to output GPU1_S1_Lane6 / 7; GPU1_S2_Lane4 and GPU1_S2_Lane5 are merged through Lane12 and Lane13 on the Host side of the re-timer to Lane10 on the Line side of the re-timer to output GPU1_S2_Lane4 / 5; GPU1_S2_Lane6 and GPU1_S2_Lane7 are merged through Lane14 and Lane15 on the Host side of the re-timer to Lane11 on the Line side of the re-timer to output GPU1_S2_Lane6 / 7.
[0068] For the first type of image processor, when the target data output channel is Lane0-7 at the output end of the re-timer, as Figure 6As shown, GPU0_S1_Lane4 and GPU0_S1_Lane5 are merged through Lane0 and Lane1 on the Host side of the re-timer to Lane0 on the Line side of the re-timer to output GPU0_S1_Lane4 / 5; GPU0_S1_Lane6 and GPU0_S1_Lane7 are merged through Lane2 and Lane3 on the Host side of the re-timer to Lane1 on the Line side of the re-timer to output GPU0_S1_Lane6 / 7; GPU0_S2_Lane4 and GPU0_S2_Lane5 are merged through Lane4 and Lane5 on the Host side of the re-timer to Lane2 on the Line side of the re-timer to output GPU0_S2_Lane4 / 5; GPU0_S2_Lane6 and GPU0_S2_Lane7 are merged through Lane6 and Lane7 on the Host side of the re-timer to Lane3 on the Line side of the re-timer to output GPU0_S2_Lane6 / 7.
[0069] GPU1_S1_Lane4 and GPU1_S1_Lane5 are merged through Lane8 and Lane9 on the Host side of the re-timer to Lane4 on the Line side of the re-timer to output GPU1_S1_Lane4 / 5; GPU1_S1_Lane6 and GPU1_S1_Lane7 are merged through Lane10 and Lane11 on the Host side of the re-timer to Lane5 on the Line side of the re-timer to output GPU1_S1_Lane6 / 7; GPU1_S2_Lane4 and GPU1_S2_Lane5 are merged through Lane12 and Lane13 on the Host side of the re-timer to Lane6 on the Line side of the re-timer to output GPU1_S2_Lane4 / 5; GPU1_S2_Lane6 and GPU1_S2_Lane7 are merged through Lane14 and Lane15 on the Host side of the re-timer to Lane7 on the Line side of the re-timer to output GPU1_S2_Lane6 / 7.
[0070] For the second type of image processor, the re-timer connects all data output channels (Lane0-15) to Lane4-5 of the S1 port and Lane4-5 of the S2 port of GPU0, and also connects to Lane4-5 of the S1 port and Lane4-5 of the S2 port of GPU1. When paired with the second type of image processor, lanes 6 / 7 of each physical port have no signal, and the serdes rate is 112 Gbps. Therefore, lane swap (transmission channel remapping) is performed within the re-timer chip. As Figure 7 shown, in the case where the target data output channels are Lane0-3 and Lane8-11 at the output end of the re-timer, GPU0_S1_Lane4 and GPU0_S1_Lane5 are mapped to Lane0 and Lane1 on the Line side of the re-timer through Lane0 and Lane1 on the Host side of the re-timer to output GPU0_S1_Lane4 and GPU0_S1_Lane5; GPU0_S2_Lane4 and GPU0_S2_Lane5 are mapped to Lane2 and Lane3 on the Line side of the re-timer through Lane2 and Lane3 on the Host side of the re-timer to output GPU0_S2_Lane4 and GPU0_S2_Lane5.
[0071] GPU1_S1_Lane4 and GPU1_S1_Lane5 are mapped to Lane8 and Lane9 on the Line side of the re-timer through Lane8 and Lane9 on the Host side of the re-timer to output GPU1_S1_Lane4 and GPU1_S1_Lane5; GPU1_S2_Lane4 and GPU1_S2_Lane5 are mapped to Lane10 and Lane11 on the Line side of the re-timer through Lane12 and Lane13 on the Host side of the re-timer to output GPU1_S2_Lane4 and GPU1_S2_Lane5.
[0072] As Figure 8As shown, when the target data output channels are the output ends of the re-timers of Lane0-7, GPU0_S1_Lane4 and GPU0_S1_Lane5 are mapped to the output of Lane0 and Lane1 on the Line side of the re-timer through Lane0 and Lane1 on the Host side of the re-timer to output GPU0_S1_Lane4 and GPU0_S1_Lane5; GPU0_S2_Lane4 and GPU0_S2_Lane5 are mapped to the output of Lane2 and Lane3 on the Line side of the re-timer through Lane2 and Lane3 on the Host side of the re-timer to output GPU0_S2_Lane4 and GPU0_S2_Lane5.
[0073] GPU1_S1_Lane4 and GPU1_S1_Lane5 are mapped to the output of Lane4 and Lane5 on the Line side of the re-timer through Lane8 and Lane9 on the Host side of the re-timer to output GPU1_S1_Lane4 and GPU1_S1_Lane5; GPU1_S2_Lane4 and GPU1_S2_Lane5 are mapped to the output of Lane6 and Lane7 on the Line side of the re-timer through Lane12 and Lane13 on the Host side of the re-timer to output GPU1_S2_Lane4 and GPU1_S2_Lane5.
[0074] Through the above connections, the compatibility of the second data output channels of the first type of image processor and the second type of image processor is achieved. There is no need to change the connection relationship between the image processor and the re-timer according to different types of image processors, avoiding the need to produce different dedicated boards for different image processors in the prior art, thereby achieving the technical effect of reducing the production cost and maintenance cost of the computer system.
[0075] Optionally, in the computer system provided in the embodiment of the present application, the computer system further includes: a plurality of high-density connectors, wherein the plurality of high-density connectors are connected between the re-timer and the switching node, and the number of the plurality of high-density connectors has a corresponding relationship with the data output channels with data output among the output ends of the plurality of re-timers.
[0076] In an alternative embodiment, in the computer system provided in the embodiment of the present application, as Figure 9As shown in the figure, the computer system further includes: a plurality of high-density connectors connected between the retimer and the switching node for transmitting the adjusted signals of the retimer to the switching node, thereby enabling interconnection between multiple computing nodes. It should be noted that the number of the plurality of high-density connectors has a corresponding relationship with the data output channels having data output among the output ends of the plurality of retimers, that is, the number of the plurality of high-density connectors needs to be adapted according to the data output channels having data output among the output ends of the plurality of retimers. For example, if there are 64 data output channels having data output, 2 8X8 high-density connectors can be selected.
[0077] By limiting the number of high-density connectors, while ensuring the signal transmission requirements, over-configuration of connectors is avoided, and the hardware cost is reduced. By reducing unnecessary high-density connectors, the material cost and assembly complexity can be effectively controlled.
[0078] Optionally, in the computer system provided in the embodiment of the present application, for the connection relationship between the plurality of high-density connectors and the plurality of retimers: all data output channels in the output ends of the first group of retimers are connected to the first high-density connector and the second high-density connector among the plurality of high-density connectors; the target data output channels in the output ends of the second group of retimers are connected to the third high-density connector among the plurality of high-density connectors.
[0079] In an optional embodiment, through the above Figures 5 - 8 analysis, it can be known that for the first data output channel of the first type of image processor and the first data output channel of the second type of image processor, the number of data output channels having data output of the retimer is different. A second type of image processor has 16 more data output channels than a first type of image processor. In order not to require an excessive number of high-density connectors when using the first type of image processor, the extra 16 data output channels need to be connected to a single high-density connector. Therefore, all data output channels in the output ends of the first group of retimers are connected to the first high-density connector and the second high-density connector among the plurality of high-density connectors; the target data output channels in the output ends of the second group of retimers are connected to the third high-density connector among the plurality of high-density connectors. For example, CH0-15 corresponding to all data output channels in the output ends of the first group of retimers is connected to the first high-density connector and the second high-density connector, while CH0-3 and CH8-11 corresponding to the target data output channels (i.e., the above-mentioned Lan0-3, Lan8-11) in the output ends of the second group of retimers are connected to the third high-density connector.
[0080] It should be noted that if the target data output channels correspond to Lan0-7, then the CH0-7 corresponding to Lan0-7 need to be connected to the third high-density connector.
[0081] By reasonably allocating the output channels of the retimer to different high-density connectors, unnecessary hardware usage is reduced, and thus the technical effect of cost reduction is achieved.
[0082] Optionally, in the computer system provided in the embodiment of the present application, regarding the connection relationship between all the data output ports in the output end of the first group of retimers and the first high-density connector and the second high-density connector among multiple high-density connectors: the target data output channel in the output end of the first group of retimers is connected to the first high-density connector; the data output channels other than the target data output channel in the output end of the first group of retimers are connected to the second high-density connector.
[0083] In an optional embodiment, the connection relationship between all the data output ports in the output end of the first group of retimers and the first high-density connector and the second high-density connector among multiple high-density connectors is as follows. The target data output channels (for example, CH0-3 and CH8-11 corresponding to Lan0-3, Lan8-11) in the output end of the first group of retimers are connected to the first high-density connector; the data output channels other than the target data output channel (for example, CH4-7 and CH12-15 corresponding to Lan4-7, Lan12-15) in the output end of the first group of retimers are connected to the second high-density connector. Through the above specific connection relationship, the wiring difficulty between the high-density connector and the retimer can be effectively reduced.
[0084] Optionally, in the computer system provided in the embodiment of the present application, when the multiple image processors are the first type of image processors, the first high-density connector and the third high-density connector are in the conducting state, and the second high-density connector is in the non-conducting state; when the multiple image processors are the second type of image processors, the first high-density connector, the second high-density connector, and the third high-density connector are all in the conducting state.
[0085] In an optional embodiment, since the extra 16 data output channels are separately connected to a high-density connector, when the multiple image processors are the first type of image processors, the first high-density connector and the third high-density connector are in the conducting state, and the second high-density connector is in the non-conducting state; when the multiple image processors are the second type of image processors, the first high-density connector, the second high-density connector, and the third high-density connector are all in the conducting state.
[0086] In an optional embodiment, taking the number of image processors as two, and the two types of image processors being the first type of image processor and the second type of image processor as an example, the topological connection relationship of the corresponding computer system is as Figure 10 shown. Among them, the two image processors are Figure 10 the OAM GPU0 and OAM GPU1 in
[0087] , and the 8 re-timers are multiple retimers, namely retimer0-retimer7. retimer0-retimer3 are a group (i.e., the first group of retimers mentioned above), and retimer4-retimer7 are a group (i.e., the second group of retimers mentioned above). For retimer0-retimer3, lanes 4-7 of the physical port are connected, and for retimer4-retimer7, lanes 0-3 of the physical port are connected. 8x8connector_1, 8x8connector_2, and 8x8connector_3 correspond to the first high-density connector, the second high-density connector, and the third high-density connector. It should be noted that for lanes 8-15 of the S1 port, it is equivalent to lanes 0-7 of a port. Figure 10As shown, lane0 - 3 and lane8 - 11 of the S1 port of OAM GPU0 and lane0 - 3 and lane8 - 11 of the S1 port of OAM GPU1 are connected to retimer4, and lane4 - 7 and lane12 - 15 of the S1 port of OAM GPU0 and lane4 - 7 and lane12 - 15 of the S1 port of OAM GPU1 are connected to retimer0; lane0 - 3 of the S2 and S3 ports of OAM GPU0 and lane0 - 3 of the S2 and S3 ports of OAM GPU1 are connected to retimer5; lane4 - 7 of the S2 and S3 ports of OAM GPU0 and lane4 - 7 of the S2 and S3 ports of OAM GPU1 are connected to retimer1; lane0 - 3 of the S4 and S5 ports of OAM GPU0 and lane0 - 3 of the S4 and S5 ports of OAM GPU1 are connected to retimer6; lane4 - 7 of the S4 and S5 ports of OAM GPU0 and lane4 - 7 of the S4 and S5 ports of OAM GPU1 are connected to retimer2; lane0 - 3 of the S6 and S7 ports of OAM GPU0 and lane0 - 3 of the S6 and S7 ports of OAM GPU1 are connected to retimer7; lane4 - 7 of the S6 and S7 ports of OAM GPU0 and lane4 - 7 of the S6 and S7 ports of OAM GPU1 are connected to retimer3. The corresponding output CH0 - CH3 and CH8 - CH11 of retimer0 - 3 are connected to 8x8connector_3, the corresponding output CH0 - CH3 and CH8 - CH11 of retimer4 - 7 are connected to 8x8connector_1, and the corresponding output CH4 - CH7 and CH12 - CH15 of retimer4 - 7 are connected to 8x8connector_2.
[0088] It should be noted that the number of multiple image processors can also be 4. The 4 image processors can be divided into 2 groups, each group consisting of 2 image processors. For the 2 image processors in each group: the connection relationship among the image processor, the retimer, and the high - density connector is the same, and can all be the topological connection relationship as Figure 10 shown.
[0089] It should be noted that for these two groups of image processors, the same type of image processor can be used. For example, the first type of image processor can be used for both groups. Different types of image processors can also be used. For example, the first group of image processors uses the first type of image processor, and the second group of image processors uses the second type of image processor. When the same type of image processor is used for these two groups of image processors (for example, the first type of re-timer is used for both), the re-timer processes the data output channels of these 4 image processors in the same way (for example, merging 2 lanes into 1 lane). When different types of image processors are used for these two groups of image processors, the re-timer needs to determine the processing method for the data output channels according to the type of the image processor. For example, the re-timer merges the data output channels for the first type of image processor, and the re-timer performs lane swap (transmission channel remapping) for the second type of image processor.
[0090] Through the topological connection relationship between the image processors (OAM GPU0 and OAM GPU1), the re-timer, and the high-density connectors (HD connectors), it is possible to support both the first type and the second type of image processors simultaneously. No matter which GPU is connected, the best data transmission performance can be achieved by adjusting the configuration of the re-timer and the conduction state of the connectors. This highly compatible and general design strategy avoids the problem of designing different computer system boards due to different types of GPUs in different computing nodes, thereby achieving the effect of reducing the cost of the computer system.
[0091] The embodiment of the present application also provides a control method for a computer system. The control method for the computer system is applied to the computer system of any one of the above, as Figure 11 shown. The control method for the computer system includes:
[0092] Step S1101, obtaining the first data transmission parameters of the physical ports of multiple image processors in the computer system and the second data transmission parameters corresponding to the switching nodes of the computer system.
[0093] Optionally, obtain the first data transmission parameters of the physical ports of multiple image processors in the computer system. The first data transmission parameters include, but are not limited to, the serdes rate (transmission rate). These parameters may vary depending on the manufacturer and model of the image processor. In an optional embodiment, the multiple image processors may include a first type of image processor and a second type of image processor, corresponding to the products of Company A and Company B respectively. For example, the GPUs in the products of Company A and Company B have different serdes rates. At the same time, it is also necessary to obtain the second data transmission parameters corresponding to the switching node in the computer system.
[0094] Step S1102, adjust the first data transmission parameters through the retimer in the computer system to adapt to the second data transmission parameters.
[0095] Optionally, after obtaining the first data transmission parameters and the second data transmission parameters, it is necessary to determine the adjustment strategy of the retimer. That is, based on the adjustment strategy, the retimer adjusts the first data transmission parameters to achieve the purpose of adapting to the second data transmission parameters. For example, if the serdes rate of multiple image processors is 56 Gbps, while the rate supported by the switching node is 112 Gbps, the retimer will combine the 4-lane signals from the GPU into 2-lane signals to increase the transmission rate. If the serdes rate of multiple image processors is 112 Gbps, but in order to be able to be compatible with multiple image processors at the same time, it is necessary to remap the data output channels of multiple image processors through the retimer.
[0096] The parameter adjustment strategy of the retimer enables the computer system to be compatible with multiple different types of image processors, such as the GPUs of Company A and Company B, reducing the hardware cost. At the same time, it avoids designing dedicated boards for different types of GPUs, further reducing the R & D and production costs. Therefore, the control method of the computer system provided in this application effectively solves the signal compatibility problem between different image processors and realizes efficient data transmission between different image processors and the switching node.
[0097] In an optional embodiment, the first data transmission parameters at least include the transmission rate, and may also include the number of channels of the transmission channels of the physical ports of the image processors. It should be noted that the transmission rate here refers to the serdes rate of the physical ports of the image processors, which is the speed at which data is transmitted between the GPU and the retimer board. Different image processors, such as the GPUs of Company A and Company B, may have different serdes rates, such as 56 Gbps or 112 Gbps.
[0098] When the first data transmission parameter includes the transmission rate, the transmission rate of the image processor is adjusted and matched to ensure compatible data exchange with the switching node in the computer system. Specifically, through a re-timer, the transmission rate of the image processor can be adjusted to meet the transmission rate requirements of the switching node. For example, if the serdes rate of the GPU is 56 Gbps and the switching node only supports a rate of 112 Gbps, the re-timer can combine the GPU's 4-lane 56 Gbps signal into a 2-lane 112 Gbps signal. If the serdes rate of the GPU is 224 Gbps, the re-timer can combine the GPU's 1-lane 224 Gbps signal into a 2-lane 112 Gbps signal, thus achieving a match in transmission rate.
[0099] Optionally, in the control method of the computer system provided in the embodiments of the present application, adjusting the first data transmission parameter through the re-timer in the computer system to adapt to the second data transmission parameter includes: determining the working mode of the re-timer according to the first data transmission parameter and the second data transmission parameter; in the working mode, adjusting the first data transmission parameter through the re-timer to adapt to the second data transmission parameter.
[0100] In an optional embodiment, adjusting the first data transmission parameter through the re-timer in the computer system includes the following steps: First, it is necessary to determine what working mode the re-timer should be in based on the first data transmission parameter of the image processor and the second data transmission parameter of the switching node. The first data transmission parameter may include the transmission rate (such as 56 Gbps or 112 Gbps), the number of lanes (the number of transmission channels), etc., while the second data transmission parameter reflects the data transmission characteristics supported by the switching node.
[0101] For example, if the transmission rate of the image processor is 56 Gbps and the switching node supports 112 Gbps, the re-timer may need to operate in Gearbox mode to combine lanes and boost the lower transmission rate to the rate required by the switching node. Gearbox mode is a technique used for signal rate adaptation in high-speed serial communication links.
[0102] After determining the working mode of the re-timer, in the determined working mode, the first data transmission parameter is adjusted through the re-timer to adapt to the second data transmission parameter. For example, multiple lower-speed lanes are combined into a higher-speed lane, thereby converting the low-speed serdes signal into a high-speed signal compatible with the switching node.
[0103] Through the retimer, the computer system can be compatible with different types and specifications of image processors, even if there are significant differences in their data transmission parameters, avoiding the need to design dedicated interconnection solutions for different image processors, reducing the hardware cost and design complexity, and improving the economic benefits.
[0104] Optionally, in the control method of the computer system provided in the embodiments of the present application, determining the working mode of the retimer according to the first data transmission parameter and the second data transmission parameter includes: judging whether the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter to obtain a judgment result; and determining the working mode of the retimer according to the judgment result.
[0105] In an optional embodiment, determining the working mode of the retimer according to the first data transmission parameter and the second data transmission parameter includes the following steps: judging the first data transmission parameter (mainly including the transmission rate) of the image processor and the second data transmission parameter of the switching node. If it is found that the transmission rates between the two are inconsistent, it means that there is a rate mismatch between the image processor and the switching node. For example, if the serdes rate of the image processor is 56 Gbps, while the expected serdes rate of the switching node is 112 Gbps, this indicates a rate difference and adjustment is required.
[0106] Based on the above judgment result, determine which working mode the retimer should enter for necessary rate adaptation.
[0107] In an optional embodiment, if the judgment result indicates that the transmission rate in the first data transmission parameter is not the same as the transmission rate in the second data transmission parameter, the working mode is determined to be the first working mode, where the first working mode is to adjust the transmission rate; if the judgment result indicates that the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter, the working mode is determined to be the second working mode, where the second working mode is to remap the transmission channels of the physical ports.
[0108] For example, if the judgment result is that the transmission rates are the same, it indicates that the image processor and the switching node can communicate directly. The working mode of the retimer can be the Retimer mode (i.e., the second working mode mentioned above), and the output of the image processor is remapped to the switching node. The Retimer mode is a technology used for signal regeneration and shaping in high-speed digital signal transmission and is widely applied in high-speed interconnection standards such as PCIe (Peripheral Component Interconnect Express), Ethernet, and Fibre Channel. In the Retimer mode, the retimer is mainly responsible for retiming the signal and reshaping the data to ensure that the signal still maintains good signal quality and integrity after long-distance transmission or passing through multiple interconnection points.
[0109] For example, if the judgment result is that the transmission rates are different, the working mode of the retimer can be the Gearbox mode (i.e., the first working mode mentioned above). If the transmission rate of the image processor is low, the Gearbox mode will merge lanes to increase the transmission rate; conversely, if the transmission rate of the image processor is high, the Gearbox mode will split lanes to reduce the transmission rate. In this way, the retimer will adjust the signal rate of the image processor to match the rate of the switching node to ensure the normal transmission of the signal.
[0110] By judging whether the transmission rates are the same and selecting the corresponding working mode (the first working mode or the second working mode), the high efficiency and compatibility of signal transmission between the image processor and the switching node are achieved, and thus a high-performance, highly compatible, and easily expandable AI computer system is realized.
[0111] Optionally, in the control method of the computer system provided in the embodiment of the present application, in the working mode, the first data transmission parameter is adjusted by the retimer in the computer system to adapt to the second data transmission parameter, including: if the working mode is the first working mode, the transmission rate in the first data transmission parameter is adjusted by adjusting the number of transmission channels in the first data transmission parameter to adapt to the second data transmission parameter.
[0112] In an optional embodiment, adjusting the first data transmission parameter by a retimer in a computer system includes the following steps: When the operating mode of the retimer is the first operating mode, the retimer changes the data transmission rate by adjusting the number of transmission channels (lanes), so as to match the rate of the switching node. For example, determine the difference in the transmission rates between the image processor and the switching node. For example, if the transmission rate of the image processor is 56 Gbps / lane and the transmission rate of the switching node is 112 Gbps / lane, then adjust the number of transmission channels of the retimer according to the difference, so that the adjusted data transmission rate is the data transmission rate corresponding to the second data transmission parameter.
[0113] By adjusting the number of lanes to achieve rate matching, the computer system can be compatible with image processors with different transmission rates, improving the flexibility and versatility of the system.
[0114] Optionally, in the control method of the computer system provided in the embodiments of the present application, adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter includes: If the transmission rate in the first data transmission parameter is less than the transmission rate in the second data transmission parameter, then perform a merging process on the transmission channels of the physical port to achieve an adjustment of the number of transmission channels in the first data transmission parameter.
[0115] In an optional embodiment, adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter includes the following steps: Confirm that there is a difference between the transmission rate in the first data transmission parameter and the transmission rate in the second data transmission parameter, and the transmission rate of the first data transmission parameter is less than the transmission rate in the second data transmission parameter. For example, the serdes transmission rate of the image processor may be 56 Gbps / lane, while the transmission rate of the switching node is 112 Gbps / lane. Enable the Gearbox function of the retimer and enter the processing mode of merging transmission channels. In the first operating mode, the retimer combines the data streams of multiple lanes received and transmits them as fewer high-speed lanes. For example, if the image processor has 8 lanes with a transmission rate of 56 Gbps per lane, then the control system may combine 4 of the lanes into 2 lanes, with the transmission rate of each lane increased to 112 Gbps to adapt to the high-speed interface of the switching node.
[0116] By merging lanes to increase the transmission rate of the image processor, the computer system can be compatible with image processors with different rates, improving the flexibility and versatility of the computer system.
[0117] Optionally, in the control method of the computer system provided in the embodiments of the present application, adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter includes: if the transmission rate in the first data transmission parameter is greater than the transmission rate in the second data transmission parameter, splitting the transmission channels of the physical port to adjust the number of transmission channels in the first data transmission parameter.
[0118] In an alternative embodiment, adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter further includes: detecting that the transmission rate in the first data transmission parameter (the transmission parameter of the image processor) is higher than the transmission rate in the second data transmission parameter (the transmission parameter of the switching node). For example, the serdes transmission rate of the image processor may be 112 Gbps / lane, while the transmission rate of the switching node is 56 Gbps / lane. To reduce the high-speed transmission rate of the image processor to a level matching that of the switching node, the retimer will split the data stream of the high-speed lane into a larger number of low-speed lanes for transmission. The splitting process involves increasing the number of transmission channels while reducing the transmission rate of each channel. For example, if the image processor has 4 lanes with a transmission rate of 112 Gbps per lane, the control system may split these 4 lanes into 8 lanes and reduce the transmission rate of each lane to 56 Gbps to adapt to the low-speed interface of the switching node.
[0119] By splitting the transmission channels of the physical port, the control method provided in the embodiments of the present application effectively solves the high-rate transmission problem between the image processor and the switching node, not only ensuring the compatibility and efficiency of data transmission, but also optimizing the hardware cost and space utilization.
[0120] Optionally, in the control method of the computer system provided in the embodiments of the present application, in the working mode, adjusting the first data transmission parameter by the retimer in the computer system to adapt to the second data transmission parameter includes: if the working mode is the second working mode, remapping the transmission channels corresponding to the first data transmission parameter to adapt to the second data transmission parameter.
[0121] In an optional embodiment, in the second working mode, adjusting the first data transmission parameter by a retimer in a computer system includes: in the remapping mode, the function of the retimer is changed to remap the transmission channels, that is, reallocate the lanes to ensure that the data stream of the image processor can match the lane configuration of the switching node. For example, the retimer can remap lane0 of the image processor to lane4, lane1 to lane5, lane2 to lane6, and lane3 to lane7.
[0122] Through lane remapping, the retimer can achieve transparent data transmission, that is, without changing the data payload and protocol information, only adjusting the lane configuration, ensuring compatibility between different hardware configurations and the continuity of data transmission.
[0123] Optionally, in the control method of the computer system provided in the embodiments of the present application, after adjusting the first data transmission parameter by a retimer in the computer system to adapt to the second data transmission parameter, the method further includes: obtaining the adjusted first data transmission parameter; determining the target number of high-density connectors to be turned on according to the adjusted first data transmission parameter, where the high-density connectors are used to implement data interaction between multiple image processors and the switching node; performing a turn-on process on the high-density connectors corresponding to the target number to implement data interaction between multiple image processors and the switching node.
[0124] In an optional embodiment, after adjusting the first data transmission parameter by a retimer in a computer system, the following steps are further included: The process of adjusting the first data transmission parameter (transmission parameter of an image processor, such as a GPU) by the retimer includes rate adjustment or lane remapping processing. After the adjustment is completed, it is necessary to determine the adjusted first data transmission parameter, and the adjusted first data transmission parameter may include detailed information such as lane configuration.
[0125] After obtaining the adjusted first data transmission parameter, determine the target number of high-density connectors required to implement data interaction between the image processor and the switching node according to the adjusted first data transmission parameter. For example, determining the target number of high-density connectors to be turned on according to the adjusted first data transmission parameter includes: determining the adjusted number of transmission channels according to the adjusted first data transmission parameter; determining the target number of high-density connectors to be turned on according to the adjusted number of transmission channels.
[0126] For example, if the adjusted first data transmission parameter indicates that the image processor needs to use more lanes to transmit data, then it may be necessary to increase the number of high-density connectors to carry these lanes.
[0127] For example, for the products of Company A, the serdes rate of the products of Company A is 56 Gbps. One OAM GPU has a total of 8 physical ports (where serdes1 is divided into serdes1L and serdes 1H), and each physical port is 8-lane. The single-lane rate of the switching node used by the switching node is 112 Gbps. The re-timer combines the 2-lane signals of the products of Company A into 1-lane signals. The serdes rate of the products of Company B is 112 Gbps. One OAM GPU is also divided into 8 physical ports, but each physical port has only 6 lanes. After being processed by the re-timer, there will be 8 more lane signals on the output of the re-timer for Company B than for Company A. Therefore, the target number of high-density connectors to be turned on when using the products of Company A is 2, and the target number of high-density connectors to be turned on when using the products of Company B is 3. By dynamically using high-density connectors, the cost of high-density connectors can be saved when matching with the GPU of Company A, reducing unnecessary hardware investment. Especially when not using all lanes, resource waste is avoided.
[0128] After determining the target number of high-density connectors, the high-density connectors with the target number will be configured and initialized to ensure that they can be correctly turned on and carry the data flow between the image processor and the switching node. This process may involve setting the physical connection, electrical characteristics, and signal path of the connectors to adapt to the adjusted data transmission parameters. Through the turn-on process, the high-density connectors can now adapt to the adjusted first data transmission parameters and realize the data interaction between the image processor and the switching node. This ensures that the data can be transmitted according to the adjusted rate and lane configuration, improving the efficiency and stability of data transmission.
[0129] In this application, under different working modes of the retimer, the data transfer rate and lane configuration between the image processor (such as GPU) and the switching node are automatically adjusted to achieve the adaptation of the rate and the number of channels. In the first working mode, it is possible to convert high-speed lane signals into lower-speed but more numerous lane signals, or vice versa, that is, convert from lower-speed lane signals to high-speed lane signals, to match the transmission rate requirements of different image processors and switching nodes. In the second working mode, through remapping processing, the correct orientation of the data stream and the accurate reception of signals are ensured. And according to the adjusted data transfer parameters of the image processor, the target number of high-density connectors to be turned on can be intelligently determined to achieve the optimal hardware configuration for data interaction between the image processor and the switching node. By dynamically adjusting the number of connectors, it is possible to adapt to the lane configuration differences between the image processor and the switching node without sacrificing data transfer efficiency and stability, while optimizing the use of hardware resources and reducing costs.
[0130] In an optional embodiment, taking three OAM GPU products included in a computer system as an example, which are the GPU of Company A (rate 56 Gbps), the GPU of Company B (rate 112 Gbps), and the GPU of Company C (rate 224 Gbps) respectively. Based on the transmission parameters of the GPU, the differences between each GPU are analyzed. For example, the number of lanes of the GPU of Company A may be 8, that of Company B is 6, and that of Company C is 3 (assuming for simplified illustration). At the same time, it is confirmed that the highest transmission rate supported by the switching node is 112 Gbps.
[0131] Compatible design for the GPU of Company A: For the GPU of Company A, the retimer adopts the Gearbox mode to convert the 56 Gbps signal of 8 lanes into the 112 Gbps signal of 4 lanes to match the transmission rate of the switching node.
[0132] Compatible design for the GPU of Company B: The transmission rate of the GPU of Company B already meets the requirements of the switching node (112 Gbps). Therefore, the retimer works in the Retimer mode in this scenario to remap the transmission channels of the GPU of Company B.
[0133] Compatible design for the GPU of Company C: The GPU of Company C has a higher transmission rate (224 Gbps). To adapt to the switching node, the retimer needs to turn on the Gearbox mode to convert the 224 Gbps signal of 3 lanes into the 112 Gbps signal of 6 lanes to achieve rate matching with the switching node.
[0134] Calculate the target number of high-density connectors required based on the adjusted number of transmission channels (lane count). For example, if Company A's GPU requires 4 lanes, Company B's GPU requires 6 lanes, and Company C's GPU requires 6 lanes, then the system may need multiple high-density connectors to carry the signals of these lanes to ensure full interconnection between each GPU and the switching node. It is also possible to intelligently plan the location and layout of high-density connectors by analyzing the distribution and requirements of lane signals to reduce the length of signal cables and improve signal transmission efficiency. For example, the connectors for GPUs processing at the same rate can be centrally arranged, or the spacing and direction between connectors can be adjusted according to the lane configuration of the GPUs.
[0135] Through the dynamic configuration of retimers and the intelligent management of high-density connectors, the compatibility issues of various OAM GPU products in computer systems are solved, and the data transmission performance is optimized. Through this compatibility process, the computer system can support GPUs with different rates and lane counts, achieving efficient and stable data transmission, demonstrating strong flexibility and adaptability, and meeting the market's support for diverse GPU products and their high-performance computing requirements.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0137] The embodiments of the present application also provide a control device for a computer system. The control device for the computer system is used to execute the control method for the computer system in any one of the above, as Figure 12 shown. The control device for the computer system includes: an acquisition unit 1201 and an adjustment unit 1202.
[0138] The acquisition unit 1201 is used to acquire the first data transmission parameters of the physical ports of multiple image processors in the computer system and the second data transmission parameters corresponding to the switching node of the computer system;
[0139] The adjustment unit 1202 is used to adjust the first data transmission parameters through the retimer in the computer system to adapt to the second data transmission parameters.
[0140] Optionally, in the control device for the computer system provided in the embodiments of the present application, the adjustment unit includes: a first determination subunit, which is used to determine the working mode of the retimer according to the first data transmission parameters and the second data transmission parameters; an adjustment subunit, which is used to adjust the first data transmission parameters through the retimer in the working mode to adapt to the second data transmission parameters.
[0141] Optionally, in the control device of the computer system provided in the embodiments of the present application, the first determination subunit includes: a judgment module, configured to judge whether the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter, and obtain a judgment result; a determination module, configured to determine the working mode of the retimer according to the judgment result.
[0142] Optionally, in the control device of the computer system provided in the embodiments of the present application, the determination module includes: a first determination sub-module, configured to determine that the working mode is the first working mode if the judgment result indicates that the transmission rate in the first data transmission parameter is different from the transmission rate in the second data transmission parameter, where the first working mode is to adjust the transmission rate; a second determination sub-module, configured to determine that the working mode is the second working mode if the judgment result indicates that the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter, where the second working mode is to remap the transmission channels of the physical port.
[0143] Optionally, in the control device of the computer system provided in the embodiments of the present application, the adjustment subunit includes: an adjustment module, configured to, if the working mode is the first working mode, adjust the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter to adapt to the second data transmission parameter.
[0144] Optionally, in the control device of the computer system provided in the embodiments of the present application, the adjustment module includes: a first processing sub-module, configured to, if the transmission rate in the first data transmission parameter is less than the transmission rate in the second data transmission parameter, merge the transmission channels of the physical port to adjust the number of transmission channels in the first data transmission parameter.
[0145] Optionally, in the control device of the computer system provided in the embodiments of the present application, the adjustment module includes: a second processing sub-module, configured to, if the transmission rate in the first data transmission parameter is greater than the transmission rate in the second data transmission parameter, split the transmission channels of the physical port to adjust the number of transmission channels in the first data transmission parameter.
[0146] Optionally, in the control device of the computer system provided in the embodiments of the present application, the adjustment subunit includes: a processing module, configured to, if the working mode is the second working mode, remap the transmission channels corresponding to the first data transmission parameter to adapt to the second data transmission parameter.
[0147] Optionally, in the control device of the computer system provided in the embodiments of the present application, the device further includes: a second acquisition unit, configured to acquire the adjusted first data transmission parameter after adjusting the first data transmission parameter through a retimer in the computer system to adapt to the second data transmission parameter; a determination unit, configured to determine the target number of high-density connectors to be turned on according to the adjusted first data transmission parameter, where the high-density connectors are used to implement data interaction between multiple image processors and a switching node; and a processing unit, configured to perform a turn-on process on the high-density connectors corresponding to the target number to implement data interaction between multiple image processors and a switching node.
[0148] Optionally, in the control device of the computer system provided in the embodiments of the present application, the determination unit includes: a second determination subunit, configured to determine the adjusted number of transmission channels according to the adjusted first data transmission parameter; and a third determination subunit, configured to determine the target number of high-density connectors to be turned on according to the adjusted number of transmission channels.
[0149] For the descriptions of the features in the embodiments corresponding to the control device of the computer system, reference may be made to the relevant descriptions in the embodiments corresponding to the control method of the computer system, which will not be elaborated here one by one.
[0150] The embodiments of the present application further provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the embodiments of the control method of the computer system described above.
[0151] The embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the embodiments of the control method of the computer system described above when running.
[0152] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0153] The embodiments of the present application further provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the embodiments of the control method of the computer system described above.
[0154] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-described embodiments of the control method of the computer system are implemented.
[0155] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0156] The above has introduced in detail a computer system and a control method of the computer system provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A computer system, characterized in that, Comprising: Computing nodes, wherein the computing nodes can deploy at least two types of multiple image processors; multiple retimers, wherein the multiple retimers are connected to the multiple image processors, the multiple retimers include a first group of retimers and a second group of retimers, and the total number of data input channels of the first group of retimers is equal to the total number of first data output channels of the multiple image processors, and the total number of data input channels of the second group of retimers is greater than or equal to the total number of second data output channels of the multiple image processors; Switching nodes, wherein the switching nodes are connected to the output ends of the multiple retimers, and the interconnection between the computing nodes and other computing nodes is realized through the switching nodes.
2. The computer system according to claim 1, characterized in that, The first group of retimers is connected to multiple first data output channels of physical ports in the multiple image processors, and the second group of retimers is connected to multiple second data output channels of physical ports in the multiple image processors, wherein, based on the number of data output channels corresponding to one logical port of the image processor, the data output channels of the physical ports in one image processor are divided into multiple first data output channels and multiple second data output channels.
3. The computer system according to claim 2, characterized in that, Regarding the connection between the first group of retimers and multiple first data output channels of physical ports in the multiple image processors: When the multiple image processors are the first type of image processors, all data input channels at the input end of the first group of retimers are connected to the multiple first data output channels, and the output signal of the first group of retimers is output from the target data output channel at the output end of the first group of retimers, wherein the output data channels of the first type of image processors are merged through the first group of retimers to adapt to the switching nodes; When the multiple image processors are the second type of image processors, all data input channels at the input end of the first group of retimers are connected to the multiple first data output channels, and the output signal of the first group of retimers is output from all data output channels at the output end of the first group of retimers.
4. The computer system according to claim 2, characterized in that, Regarding the connection between the second group of retimers and multiple second data output channels of physical ports in the multiple image processors: When the multiple image processors are the first type of image processors, all data input channels at the input end of the second group of retimers are connected to the multiple second data output channels, and the output signal of the first group of retimers is output from the target data output channel at the output end of the second group of retimers, wherein the data output channels of the first type of image processors are merged through the second group of retimers to adapt to the switching nodes; When the multiple image processors are of the second type, the target data input channel in the input ends of the second group of retimers is connected to the multiple second data output channels, and the output signal of the second group of retimers is output by the target data output channel in the output ends of the second group of retimers.
5. The computer system according to claim 4, wherein the computer system further includes: a plurality of high-density connectors, wherein the plurality of high-density connectors are connected between the retimers and the switching node, and the number of the plurality of high-density connectors has a corresponding relationship with the data output channels having data output in the output ends of the plurality of retimers.
6. The computer system according to claim 5, wherein for the connection relationship between the plurality of high-density connectors and the plurality of retimers: all the data output channels in the output ends of the first group of retimers are connected to the first high-density connector and the second high-density connector among the plurality of high-density connectors; the target data output channel in the output ends of the second group of retimers is connected to the third high-density connector among the plurality of high-density connectors.
7. The computer system according to claim 6, wherein for the connection relationship between all the data output ports in the output ends of the first group of retimers and the first high-density connector and the second high-density connector among the plurality of high-density connectors: the target data output channel in the output ends of the first group of retimers is connected to the first high-density connector; the data output channels other than the target data output channel in the output ends of the first group of retimers are connected to the second high-density connector.
8. The computer system according to claim 6, wherein when the multiple image processors are of the first type, the first high-density connector and the third high-density connector are in a conducting state, and the second high-density connector is in a non-conducting state; when the multiple image processors are of the second type, the first high-density connector, the second high-density connector and the third high-density connector are all in a conducting state.
9. A control method for a computer system, wherein the control method of the computer system is applied to the computer system according to any one of claims 1 to 8, and includes: acquiring first data transmission parameters of physical ports of multiple image processors in the computer system and second data transmission parameters corresponding to a switching node of the computer system; adjusting the first data transmission parameters through a retimer in the computer system to adapt to the second data transmission parameters.
10. The control method according to claim 9, wherein adjusting the first data transmission parameters through a retimer in the computer system to adapt to the second data transmission parameters includes: determining an operating mode of the retimer according to the first data transmission parameters and the second data transmission parameters; In the working mode, the first data transmission parameter is adjusted by the retimer to adapt to the second data transmission parameter.
11. The control method according to claim 10, wherein Determining the working mode of the retimer based on the first data transmission parameter and the second data transmission parameter includes: Judging whether the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter to obtain a judgment result; Determining the working mode of the retimer according to the judgment result.
12. The control method according to claim 11, wherein Determining the working mode of the retimer according to the judgment result includes: If the judgment result indicates that the transmission rate in the first data transmission parameter is different from the transmission rate in the second data transmission parameter, it is determined that the working mode is the first working mode, wherein the first working mode is to adjust the transmission rate; If the judgment result indicates that the transmission rate in the first data transmission parameter is the same as the transmission rate in the second data transmission parameter, it is determined that the working mode is the second working mode, wherein the second working mode is to remap the transmission channels of the physical port.
13. The control method according to claim 10, wherein In the working mode, adjusting the first data transmission parameter by the retimer in the computer system to adapt to the second data transmission parameter includes: If the working mode is the first working mode, the transmission rate in the first data transmission parameter is adjusted by adjusting the number of transmission channels in the first data transmission parameter to adapt to the second data transmission parameter.
14. The control method according to claim 13, wherein Adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter includes: If the transmission rate in the first data transmission parameter is less than the transmission rate in the second data transmission parameter, the transmission channels of the physical port are merged to adjust the number of transmission channels in the first data transmission parameter.
15. The control method according to claim 13, wherein Adjusting the transmission rate in the first data transmission parameter by adjusting the number of transmission channels in the first data transmission parameter includes: If the transmission rate in the first data transmission parameter is greater than the transmission rate in the second data transmission parameter, the transmission channels of the physical port are split to adjust the number of transmission channels in the first data transmission parameter.
16. The control method according to claim 10, wherein In the working mode, adjusting the first data transmission parameter by the retimer in the computer system to adapt to the second data transmission parameter includes: If the working mode is the second working mode, remap the transmission channel corresponding to the first data transmission parameter to adapt to the second data transmission parameter.
17. The control method according to claim 9, characterized in that after adjusting the first data transmission parameter to adapt to the second data transmission parameter through the retimer in the computer system, the method further includes: obtaining the adjusted first data transmission parameter; determining the target number of high-density connectors to be turned on according to the adjusted first data transmission parameter, wherein the high-density connectors are used to implement data interaction between the multiple image processors and the switching node; performing a turn-on process on the high-density connectors corresponding to the target number to implement data interaction between the multiple image processors and the switching node.
18. The control method according to claim 17, characterized in that determining the target number of high-density connectors to be turned on according to the adjusted first data transmission parameter includes: determining the adjusted number of transmission channels according to the adjusted first data transmission parameter; determining the target number of high-density connectors to be turned on according to the adjusted number of transmission channels.
19. A control device for a computer system, characterized in that it includes: an acquisition unit for acquiring the first data transmission parameter of the physical ports of multiple image processors of the computer system and the second data transmission parameter corresponding to the switching node of the computer system; an adjustment unit for adjusting the first data transmission parameter through the retimer in the computer system to adapt to the second data transmission parameter.
20. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method described in any one of claims 9 to 18 are implemented.
21. A computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method described in any one of claims 9 to 18 are implemented.
Citation Information
Cited By
Bus communication system, signal aggregation method and server
CN121051049A