Power consumption optimization method for switching chip, and related device thereof
By sampling and detecting the traffic forwarded by the switch chip and adjusting its power consumption to match the changes in the working stage of the AI chip, the problem of excessive power consumption of the switch chip is solved, and the optimization of power consumption and efficiency improvement is achieved.
Patent Information
- Application Number
- PCT/CN2024/115353
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2024-08-29
- Publication Date
- 2025-09-04
AI Technical Summary
The switching chip is always in a full-load working state in the AI training cluster, resulting in excessive power consumption and cannot match the actual workload demand.
By sampling the traffic forwarded by the switch chip, it detects whether the AI chip switches between the computing stage and the communication stage, and adjusts power consumption according to the stage changes, including reducing or increasing the operating frequency, voltage and the status of the serializer/deserializer to match the actual workload.
The power consumption of the switch chip is optimized, avoiding the inflated power consumption, and improving the matching and efficiency of power consumption.
Smart Images

Figure CN2024115353_04092025_PF_FP_ABST
Abstract
Description
A power consumption optimization method for switching chips and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on February 28, 2024, with application number 202410223654.4 and application name “A method for power consumption optimization of switching chips and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of chip technology, and in particular to a power consumption optimization method for a switching chip and related equipment. Background Art
[0003] With the popularization of artificial intelligence (AI) training clusters, in order to improve the training efficiency of AI training clusters for neural network models, an ultra-high bandwidth interconnection infrastructure device can be introduced into the AI training cluster. This device is used as an interconnection node to connect multiple computing nodes in the cluster, thereby reducing the communication delay between computing nodes and improving the training efficiency of the AI training cluster.
[0004] The AI training cluster provided by the related technology may include multiple computing nodes and interconnecting nodes. The multiple computing nodes may be presented as multiple AI chips, and the interconnecting nodes may be presented as switching chips, wherein the receiving end of the switching chip is connected to a portion of the AI chips, and the transmitting end of the switching chip is connected to another portion of the AI chips. The model training performed by multiple AI chips includes multiple rounds of iterations, each of which consists of a computing phase and a communication phase. During the computing phase, the communication volume between AI chips is close to zero, while during the communication phase, the communication volume between AI chips is very large. Whether in the computing phase or the communication phase, the switching chip can ensure normal communication between AI chips.
[0005] However, in order to ensure the communication quality between AI chips, the switching chip often needs to be in a full-load working state all the time, resulting in excessive power consumption of the switching chip.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a power consumption optimization method for a switching chip and related equipment, which can make the power consumption of the switching chip match the actual workload, thereby avoiding the situation where the power consumption of the switching chip is inflated and achieving power consumption optimization.
[0008] The first aspect of the embodiments of the present application provides a power consumption optimization method for a switching chip, which can be used to forward traffic between multiple AI chips, that is, forwarding the traffic of one part of the AI chip to another part of the AI chip for processing. Multiple AI chips can be used to train neural networks. During the training process of the neural network, multiple AI chips can be in a computing stage or a communication stage (that is, the working state of multiple AI chips can be a computing state or a communication state). When in the computing stage, the traffic between multiple AI chips is close to zero, and when in the communication stage, the traffic between multiple AI chips is very large. The method includes:
[0009] During the operation of the switch chip, the switch chip can sample the traffic forwarded by the switch chip to obtain the sampled traffic. The switch chip can then use the sampled traffic to detect whether multiple AI chips are switching between the computing phase and the communication phase, and determine whether to adjust the chip power consumption for the traffic to be forwarded.
[0010] If it is determined based on the sampled traffic that multiple AI chips have not switched between the computing and communication phases, the switching chip will not adjust the power consumption of the traffic to be forwarded, and will subsequently forward the traffic to be forwarded according to the current power consumption. If it is determined based on the sampled traffic that multiple AI chips have just switched from the communication phase to the computing phase, the switching chip will reduce the power consumption of the traffic to be forwarded, and will subsequently forward the traffic to be forwarded according to the reduced power consumption. If it is determined based on the sampled traffic that multiple AI chips have just switched from the computing phase to the communication phase, the switching chip will increase the power consumption of the traffic to be forwarded, and will subsequently forward the traffic to be forwarded according to the increased power consumption.
[0011] The above method demonstrates that the switching chip has the ability to adjust power consumption. It can identify changes in the AI chip's phase based on the forwarded traffic and adaptively adjust the power consumption of the transmitted traffic based on these changes. This means that when the AI chip is in the communication phase, the switching chip can operate at high power consumption to forward traffic between AI chips. When the AI chip is in the computation phase, the switching chip can operate at low power consumption to forward traffic between AI chips. This allows the switching chip's power consumption to match the actual workload, preventing inflated power consumption and achieving optimized power consumption.
[0012] In one possible implementation, the training process includes multiple sampling cycles, the multiple sampling cycles include the current sampling cycle, the sampled traffic is the traffic sampled in the current sampling cycle, and the switching chip detects whether the multiple AI chips are switching between the computing phase and the communication phase based on the sampled traffic, including: the switching chip detects whether the multiple AI chips are switching between the computing phase and the communication phase based on the traffic sampled in the previous sampling cycle and the traffic sampled in the current sampling cycle. In the aforementioned implementation, the entire training process can be divided into multiple statistical cycles, each of which can include multiple sampling cycles. Assuming that the current sampling cycle in the current statistical cycle has been reached, the switching chip can sample the traffic being transmitted by the switching chip in the current sampling cycle to obtain the traffic sampled in the current sampling cycle. Since the traffic sampled in the previous sampling cycle is known, the switching chip can compare the traffic sampled in the previous sampling cycle with the traffic sampled in the current sampling cycle to accurately determine whether the multiple AI chips are switching between the computing phase and the communication phase.
[0013] In one possible implementation, if multiple AI chips switch from the communication phase to the computation phase, the switching chip reduces power consumption for the traffic to be forwarded, including: if the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period, the switching chip determines that the multiple AI chips have switched from the communication phase to the computation phase, and reduces power consumption for the traffic to be forwarded. In the aforementioned process, if the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period to a certain extent, the switching chip determines that the multiple AI chips have just switched from the communication phase to the computation phase, and the switching chip reduces power consumption for the traffic to be forwarded, and subsequently forwards the traffic to be forwarded using the reduced power consumption.
[0014] In one possible implementation, if multiple AI chips switch from the computation phase to the communication phase, the switching chip increases power consumption for the traffic to be forwarded, including: if the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period, the switching chip determines that the multiple AI chips have switched from the computation phase to the communication phase, and increases power consumption for the traffic to be forwarded. In the aforementioned process, if the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period to a certain extent, the switching chip determines that the multiple AI chips have just switched from the computation phase to the communication phase, and the switching chip increases power consumption for the traffic to be forwarded, and subsequently forwards the traffic to be forwarded using the increased power consumption.
[0015] In one possible implementation, the switching chip reduces power consumption for traffic to be forwarded, including: the switching chip performs at least one of the following: gradually reducing the operating frequency of the switching chip, gradually reducing the operating voltage of the switching chip, and shutting down the serializer and deserializer (serdes) of the switching chip, wherein the priority of gradually reducing the operating frequency of the switching chip is higher than the priority of gradually reducing the operating voltage of the switching chip, and the priority of gradually reducing the operating voltage of the switching chip is higher than the priority of shutting down the serdes of the switching chip. In the aforementioned implementation, when the switching chip needs to reduce power consumption for subsequent traffic, the switching chip can achieve this by reducing the operating frequency, reducing the operating voltage, and shutting down the serdes. Since the switching chip adopts a solution of gradually reducing the operating frequency, gradually reducing the operating voltage, and shutting down some serdes in this process, when multiple AI chips switch from the communication phase to the computing phase, the switching chip can achieve smooth adjustment of power consumption, so that the traffic between the AI chips it forwards shows a continuous reduction and smooth transition.
[0016] In one possible implementation, the switching chip increases the power consumption for the traffic to be forwarded, including: the switching chip performs at least one of the following at the moment when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, the switching chip performs at least one of the following at a target moment: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip, and the target moment is before the moment when the traffic to be forwarded arrives. In the aforementioned implementation, when the switching chip needs to increase the power consumption for subsequent traffic, the switching chip can achieve this by starting serdes, increasing the operating voltage, and increasing the operating frequency. Since in this process the switching chip can increase the power consumption instantly (increase the power consumption when the subsequent traffic arrives) or increase the power consumption in advance (increase the power consumption at the target moment) according to actual sampling, it can avoid the long reset (startup) time of the serdes in the switching chip affecting the communication between AI chips.
[0017] In one possible implementation, the training process includes multiple statistical cycles, the multiple statistical cycles include the previous statistical cycle and the current statistical cycle, the current statistical cycle includes the previous sampling cycle and the current sampling cycle, and the switching chip executes at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip. The method includes: if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle is zero, the switching chip executes at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle and the average duration of the multiple AI chips in the computing phase in the previous statistical cycle are zero, the switching chip executes at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; The average duration of the phase is not zero, and the standard deviation of the duration of multiple AI chips in the communication phase in the previous statistical cycle is greater than or equal to a preset first threshold, the switching chip executes at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, if the average duration of multiple AI chips in the communication phase in the previous statistical cycle and the average duration of multiple AI chips in the calculation phase in the previous statistical cycle are not zero, and the standard deviation of the duration of multiple AI chips in the calculation phase in the previous statistical cycle is greater than or equal to a preset second threshold, the switching chip executes at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip. In the above process, the multiple statistical cycles also include the previous statistical cycle. The switching chip can obtain traffic characteristic parameters used to generate the power consumption adjustment strategy for the current statistical period in the previous statistical period. For example, in the previous statistical period, the average duration of multiple AI chips in the communication phase, the average duration of multiple AI chips in the calculation phase, the standard deviation of the duration of multiple AI chips in the communication phase, and the standard deviation of the duration of multiple AI chips in the calculation phase. Then, in the current statistical phase, when the switching chip needs to increase power consumption, the switching chip can determine whether the current traffic waveform conforms to the typical AI traffic model based on these traffic characteristic parameters. If it is determined based on these traffic characteristic parameters that the current traffic waveform does not conform to the typical AI traffic model, the switching chip will choose to immediately increase power consumption to achieve the increase in power consumption.
[0018] In one possible implementation, the switching chip performs at least one of the following at the target moment: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip, including: if the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period is less than a preset second threshold, the switching chip performs at least one of the following at the target moment: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip. In the aforementioned implementation, if it is judged based on these traffic characteristic parameters that the current traffic waveform conforms to the typical AI traffic model, the switching chip chooses to increase power consumption in advance to achieve an increase in power consumption, thereby avoiding the long reset (startup) time of the serdes in the switching chip affecting the communication between the AI chips.
[0019] In one possible implementation, the average duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; the average duration of time that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period; the standard deviation of the duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; and the standard deviation of the duration that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period. In the above-mentioned implementation, the switching chip may sample the traffic being transmitted by the switching chip in multiple sampling periods of the previous statistical period, and determine, based on the traffic sampled in multiple sampling periods, the number of times multiple AI chips are in the communication stage, the number of times multiple AI chips are in the calculation stage, the duration of each time multiple AI chips are in the communication stage, and the duration of each time multiple AI chips are in the calculation stage in the previous statistical period, so as to calculate, based on these traffic characteristic parameters, the average duration of multiple AI chips in the communication stage, the average duration of multiple AI chips in the calculation stage, the standard deviation of the duration of multiple AI chips in the communication stage, and the standard deviation of the duration of multiple AI chips in the calculation stage in the previous statistical period.
[0020] In one possible implementation, the difference between the target time and the time when the traffic to be forwarded arrives is calculated based on the average time that multiple AI chips were in the calculation phase in the previous statistical period and the standard deviation of the time that multiple AI chips were in the calculation phase in the previous statistical period.
[0021] A second aspect of an embodiment of the present application provides a switching chip, which is used to forward traffic between multiple AI chips. The multiple AI chips are used to train a neural network. During the training process of the neural network, the multiple AI chips are in a computing phase or a communication phase. The switching chip includes: a sampling module, which is used to sample the traffic forwarded by the switching chip to obtain the sampled traffic; a detection module, which is used to detect whether the multiple AI chips are switching between the computing phase and the communication phase based on the sampled traffic; a first adjustment module, which is used to reduce the power consumption of the traffic to be forwarded if the multiple AI chips switch from the communication phase to the computing phase; or, a second adjustment module, which is used to increase the power consumption of the traffic to be forwarded if the multiple AI chips switch from the computing phase to the communication phase.
[0022] In one possible implementation, the training process includes multiple sampling cycles, the multiple sampling cycles include the current sampling cycle, the sampled traffic is the traffic sampled in the current sampling cycle, and the detection module is used to detect whether multiple AI chips switch between the computing stage and the communication stage based on the traffic sampled in the previous sampling cycle and the traffic sampled in the current sampling cycle.
[0023] In one possible implementation, the first adjustment module is used to determine that multiple AI chips switch from the communication stage to the computing stage and reduce power consumption for the traffic to be forwarded if the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period.
[0024] In one possible implementation, the second adjustment module is used to determine that multiple AI chips switch from the computing stage to the communication stage and increase the power consumption of the traffic to be forwarded if the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period.
[0025] In one possible implementation, the first adjustment module is used to perform at least one of the following: gradually reducing the operating frequency of the switching chip, gradually reducing the operating voltage of the switching chip, and shutting down the serializer and deserializer Serdes of the switching chip, wherein the priority of gradually reducing the operating frequency of the switching chip is higher than the priority of gradually reducing the operating voltage of the switching chip, and the priority of gradually reducing the operating voltage of the switching chip is higher than the priority of shutting down the Serdes of the switching chip.
[0026] In one possible implementation, the second adjustment module is used to: execute at least one of the following at the moment when the traffic to be forwarded arrives: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or, execute at least one of the following at a target moment: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip, and the target moment is before the moment when the traffic to be forwarded arrives.
[0027] In one possible implementation, the training process includes multiple statistical cycles, the multiple statistical cycles include the previous statistical cycle and the current statistical cycle, the current statistical cycle includes the previous sampling cycle and the current sampling cycle, and the second adjustment module is used to: if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle is zero, at the moment when the traffic to be forwarded arrives, perform at least one of the following: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or, if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle and the average duration of the multiple AI chips in the calculation phase in the previous statistical cycle are not zero, and the multiple AI chips in the previous statistical cycle If the standard deviation of the duration in the communication phase is greater than or equal to a preset first threshold, at least one of the following items is performed when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, if the average duration of multiple AI chips in the communication phase in the previous statistical period and the average duration of multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of multiple AI chips in the calculation phase in the previous statistical period is greater than or equal to a preset second threshold, at least one of the following items is performed when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip.
[0028] In one possible implementation, the second adjustment module is configured to, if the average duration of time that multiple AI chips were in the communication phase in a previous statistical period and the average duration that multiple AI chips were in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration that multiple AI chips were in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration that multiple AI chips were in the calculation phase in the previous statistical period is less than a preset second threshold, perform at least one of the following at a target time: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip.
[0029] In one possible implementation, the average duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; the average duration of time that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period; the standard deviation of the duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; and the standard deviation of the duration that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period.
[0030] In one possible implementation, the difference between the target time and the time when the traffic to be forwarded arrives is calculated based on the average time that multiple AI chips were in the calculation phase in the previous statistical period and the standard deviation of the time that multiple AI chips were in the calculation phase in the previous statistical period.
[0031] A third aspect of an embodiment of the present application provides an interconnection device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the device executes the method described in the first aspect or any possible implementation method of the first aspect.
[0032] A fourth aspect of an embodiment of the present application provides a computer storage medium storing one or more instructions, which, when executed by one or more computers, enables the one or more computers to implement the method described in the first aspect or any possible implementation method of the first aspect.
[0033] A fifth aspect of the embodiments of the present application provides a computer program product, which stores instructions. When the instructions are executed by a computer, the computer implements the method described in the first aspect or any possible implementation method of the first aspect.
[0034] In an embodiment of the present application, a switching chip can be used to transmit traffic between multiple AI chips, which can be used to train a neural network model. During the training process, the multiple AI chips can switch back and forth between the computation phase and the communication phase. During operation, the switching chip can sample the traffic being forwarded by the switching chip to obtain the sampled traffic. The switching chip can then use the sampled traffic to detect whether the multiple AI chips are switching between the computation phase and the communication phase, thereby determining whether to adjust the power consumption of the subsequent traffic to be forwarded. If the multiple AI chips switch from the communication phase to the computation phase, the switching chip can reduce the power consumption of the traffic to be forwarded, so that the traffic to be forwarded can be forwarded at the reduced power consumption. If the multiple AI chips switch from the computation phase to the communication phase, the switching chip can increase the power consumption of the traffic to be forwarded, so that the traffic to be forwarded can be forwarded at the increased power consumption. In the aforementioned process, the switching chip has power consumption adjustment capabilities, which can identify the changes in the AI chip phase based on the forwarded traffic and adaptively adjust the power consumption of the transmitted traffic based on these changes. It can be seen that when the AI chip is in the communication stage, the switching chip can be in a high-power state to forward traffic between AI chips. When the AI chip is in the computing stage, the switching chip can be in a low-power state to forward traffic between AI chips. In this way, the power consumption of the switching chip can match the actual workload, thereby avoiding the situation where the power consumption of the switching chip is inflated and achieving power consumption optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] FIG1 is a schematic diagram of the structure of a training cluster provided in an embodiment of the present application;
[0036] FIG2 is a schematic diagram of an AI traffic model provided in an embodiment of the present application;
[0037] FIG3 is a schematic diagram of the structure of a switching chip provided in an embodiment of the present application;
[0038] FIG4 is a flow chart of a method for optimizing power consumption of a switching chip according to an embodiment of the present application;
[0039] FIG5 is a schematic diagram of a statistical period and a sampling period provided in an embodiment of the present application;
[0040] FIG6 is another schematic flow chart of a method for optimizing power consumption of a switching chip according to an embodiment of the present application;
[0041] FIG7 is another schematic diagram of a statistical period and a sampling period provided in an embodiment of the present application;
[0042] FIG8 is a schematic structural diagram of a switching chip provided in an embodiment of the present application;
[0043] FIG9 is a schematic structural diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The embodiments of the present application provide a power consumption optimization method for a switching chip and related equipment, which can make the power consumption of the switching chip match the actual workload, thereby avoiding the situation where the power consumption of the switching chip is inflated and achieving power consumption optimization.
[0045] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0046] With the popularization of AI training clusters, in order to improve the training efficiency of AI training clusters for neural network models, an ultra-high bandwidth interconnection infrastructure device can be introduced into the AI training cluster. This device is used as an interconnection node to connect multiple computing nodes in the cluster, thereby reducing the communication delay between computing nodes and improving the training efficiency of the AI training cluster.
[0047] The AI training cluster provided by the related technology may include multiple computing nodes and interconnection nodes. The multiple computing nodes may be presented as multiple AI chips, and the interconnection nodes may be presented as switching chips, wherein the receiving end of the switching chip is connected to a part of the AI chips, and the transmitting end of the switching chip is connected to another part of the AI chips. The model training performed by multiple AI chips includes multiple rounds of iterations, and each round of iteration consists of a computing phase and a communication phase. In the computing phase, the amount of communication between AI chips is close to zero. In this case, the traffic between AI chips that the switching chip needs to forward is close to zero. In the communication phase, the amount of communication between AI chips is quite large. In this case, the traffic between AI chips that the switching chip needs to forward is very large. Since the switching chip has ultra-high bandwidth, it can smoothly forward large traffic from some AI chips to other AI chips, thereby ensuring normal communication between AI chips.
[0048] However, to ensure the communication quality between AI chips, the switching chip often needs to be in a full-load working state. During the communication stage, the switching chip needs to forward a lot of traffic, and the power consumption of the switching chip can match the actual traffic forwarding requirements. However, in the calculation stage, the switching chip needs to forward very little traffic, and the power consumption of the switching chip is much greater than the actual traffic forwarding requirements, which will lead to inflated power consumption of the switching chip.
[0049] In order to solve the above problems, an embodiment of the present application provides a power consumption optimization method for a switching chip, which can be implemented in combination with artificial intelligence (AI) technology. AI technology is a technical discipline that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by perceiving the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Using artificial intelligence for data processing is a common application of artificial intelligence.
[0050] In actual applications, AI technology can complete user tasks through neural network models to meet user needs. For example, an image from a user can be input into a neural network model to determine the category of the image. For another example, a voice from a user can be input into a neural network model to obtain corresponding text. For another example, a question from a user can be input into a neural network model to obtain a corresponding answer, and so on. The neural network models in these actual application scenarios are all trained neural network models. In order to complete the training of these models, it can be achieved through a training cluster. Figure 1 is a structural diagram of a training cluster provided in an embodiment of the present application. As shown in Figure 1, the training cluster includes a switching chip and multiple AI chips. The AI chip and the switching chip are introduced below respectively:
[0051] AI chips can be presented in many ways. For example, an AI chip can be a central processing unit (CPU), or a graphics processing unit (GPU), or a neural processing unit (NPU), or a tensor processing unit (TPU), etc. In a training cluster, multiple AI chips can be used to jointly implement AI training, that is, multiple AI chips can jointly train a neural network model. The training process of the neural network model can include multiple rounds of iterations, and each round of iteration includes a computing phase and a communication phase. When multiple AI chips are in the computing phase, the communication volume between the multiple AI chips is very small, and the traffic between the multiple AI chips is close to zero. When multiple AI chips are in the communication phase, the communication volume between the multiple AI chips is large, and the traffic between the multiple AI chips is much greater than zero. It can be seen that during the entire training process, multiple AI chips are either in the computing phase or in the communication phase, so multiple AI chips are constantly switching between the computing phase and the communication phase. In different computing stages, the traffic required to be transmitted between multiple AI chips is very close. Similarly, in different communication stages, the traffic required to be transmitted between multiple AI chips is also very close. Therefore, during the entire training process, the traffic required to be transmitted between multiple AI chips can be presented in a typical square wave shape. This square wave can also be called an AI traffic model, as shown in Figure 2 (Figure 2 is a schematic diagram of the AI traffic model provided in an embodiment of the present application).
[0052] A switch chip can connect multiple AI chips. Generally, the receiving end of the switch chip is connected to a portion of the multiple AI chips, and the transmitting end of the switch chip is connected to another portion of the multiple AI chips. Therefore, the switch chip can be used to forward traffic (e.g., the number of received packets, received bytes, transmitted packets, and transmitted bytes, etc.) between the multiple AI chips. In other words, the switch chip can forward traffic from the receiving AI chips to the transmitting AI chips. Because the amount of traffic required to be sent from the receiving AI chip to the transmitting AI chip is very similar and small during different computing phases, and very similar and large during different communication phases, if the switch chip is constantly operating at high load, its power consumption will be extremely high. To optimize the power consumption of the switch chip, improvements can be made to the switch chip so that it can adjust its power consumption based on the phase of the multiple chips.
[0053] As shown in Figure 3 (Figure 3 is a structural diagram of a switching chip provided by an embodiment of the present application), the switching chip provided by an embodiment of the present application includes: a sampling module, a detection module and an adjustment module. Among them, the sampling module can also be called a port monitoring circuit. The sampling module is connected to the port of the receiving end and the port of the sending end, and can sample the traffic forwarded by the switching chip (for example, the traffic forwarded by the switching chip in a certain time period) in real time to obtain the sampled traffic. The detection module can detect whether multiple AI chips are switching between the calculation phase and the communication phase based on the sampled traffic, and notify the adjustment module of the detection results. The adjustment module includes configuration registers that store various traffic characteristic parameters (these traffic characteristic parameters are often calculated based on the traffic obtained by previous sampling and stored in registers). These traffic characteristic parameters can be used as the current power consumption adjustment strategy. If the detection result indicates that multiple AI chips switch from the communication phase to the calculation phase, the adjustment module reduces the power consumption of the traffic to be transmitted based on the power consumption adjustment strategy. If the detection result indicates that multiple AI chips switch from the calculation phase to the communication phase, the adjustment module increases the power consumption of the traffic to be transmitted based on the power consumption adjustment strategy. This shows that the switching chip can adjust its power consumption in forwarding traffic based on the stage changes of multiple AI chips, thereby optimizing its own performance.
[0054] In order to further understand the power consumption adjustment process of the switching chip, the following will further introduce the process in conjunction with specific embodiments. It should be noted that the switching chip can divide the entire training process into multiple statistical cycles (the duration of each statistical cycle is fixed, and its size can be set according to actual needs, and there is no restriction here), and each statistical cycle contains multiple sampling cycles (the duration of each sampling cycle is fixed, and its size can be set according to actual needs, and there is no restriction here). Since the operation of the switching chip in each statistical cycle is similar, any one of the multiple statistical cycles will be selected for introduction below, and the statistical cycle will be regarded as the current statistical cycle. Since the current power consumption adjustment strategy applied by the switching chip in the current statistical cycle comes from the multiple parameters calculated by the switching chip in the previous statistical cycle, the operations performed by the switching chip in the previous statistical cycle will be introduced below. Figure 4 is a flow chart of the power consumption optimization method of the switching chip provided in an embodiment of the present application. As shown in Figure 4, the method includes:
[0055] 401. In the previous statistical period, the switching chip determines the number of times the multiple AI chips are in the computing phase, the number of times the multiple AI chips are in the communication phase, the duration of each time the multiple AI chips are in the computing phase, and the duration of each time the multiple AI chips are in the communication phase.
[0056] In this embodiment, after entering the previous statistical period, since the previous statistical period can be divided into multiple sampling periods, the switching chip can sample the traffic forwarded by the switching chip in each sampling period, thereby obtaining the traffic sampled by the switching chip in each sampling period. Then, based on the traffic sampled by the switching chip in multiple sampling periods of the previous statistical period, the switching chip can perform the following operations:
[0057] If the traffic sampled by the switching chip in several consecutive sampling periods is less than the preset traffic threshold (the size of this threshold can be set according to actual needs and is not restricted here), it can be considered that the switching chip forwarded a small amount of traffic (usually close to zero) between multiple AI chips in these consecutive sampling periods. The switching chip can determine that multiple AI chips are in the calculation phase at a time and use the number of these consecutive sampling periods and the duration of each sampling period to calculate the duration of the multiple AI chips in the calculation phase at a time.
[0058] If the traffic sampled by the switching chip in several consecutive sampling periods is greater than or equal to the preset traffic threshold, it can be considered that the switching chip forwarded large traffic between multiple AI chips in these consecutive sampling periods. The switching chip can determine that multiple AI chips are in the communication phase at a time and use the number of these consecutive sampling periods and the duration of each sampling period to calculate the duration of multiple AI chips in the communication phase at a time.
[0059] The above operations are repeated continuously until the end of the previous statistical period. The switching chip can determine the number of times the multiple AI chips are in the calculation phase, the number of times the multiple AI chips are in the communication phase, the duration of each time the multiple AI chips are in the calculation phase, and the duration of each time the multiple AI chips are in the communication phase in the previous statistical period.
[0060] For example, as shown in FIG5 (FIG5 is a schematic diagram of a statistical period and a sampling period provided in an embodiment of the present application), assuming that the statistical period is T sample-period , the sampling period is T sample-interval , a T sample-period Can be divided into T sample-interval . Let the previous statistical period be T sample-period-0 , in T sample-period-0 In each T sample-interval The forwarded traffic is sampled, so that the sample-interval , the traffic sampled in the ,. The switching chip can make the following judgments:
[0061] If several consecutive T sample-interval The traffic obtained by sampling is all large traffic. The switching chip can record it as multiple AI chips in the communication phase at a time and store these multiple T sample-interval The number of times T sample-interval The duration Ti of multiple AI chips in the communication phase is obtained.
[0062] If several consecutive T sample-interval The traffic obtained by sampling is close to 0. The switching chip can record it as multiple AI chips in the calculation stage at a time and store these T sample-interval The number of times T sample-interval The duration of time that multiple AI chips are in the calculation stage at a time is obtained.
[0063] Then, the switching chip can calculate the sample-period-0 In the example, the number of times the multiple AI chips are in the communication phase is n (n is usually an integer greater than or equal to 2), the number of times the multiple AI chips are in the computing phase is m (m is usually an integer greater than or equal to 2), and the duration of each communication phase of the multiple AI chips is T1, T2, ..., T n , and the duration t1, t2, ..., t of each calculation phase of multiple AI chips m .
[0064] 402. The switching chip calculates the number of times the multiple AI chips are in the communication phase in the previous statistical period and the duration of each time the multiple AI chips are in the communication phase in the previous statistical period, to obtain an average duration that the multiple AI chips are in the communication phase in the previous statistical period.
[0065] 403. The switching chip calculates the number of times the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period to obtain an average duration of the multiple AI chips in the calculation phase in the previous statistical period.
[0066] 404. The switching chip calculates the average duration of the multiple AI chips in the communication phase in the previous statistical period and the duration of each communication phase of the multiple AI chips in the previous statistical period to obtain a standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period.
[0067] 405. The switching chip calculates the average duration of the calculation phase of the multiple AI chips in the previous statistical period and the duration of each calculation phase of the multiple AI chips in the previous statistical period to obtain the standard deviation of the duration of the calculation phase of the multiple AI chips in the previous statistical period.
[0068] After obtaining the number of times the multiple AI chips were in the calculation phase in the previous statistical cycle, the number of times the multiple AI chips were in the communication phase in the previous statistical cycle, the duration of each time the multiple AI chips were in the calculation phase in the previous statistical cycle, and the duration of each time the multiple AI chips were in the communication phase in the previous statistical cycle, the switching chip can detect whether the number of times the multiple AI chips were in the calculation phase in the previous statistical cycle is close to the number of times the multiple AI chips were in the communication phase in the previous statistical cycle. If the two are close (for example, they are equal or the difference between the two is one), the switching chip can perform the following calculation to obtain corresponding traffic characteristic parameters:
[0069] In the previous statistical cycle, if the switching chip can average the number of times multiple AI chips are in the communication phase and the duration of each communication phase, thereby obtaining the average duration of the multiple AI chips in the communication phase, then the switching chip can also average the number of times multiple AI chips are in the calculation phase and the duration of each calculation phase, thereby obtaining the average duration of the multiple AI chips in the calculation phase.
[0070] In addition, the switching chip can also calculate the average duration of multiple AI chips in the communication phase, the number of times multiple AI chips are in the communication phase, and the duration of each time multiple AI chips are in the communication phase, thereby obtaining the standard deviation of the duration of multiple AI chips in the communication phase. Similarly, the switching chip can calculate the average duration of multiple AI chips in the calculation phase, the number of times multiple AI chips are in the calculation phase, and the duration of each time multiple AI chips are in the calculation phase, thereby obtaining the standard deviation of the duration of multiple AI chips in the calculation phase. After obtaining these traffic characteristic parameters, the switching chip can store these traffic characteristic parameters in a register as a power consumption adjustment strategy for the switching chip in the current statistical period.
[0071] Still as in the above example, statistics show that in T sample-period-0 In the process, after the number of times n that multiple AI chips are in the communication phase and the number of times m that multiple AI chips are in the calculation phase, the switching chip can detect whether n is close to m. If the two are very close, the switching chip can calculate the number of times in T using the following formula: sample-period-0 The average duration of multiple AI chips in the communication phase is:
[0072] Similarly, the switching chip can also be calculated by the following formula to obtain the value of T sample-period-0 The average length of time that multiple AI chips are in the computing phase is:
[0073] In addition, the switching chip can also be calculated by the following formula to obtain the T sample-period-0 The standard deviation of the duration that multiple AI chips spend in the communication phase is:
[0074] Similarly, the switching chip can also be calculated by the following formula to obtain the value of T sample-period-0 The standard deviation of the time that multiple AI chips spend in the computing phase is:
[0075] After obtaining these traffic characteristic parameters, the switching chip can store these traffic characteristic parameters in the register as T sample-period-1 The power consumption adjustment strategy in .
[0076] The following further introduces the operations performed by the switching chip in the current statistical period. FIG6 is another flowchart of the power consumption optimization method of the switching chip provided in an embodiment of the present application. As shown in FIG6 , the method includes:
[0077] 601. The switching chip samples the traffic forwarded by the switching chip to obtain sampled traffic.
[0078] In this embodiment, after entering the current statistical period, the switching chip may sample the traffic forwarded by the switching chip to obtain the sampled traffic. It should be noted that the current statistical period may include multiple sampling periods. Since the operations performed by the switching chip in each sampling period of the current statistical period are similar, the following description uses any sampling period in the current statistical period for schematic illustration, and this sampling period is referred to as the current sampling period. Therefore, the switching chip may sample the traffic being forwarded by the switching chip in the current sampling period, thereby obtaining the traffic sampled in the current sampling period.
[0079] 602. The switching chip detects whether the multiple AI chips are switching between the computing phase and the communication phase based on the sampled traffic.
[0080] 603. If multiple AI chips switch from the communication phase to the computing phase, the switching chip reduces power consumption for the traffic to be forwarded.
[0081] 604. If multiple AI chips switch from the computing phase to the communication phase, the switching chip increases the power consumption for the traffic to be forwarded.
[0082] After obtaining the sampled traffic, the switching chip can use the sampled traffic to detect whether multiple AI chips are switching between the computing phase and the communication phase to determine whether to adjust the power consumption of the traffic to be forwarded.
[0083] If it is determined based on the sampled traffic that multiple AI chips have not switched between the computing phase and the communication phase, that is, multiple AI chips are still in the computing phase or the communication phase, the switching chip does not adjust the power consumption for the traffic to be forwarded, and forwards the traffic to be forwarded in the future according to the current power consumption.
[0084] If it is determined based on the sampled traffic that multiple AI chips have just switched from the communication phase to the computing phase, that is, multiple AI chips have entered the computing phase from the communication phase, the switching chip will reduce the power consumption of the traffic to be forwarded according to the power consumption adjustment strategy, so as to forward the traffic to be forwarded subsequently according to the reduced power consumption.
[0085] If it is determined based on the sampled traffic that multiple AI chips have just switched from the computing stage to the communication stage, that is, multiple AI chips have entered the communication stage from the computing stage, the switching chip will increase the power consumption of the traffic to be forwarded according to the power consumption adjustment strategy, so as to forward the traffic to be forwarded subsequently according to the increased power consumption.
[0086] Specifically, the switching chip can detect whether multiple AI chips are switching between the computing phase and the communication phase in the following ways:
[0087] After obtaining the traffic sampled in the current sampling period, since the traffic sampled in the previous sampling period has also been obtained, the switching chip can compare the traffic sampled in the previous sampling period and the traffic sampled in the current sampling period, thereby detecting whether multiple AI chips are switching between the computing phase and the communication phase.
[0088] If the traffic sampled in the previous sampling period is equal to or similar to the traffic sampled in the current sampling period, the switching chip determines that the multiple AI chips have not switched between the computing phase and the communication phase. In other words, the multiple AI chips are still in the computing phase or the communication phase. The switching chip does not adjust the power consumption of the traffic to be forwarded, and forwards the traffic to be forwarded in the future with the current power consumption. For example, as shown in Figure 7 (Figure 7 is another schematic diagram of the statistical period and sampling period provided in the embodiment of the present application), in T sample- period-1 In the example, let T sample-interval-2 is the current sampling period, T sample-interval-1 is the previous sampling period, since at T sample- The flow rate collected in interval-1 is equal to that collected in T sample-interval-2 The traffic collected in the figure is large, so the switching chip can determine that multiple AI chips have not switched between phases and are still in the communication phase. Therefore, the power consumption of subsequent forwarding traffic is not adjusted.
[0089] If the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period to a certain extent, the switching chip determines that multiple AI chips have just switched from the communication phase to the calculation phase. In other words, multiple AI chips have entered the calculation phase from the communication phase. The switching chip reduces the power consumption of the traffic to be forwarded according to the power consumption adjustment strategy, and forwards the traffic to be forwarded in the future with the reduced power consumption. sample-period-1 In the example, let T sample-interval-4 is the current sampling period, T sample- Interval-3 is the previous sampling period, because in T sample-interval-3 The flow rate collected in T sample-interval-4 For the traffic collected in the process, the switching chip can determine multiple AI chips and enter the computing stage from the communication stage, which can reduce the power consumption of subsequent forwarding traffic.
[0090] If the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period to a certain extent, the switching chip determines that multiple AI chips have just switched from the computing phase to the communication phase. In other words, multiple AI chips have entered the communication phase from the computing phase. The switching chip increases the power consumption of the traffic to be forwarded according to the power consumption adjustment strategy, and forwards the traffic to be forwarded in the future with the increased power consumption. sample-period-1 In the example, let T sample-interval-6 is the current sampling period, T sample- Interval-5 is the previous sampling period, because in T sample-interval-5 The flow rate collected in T sample-interval-6 For the traffic collected in the process, the switching chip can determine multiple AI chips and enter the communication stage from the calculation stage, which can improve the power consumption of subsequent forwarding traffic.
[0091] More specifically, the switch chip can reduce power consumption for forwarded traffic by:
[0092] After determining that multiple AI chips have switched from the communication phase to the computing phase, the switching chip performs at least one of the following operations: (1) gradually reducing the operating frequency of the switching chip, for example, reducing the operating frequency of the switching chip by one level every 100 us, etc. (2) gradually reducing the operating voltage of the switching chip, for example, reducing the operating voltage of the switching chip by one level every 1000 us, etc. (3) turning off the serializer and deserializer (serdes) of the switching chip, for example, turning off serdes for 1ms, etc. In this way, the switching chip transmits traffic after reducing the operating frequency / operating voltage / turning off serdes.
[0093] Gradually reducing the operating frequency of the switch chip has a higher priority than gradually reducing the operating voltage of the switch chip, which in turn has a higher priority than shutting down the switch chip's SerDes (SerDes). For example, the switch chip continuously reduces its operating frequency, and when the operating frequency cannot be reduced any further, it reduces its operating voltage. Another example is that the switch chip continuously reduces its operating voltage, and when the operating voltage cannot be reduced any further, it shuts down its SerDes (SerDes).
[0094] More specifically, the switch chip can reduce power consumption for forwarded traffic in the following ways:
[0095] (a) After determining that multiple AI chips have switched from the computing phase to the communication phase, the switching chip performs at least one of the following when traffic to be forwarded arrives: (1) increasing the operating frequency of the switching chip, for example, increasing the operating frequency of the switching chip to the highest operating frequency. (2) increasing the operating voltage of the switching chip, for example, increasing the operating voltage of the switching chip to the highest operating voltage. (3) starting the serdes of the switching chip, for example, starting each serdes directly by the switching chip.
[0096] (b) Since the serdes startup time is long (usually 500us), it will affect the efficiency of the switch chip in forwarding traffic. Therefore, the switch chip can increase its power consumption in advance. That is, the switch chip can perform at least one of the following at a certain time before the arrival of the traffic to be forwarded, that is, at the target time: (1) Increase the operating frequency of the switch chip, for example, directly increase the operating frequency of the switch chip itself to the highest operating frequency, etc. (2) Increase the operating voltage of the switch chip, for example, directly increase the operating voltage of the switch chip itself to the highest operating voltage, etc. (3) Start the serdes of the switch chip, for example, directly start each serdes of the switch chip, etc.
[0097] More specifically, the switch chip can determine whether to increase power consumption immediately or in advance by:
[0098] Since the switching chip has stored the various traffic characteristic parameters calculated in the previous statistical period, that is, the power consumption adjustment strategy, when the switching chip needs to increase power consumption in the current statistical period, it can determine whether to increase power consumption immediately or in advance based on these traffic characteristic parameters.
[0099] If the average duration of multiple AI chips in the communication phase in the previous statistical cycle is zero, it means that multiple AI chips are in a dormant state in the previous statistical cycle. The switching chip can immediately increase power consumption. Therefore, the switching chip can perform at least one of the following when the traffic to be forwarded arrives: start the serdes of the switching chip, increase the operating voltage of the switching chip, and increase the operating frequency of the switching chip (generally, starting the serdes of the switching chip has a higher priority than increasing the operating voltage of the switching chip, and increasing the operating voltage of the switching chip has a higher priority than increasing the operating frequency of the switching chip). Still as in the above example, if T average If the value is 0, the switching chip can increase the power consumption for subsequent traffic when the traffic arrives, so as to forward the traffic according to the increased power consumption.
[0100] If the average duration of multiple AI chips in the communication phase in the previous statistical period and the average duration of multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of multiple AI chips in the communication phase in the previous statistical period is greater than or equal to a preset first threshold (the size of the threshold can be set according to actual needs and is not limited here), it means that the traffic between multiple AI chips does not conform to the typical AI traffic model (that is, the duration of different communication phases is not equal. In a typical AI traffic model, the duration of different communication phases is usually equal). The switching chip can increase power consumption immediately. At the moment when the traffic to be forwarded arrives, at least one of the following items is executed: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip. Still as in the above example, if T average >0,t average >0 and T standard-deviation >T sd-threshold (the aforementioned first threshold), the switching chip can increase the power consumption for subsequent traffic when the traffic arrives, and forward the traffic with the increased power consumption.
[0101] If the average duration of multiple AI chips in the communication phase in the previous statistical period and the average duration of multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of multiple AI chips in the calculation phase in the previous statistical period is greater than or equal to a preset second threshold (the size of the threshold can be set according to actual needs and is not limited here), it means that the traffic between multiple AI chips does not conform to the typical AI traffic model (that is, the duration of different calculation phases is not equal. In the typical AI traffic model, the duration of different communication phases is usually equal). The switching chip can increase power consumption immediately. At the moment when the traffic to be forwarded arrives, at least one of the following items is executed: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip. Still as in the above example, if T average >0,t average >0 and tstandard-deviation>tsd-threshold (the aforementioned second threshold), the switching chip can increase the power consumption for the subsequent traffic when the traffic arrives, and forward the traffic with the increased power consumption.
[0102] If the average duration of multiple AI chips in the communication phase in the previous statistical period and the average duration of multiple AI chips in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration of multiple AI chips in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration of multiple AI chips in the calculation phase in the previous statistical period is less than a preset second threshold, it indicates that the traffic between the multiple AI chips conforms to the typical AI traffic model and the switching chip can increase power consumption in advance. In this case, at least one of the following is performed at the target time: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the SerDes (Serdes) of the switching chip.
[0103] Among them, the difference between the target time and the arrival time of the traffic to be forwarded is calculated based on the average time that multiple AI chips are in the calculation stage in the previous statistical period and the standard deviation of the time that multiple AI chips are in the calculation stage in the previous statistical period.
[0104] Still as in the above example, T average >0,t average >0,T standard-deviation <T sd-threshold And tstandard-deviation<tsd-threshold, the switching chip can increase the power consumption for subsequent traffic at a certain time in advance and forward these traffic with the increased power consumption. Among them, the time is t0, t0=t1+(t average -tstandard-deviation-T serdes-up ), since the increase in power consumption is executed when the Ai chip enters the communication stage from the calculation stage, t1 is the initial moment when the Ai chip is in the calculation stage for the last time. It can be seen that t0 is between t1 and t2, and t2 is the moment when the Ai chip enters the current communication stage from the most recent calculation stage (that is, the initial moment when the Ai chip is in the calculation stage for the last time). After entering the communication stage, at a certain moment after t2, the subsequent traffic arrives.
[0105] In an embodiment of the present application, a switching chip can be used to transmit traffic between multiple AI chips, which can be used to train a neural network model. During the training process, the multiple AI chips can switch back and forth between the computation phase and the communication phase. During operation, the switching chip can sample the traffic being forwarded by the switching chip to obtain the sampled traffic. The switching chip can then use the sampled traffic to detect whether the multiple AI chips are switching between the computation phase and the communication phase, thereby determining whether to adjust the power consumption of the subsequent traffic to be forwarded. If the multiple AI chips switch from the communication phase to the computation phase, the switching chip can reduce the power consumption of the traffic to be forwarded, so that the traffic to be forwarded can be forwarded at the reduced power consumption. If the multiple AI chips switch from the computation phase to the communication phase, the switching chip can increase the power consumption of the traffic to be forwarded, so that the traffic to be forwarded can be forwarded at the increased power consumption. In the aforementioned process, the switching chip has power consumption adjustment capabilities, which can identify the changes in the AI chip phase based on the forwarded traffic and adaptively adjust the power consumption of the transmitted traffic based on these changes. It can be seen that when the AI chip is in the communication stage, the switching chip can be in a high-power state to forward traffic between AI chips. When the AI chip is in the computing stage, the switching chip can be in a low-power state to forward traffic between AI chips. In this way, the power consumption of the switching chip can match the actual workload, thereby avoiding the situation where the power consumption of the switching chip is inflated and achieving power consumption optimization.
[0106] Furthermore, in an embodiment of the present application, when the switching chip needs to reduce power consumption, a solution of gradually reducing the operating frequency, gradually reducing the operating voltage, and shutting down some serdes is adopted to achieve this. When switching from the communication stage to the computing stage, the switching chip can achieve smooth adjustment of power consumption, so that the traffic it forwards shows continuous reduction and smooth transition.
[0107] Furthermore, in an embodiment of the present application, when the switching chip needs to increase power consumption, it determines whether to increase power consumption immediately or in advance by judging whether the waveform of the current traffic conforms to the typical AI traffic model, thereby avoiding the long reset (startup) time of the serdes in the switching chip affecting the communication between AI chips.
[0108] The above is a detailed description of the power consumption optimization method of the switching chip provided in the embodiment of the present application. The switching chip provided in the embodiment of the present application will be introduced below. Figure 8 is a structural schematic diagram of the switching chip provided in the embodiment of the present application. As shown in Figure 8, the switching chip is used to forward traffic between multiple AI chips, and multiple AI chips are used to train neural networks. During the training process of the neural network, multiple AI chips can be in the computing stage or the communication stage. The switching chip includes:
[0109] The sampling module 801 is configured to sample the traffic forwarded by the switching chip to obtain sampled traffic;
[0110] A detection module 802 is configured to detect, based on the sampled traffic, whether the multiple AI chips are switching between the computing phase and the communication phase;
[0111] The first adjustment module 803 is configured to reduce the power consumption of traffic to be forwarded when multiple AI chips switch from the communication phase to the computing phase; or
[0112] The second adjustment module 804 is configured to increase the power consumption of the traffic to be forwarded when multiple AI chips switch from the computing phase to the communication phase.
[0113] In one possible implementation, the training process includes multiple sampling cycles, the multiple sampling cycles include the current sampling cycle, the sampled traffic is the traffic sampled in the current sampling cycle, and the detection module is used to detect whether multiple AI chips switch between the computing stage and the communication stage based on the traffic sampled in the previous sampling cycle and the traffic sampled in the current sampling cycle.
[0114] In one possible implementation, the first adjustment module is used to determine that multiple AI chips switch from the communication stage to the computing stage and reduce power consumption for the traffic to be forwarded if the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period.
[0115] In one possible implementation, the second adjustment module is used to determine that multiple AI chips switch from the computing stage to the communication stage and increase the power consumption of the traffic to be forwarded if the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period.
[0116] In one possible implementation, the first adjustment module is used to perform at least one of the following: gradually reducing the operating frequency of the switching chip, gradually reducing the operating voltage of the switching chip, and shutting down the serializer and deserializer Serdes of the switching chip, wherein the priority of gradually reducing the operating frequency of the switching chip is higher than the priority of gradually reducing the operating voltage of the switching chip, and the priority of gradually reducing the operating voltage of the switching chip is higher than the priority of shutting down the Serdes of the switching chip.
[0117] In one possible implementation, the second adjustment module is used to: execute at least one of the following at the moment when the traffic to be forwarded arrives: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or, execute at least one of the following at a target moment: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip, and the target moment is before the moment when the traffic to be forwarded arrives.
[0118] In one possible implementation, the training process includes multiple statistical cycles, the multiple statistical cycles include the previous statistical cycle and the current statistical cycle, the current statistical cycle includes the previous sampling cycle and the current sampling cycle, and the second adjustment module is used to: if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle is zero, at the moment when the traffic to be forwarded arrives, perform at least one of the following: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or, if the average duration of the multiple AI chips in the communication phase in the previous statistical cycle and the average duration of the multiple AI chips in the calculation phase in the previous statistical cycle are not zero, and the multiple AI chips in the previous statistical cycle If the standard deviation of the duration in the communication phase is greater than or equal to a preset first threshold, at least one of the following items is performed when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, if the average duration of multiple AI chips in the communication phase in the previous statistical period and the average duration of multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of multiple AI chips in the calculation phase in the previous statistical period is greater than or equal to a preset second threshold, at least one of the following items is performed when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip.
[0119] In one possible implementation, the second adjustment module is configured to, if the average duration of time that multiple AI chips were in the communication phase in a previous statistical period and the average duration that multiple AI chips were in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration that multiple AI chips were in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration that multiple AI chips were in the calculation phase in the previous statistical period is less than a preset second threshold, perform at least one of the following at a target time: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip.
[0120] In one possible implementation, the average duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; the average duration of time that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the number of times the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period; the standard deviation of the duration of time that multiple AI chips were in the communication phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the communication phase in the previous statistical period and the duration of each time the multiple AI chips were in the communication phase in the previous statistical period; and the standard deviation of the duration that multiple AI chips were in the calculation phase in the previous statistical period is calculated based on the average duration of time that the multiple AI chips were in the calculation phase in the previous statistical period and the duration of each time the multiple AI chips were in the calculation phase in the previous statistical period.
[0121] In one possible implementation, the difference between the target time and the time when the traffic to be forwarded arrives is calculated based on the average time that multiple AI chips were in the calculation phase in the previous statistical period and the standard deviation of the time that multiple AI chips were in the calculation phase in the previous statistical period.
[0122] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.
[0123] An embodiment of the present application also relates to an interconnection device, which can be presented as a chip, and the chip includes: a processing unit and a communication unit, wherein the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer-executable instructions stored in the storage unit so that the chip executes the power consumption optimization method described in the above embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0124] Specifically, please refer to Figure 9, which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be expressed as a neural network processor switching chip 900. The switching chip 900 is mounted on the AI chip as a coprocessor, and the AI chip assigns tasks (for example, forwarding traffic from the AI chip). The core part of the switching chip is the arithmetic circuit 903, which is controlled by the controller 904 to extract matrix data from the memory and perform multiplication operations.
[0125] In some implementations, the arithmetic circuit 903 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 903 is a two-dimensional systolic array. The arithmetic circuit 903 may also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 903 is a general-purpose matrix processor.
[0126] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 902 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 901 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 908.
[0127] Unified memory 906 is used to store input and output data. Weight data is directly transferred to weight memory 902 through the Direct Memory Access Controller (DMAC) 905. Input data is also transferred to unified memory 906 through the DMAC.
[0128] BIU stands for Bus Interface Unit, i.e., bus interface unit 913 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 909 .
[0129] The bus interface unit 913 (BIU) is used for the instruction fetch memory 909 to obtain instructions from the external memory, and is also used for the storage unit access controller 905 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0130] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 906 or transfer weight data to the weight memory 902 or transfer input data to the input memory 901.
[0131] The vector calculation unit 907 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 903, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used to cooperate with the operation circuit 903 to complete the parameter calculation of the power consumption adjustment strategy, etc.
[0132] In some implementations, the vector calculation unit 907 can store the processed output vector to the unified memory 906. For example, the vector calculation unit 907 can apply a linear function or a nonlinear function to the output of the operation circuit 903.
[0133] An instruction fetch buffer 909 connected to the controller 904 is used to store instructions used by the controller 904;
[0134] The unified memory 906, input memory 901, weight memory 902 and instruction fetch memory 909 are all on-chip memories. External memories are private to the switch chip hardware architecture.
[0135] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0136] An embodiment of the present application also relates to a computer storage medium, in which a program for signal processing is stored. When the program is run on a computer, the computer executes the steps executed by the aforementioned switching chip.
[0137] An embodiment of the present application also relates to a computer program product, which stores instructions that, when executed by a computer, enable the computer to perform the steps performed by the aforementioned switching chip.
[0138] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0139] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0140] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0141] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0142] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
Claims
1. A method for optimizing power consumption of a switching chip, characterized in that: The switching chip is used to forward traffic between multiple artificial intelligence (AI) chips, and the multiple AI chips are used to train a neural network. During the training of the neural network, the multiple AI chips are in a computing phase or a communication phase. The method includes: The switching chip samples the traffic forwarded by the switching chip to obtain sampled traffic; The switching chip detects, based on the sampled traffic, whether the multiple AI chips switch between the computing phase and the communication phase; If the multiple AI chips switch from the communication phase to the calculation phase, the switching chip reduces power consumption for the traffic to be forwarded; or, If the multiple AI chips switch from the computing phase to the communication phase, the switching chip increases power consumption for the traffic to be forwarded.
2. The method according to claim 1, characterized in that The training process includes multiple sampling cycles, the multiple sampling cycles include a current sampling cycle, the sampled traffic is traffic sampled in the current sampling cycle, and the switching chip detects whether the multiple AI chips switch between the calculation phase and the communication phase based on the sampled traffic, including: The switching chip detects whether the multiple AI chips switch between the calculation phase and the communication phase based on the traffic sampled in the previous sampling period and the traffic sampled in the current sampling period.
3. The method according to claim 2, characterized in that If the multiple AI chips switch from the communication phase to the calculation phase, the switching chip reduces power consumption of traffic to be forwarded, including: If the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period, the switching chip determines that the multiple AI chips switch from the communication phase to the calculation phase and reduces power consumption for the traffic to be forwarded.
4. The method according to claim 2, characterized in that If the multiple AI chips switch from the computing phase to the communication phase, the switching chip increases power consumption of traffic to be forwarded, including: If the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period, the switching chip determines that the multiple AI chips switch from the calculation stage to the communication stage and increases the power consumption of the traffic to be forwarded.
5. The method according to claim 3, characterized in that The switching chip reduces power consumption of traffic to be forwarded by: The switching chip performs at least one of the following: gradually reducing the operating frequency of the switching chip, gradually reducing the operating voltage of the switching chip, and shutting down the serializer and deserializer serdes of the switching chip, wherein the priority of gradually reducing the operating frequency of the switching chip is higher than the priority of gradually reducing the operating voltage of the switching chip, and the priority of gradually reducing the operating voltage of the switching chip is higher than the priority of shutting down the serdes of the switching chip.
6. The method according to claim 4, characterized in that The switching chip improves power consumption of traffic to be forwarded by: The switching chip performs at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or The switching chip performs at least one of the following at a target time: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip. The target time is before the time when the traffic to be forwarded arrives.
7. The method according to claim 6, characterized in that The training process includes multiple statistical cycles, the multiple statistical cycles include a previous statistical cycle and a current statistical cycle, the current statistical cycle includes the previous sampling cycle and the current sampling cycle, and the switching chip performs at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip, including: If the average duration of the multiple AI chips in the communication phase in the previous statistical period is zero, the switching chip performs at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or If the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period is greater than or equal to a preset first threshold, the switching chip performs at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or, If the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, and the multiple AI chips in the previous statistical period are The standard deviation of the duration of the calculation phase is greater than or equal to a preset second threshold, and the switching chip performs at least one of the following when the traffic to be forwarded arrives: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip.
8. The method according to claim 6, characterized in that The switching chip performs at least one of the following at the target time: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the SerDes of the switching chip, including: If the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period is less than a preset second threshold, the switching chip performs at least one of the following at the target time: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip.
9. The method according to claim 7 or 8, characterized in that The average duration that the multiple AI chips are in the communication phase in the previous statistical period is calculated based on the number of times the multiple AI chips are in the communication phase in the previous statistical period and the duration each time the multiple AI chips are in the communication phase in the previous statistical period; The average duration that the multiple AI chips are in the calculation phase in the previous statistical period is calculated based on the number of times the multiple AI chips are in the calculation phase in the previous statistical period and the duration each time the multiple AI chips are in the calculation phase in the previous statistical period; The standard deviation of the durations during which the multiple AI chips are in the communication phase in the previous statistical period is calculated based on the average duration during which the multiple AI chips are in the communication phase in the previous statistical period and the duration during which the multiple AI chips are in the communication phase each time in the previous statistical period. The standard deviation of the duration that the multiple AI chips are in the calculation stage in the previous statistical period is calculated based on the average duration that the multiple AI chips are in the calculation stage in the previous statistical period and the duration that the multiple AI chips are in the calculation stage each time in the previous statistical period.
10. The method according to claim 8, characterized in that The difference between the target time and the time when the traffic to be forwarded arrives is calculated based on the average duration of the multiple AI chips in the calculation phase in the previous statistical period and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period.
11. A switching chip, characterized in that: The switching chip is used to forward traffic between multiple AI chips, and the multiple AI chips are used to train a neural network. During the training process of the neural network, the multiple AI chips are in a computing phase or a communication phase. The switching chip includes: a sampling module, configured to sample the traffic forwarded by the switching chip to obtain sampled traffic; a detection module, configured to detect, based on the sampled traffic, whether the plurality of AI chips are switching between the computing phase and the communication phase; a first adjustment module, configured to reduce power consumption of traffic to be forwarded if the multiple AI chips switch from the communication phase to the calculation phase; or The second adjustment module is configured to increase the power consumption of the traffic to be forwarded if the multiple AI chips switch from the computing phase to the communication phase.
12. The switching chip according to claim 11, characterized in that: The training process includes multiple sampling cycles, the multiple sampling cycles include the current sampling cycle, the sampled traffic is the traffic sampled in the current sampling cycle, and the detection module is used to detect whether the multiple AI chips switch between the computing stage and the communication stage based on the traffic sampled in the previous sampling cycle and the traffic sampled in the current sampling cycle.
13. The switching chip according to claim 12, characterized in that: The first adjustment module is configured to determine that the multiple AI chips switch from the communication phase to the calculation phase if the traffic sampled in the previous sampling period is less than the traffic sampled in the current sampling period, and reduce power consumption for the traffic to be forwarded.
14. The switching chip according to claim 12, characterized in that: The second adjustment module is configured to determine that the multiple AI chips switch from the calculation phase to the communication phase if the traffic sampled in the previous sampling period is greater than the traffic sampled in the current sampling period, and to increase power consumption for the traffic to be forwarded.
15. The switching chip according to claim 13, characterized in that: The first adjustment module is used to perform at least one of the following: gradually reducing the operating frequency of the switching chip, gradually reducing the operating voltage of the switching chip, and shutting down the serializer and deserializer Serdes of the switching chip, wherein the priority of gradually reducing the operating frequency of the switching chip is higher than the priority of gradually reducing the operating voltage of the switching chip, and the priority of gradually reducing the operating voltage of the switching chip is higher than the priority of shutting down the Serdes of the switching chip.
16. The switching chip according to claim 14, characterized in that: The second adjustment module is configured to: When the traffic to be forwarded arrives, at least one of the following is performed: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip; or At a target time, at least one of the following is performed: increasing the operating frequency of the switching chip, increasing the operating voltage of the switching chip, and starting the serdes of the switching chip. The target time is before the time when the traffic to be forwarded arrives.
17. The switching chip according to claim 16, characterized in that: The training process includes multiple statistical cycles, the multiple statistical cycles include a previous statistical cycle and a current statistical cycle, the current statistical cycle includes the previous sampling cycle and the current sampling cycle, and the second adjustment module is used to: If the average duration of the multiple AI chips in the communication phase in the previous statistical period is zero, at the moment when the traffic to be forwarded arrives, perform at least one of the following: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or If the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period is greater than or equal to a preset first threshold, at the moment when the traffic to be forwarded arrives, perform at least one of the following: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip; or If the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period is greater than or equal to a preset second threshold, at the moment when the traffic to be forwarded arrives, perform at least one of the following: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the serdes of the switching chip.
18. The switching chip according to claim 16, characterized in that: The second adjustment module is configured to, if the average duration of the multiple AI chips in the communication phase in the previous statistical period and the average duration of the multiple AI chips in the calculation phase in the previous statistical period are not zero, the standard deviation of the duration of the multiple AI chips in the communication phase in the previous statistical period is less than a preset first threshold, and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period is less than a preset second threshold, perform at least one of the following at a target time: increase the operating frequency of the switching chip, increase the operating voltage of the switching chip, and start the SerDes of the switching chip.
19. The switching chip according to claim 17 or 18, characterized in that: The average duration that the multiple AI chips are in the communication phase in the previous statistical period is calculated based on the number of times the multiple AI chips are in the communication phase in the previous statistical period and the duration each time the multiple AI chips are in the communication phase in the previous statistical period; The average duration that the multiple AI chips are in the calculation phase in the previous statistical period is calculated based on the number of times the multiple AI chips are in the calculation phase in the previous statistical period and the duration each time the multiple AI chips are in the calculation phase in the previous statistical period; The standard deviation of the durations during which the multiple AI chips are in the communication phase in the previous statistical period is calculated based on the average duration during which the multiple AI chips are in the communication phase in the previous statistical period and the duration during which the multiple AI chips are in the communication phase each time in the previous statistical period. The standard deviation of the duration that the multiple AI chips are in the calculation stage in the previous statistical period is calculated based on the average duration that the multiple AI chips are in the calculation stage in the previous statistical period and the duration that the multiple AI chips are in the calculation stage each time in the previous statistical period.
20. The switching chip according to claim 18, characterized in that: The difference between the target time and the time when the traffic to be forwarded arrives is calculated based on the average duration of the multiple AI chips in the calculation phase in the previous statistical period and the standard deviation of the duration of the multiple AI chips in the calculation phase in the previous statistical period.
21. An interconnection device, characterized in that: The device includes a memory and a processor; the memory stores codes, and the processor is configured to execute the codes. When the codes are executed, the summary generating device performs the method according to any one of claims 1 to 10.
22. A computer storage medium, characterized in that The computer storage medium stores one or more instructions, which, when executed by one or more computers, enable the one or more computers to implement the method according to any one of claims 1 to 10.
23. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, enable the computer to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Power consumption optimization method for switch chip and related equipment thereof
CN120567780A
Systems and methods for dynamic temporal power steering
CN107003686A
Data exchange chip and server
CN112148663A
Network switching frequency dynamic adjustment method and system based on traffic sensing, and network switching chip structure
CN113132272A
Proactive management of inter-GPU network links
US20210064444A1