Uci e-based traffic control method, artificial intelligence chip and storage medium

By dynamically adjusting the credit value of the receiver's data buffer, the problem of data transmission failure caused by fixed reserved link margin in the UCIe protocol layer is solved, adaptive flow control is achieved, and the stability and efficiency of data transmission are improved.

CN121193679BActive Publication Date: 2026-02-06SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511746703.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-06
Estimated Expiration
2045-11-26

AI Technical Summary

Technical Problem

In the existing UCIe protocol layer overflow buffer flow control mechanism, the reserved link margin is a fixed value, which makes it easy for the receiving end data buffer to overflow when the link delay changes, causing data transmission failure and global reset.

Method used

Adaptive flow control is achieved by dynamically adjusting the credit value of the receiver's data buffer, utilizing multiple register configuration values ​​and reserved depth, and calculating the current credit value based on the remaining available depth.

Benefits of technology

It effectively reduces the risk of data buffer overflow at the receiving end caused by changes in link latency, and improves the stability and efficiency of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121193679B_ABST
    Figure CN121193679B_ABST
Patent Text Reader

Abstract

The application provides a UCIe-based flow control method, an artificial intelligence chip and a storage medium. The UCIe-based flow control method is suitable for a first core particle and a second core particle. The first core particle communicates with the second core particle. The first core particle comprises a receiving end data buffer. The UCIe-based flow control method comprises the following steps: reading, by the first core particle, an initial reserved link margin corresponding to the receiving end data buffer; reading, by the first core particle, a plurality of register configuration values and a plurality of reserved depths corresponding to the receiving end data buffer; obtaining, by the first core particle, a remaining available depth of the receiving end data buffer; and comparing, by the first core particle, the remaining available depth of the receiving end data buffer with the plurality of register configuration values to determine whether to select one of the plurality of reserved depths to calculate a current credit value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chip technology, and in particular to a flow control method based on UCIe, an artificial intelligence chip, and a computer-readable storage medium. Background Technology

[0002] The flow control mechanism for multiple virtual channels (VCs) between chiplets based on the Universal Chiplet Interconnect Express (UCIe) primarily utilizes a protocol layer design for credit-based flow control. Specifically, the receiver can pre-announce the buffer capacity (credit value) of each of its virtual channels to the sender. This allows the sender to transmit data within the credit range corresponding to each virtual channel, with each data transmission consuming a corresponding amount of credit. Upon confirmation by the receiver, the credit for the corresponding virtual channel can be replenished.

[0003] However, because the reserved link margin in the existing protocol layer overflow buffer flow control mechanism is a fixed value, in actual operation, when the link delay is greater than the reserved link delay for a certain period of time (for example, during a certain period of time, due to poor link quality, the probability of triggering retraining is high, which may result in the link being able to transmit normally to the other end, but the actual link delay may be greater than the reserved link delay), there is a probability that the buffered data of the receiving end data buffer will overflow, resulting in data transmission failure and triggering a global reset. Summary of the Invention

[0004] This invention relates to a flow control method, storage medium, and artificial intelligence chip based on UCIe, which can effectively update the credit values ​​of multiple virtual channels between two chips to achieve a good data flow control mechanism between the two chips.

[0005] According to embodiments of the present invention, the UCIe-based flow control method is applicable to a first core and a second core. The first core communicates with the second core. The first core includes a receiver data buffer. The UCIe-based flow control method includes the following steps: reading the initial reserved link margin corresponding to the receiver data buffer through the first core; reading multiple register configuration values ​​and multiple reserved depths corresponding to the receiver data buffer through the first core; obtaining the remaining available depth of the receiver data buffer through the first core; and comparing the remaining available depth of the receiver data buffer with the multiple register configuration values ​​through the first core to determine whether to select one of the multiple reserved depths to calculate the current credit value.

[0006] In the flow control method according to an embodiment of the present invention, the plurality of register configuration values ​​include a first register configuration value. The step of calculating the current credit value includes: when the first core determines that the remaining available depth is greater than the first register configuration value, subtracting the initial reserved link margin from the remaining available depth through the first core to obtain the current credit value.

[0007] In the flow control method according to an embodiment of the present invention, the plurality of reserved depths includes a first reserved depth. The step of calculating the current credit value further includes: when the first core determines that the remaining available depth is less than or equal to the first register configuration value, subtracting the initial reserved link margin from the remaining available depth through the first core, and subtracting the first reserved depth, to obtain the current credit value.

[0008] In the flow control method according to an embodiment of the present invention, a plurality of register configuration values ​​include a second register configuration value. The second register configuration value is less than a first register configuration value. A plurality of reserved depths include a second reserved depth. The step of calculating the current credit value includes: when a first core determines that the remaining available depth is less than or equal to the second register configuration value, subtracting the initial reserved link margin from the remaining available depth through the first core, and subtracting the second reserved depth, to obtain the current credit value.

[0009] In the flow control method according to an embodiment of the present invention, the second reserved depth is greater than the first reserved depth.

[0010] In the flow control method according to an embodiment of the present invention, the plurality of register configuration values ​​includes a third register configuration value. The third register configuration value is less than the second register configuration value. The step of calculating the current credit value includes: when the first core determines that the remaining available depth is less than or equal to the third register configuration value, setting the current credit value to 0 through the first core.

[0011] In the flow control method according to an embodiment of the present invention, the first chip includes a plurality of receiver data buffers corresponding to different virtual channels. The UCIe-based flow control method further includes: calculating a plurality of current credit values ​​corresponding to the plurality of receiver data buffers through the first chip; providing the plurality of current credit values ​​to a second chip through the first chip; and individually determining, through the second chip, whether to output data to the first chip via the corresponding virtual channel based on the plurality of current credit values.

[0012] According to an embodiment of the present invention, the UCIe-based artificial intelligence chip includes a first chip and a second chip. The first chip includes a plurality of receiver data buffers. The second chip is coupled to the first chip and is used to communicate with the first chip. The first chip reads the initial reserved link margin corresponding to the receiver data buffer, and reads a plurality of register configuration values ​​and a plurality of reserved depths corresponding to the receiver data buffer. The first chip obtains the remaining available depth of the receiver data buffer and compares the remaining available depth of the receiver data buffer with the plurality of register configuration values ​​to determine whether to select one of the plurality of reserved depths to calculate the current credit value.

[0013] In the artificial intelligence chip according to an embodiment of the present invention, multiple register configuration values ​​include a first register configuration value. When the first chip determines that the remaining available depth is greater than the first register configuration value, the first chip subtracts the remaining available depth from the initial reserved link margin to obtain the current credit value.

[0014] In the artificial intelligence chip according to an embodiment of the present invention, the multiple reserved depths include a first reserved depth and a second reserved depth. The multiple register configuration values ​​also include a second register configuration value, which is less than the first register configuration value. When the first chip determines that the remaining available depth is less than or equal to the first register configuration value and greater than the second register configuration value, the first chip subtracts the initial reserved link margin from the remaining available depth and subtracts the first reserved depth to obtain the current credit value.

[0015] In the artificial intelligence chip according to an embodiment of the present invention, the multiple register configuration values ​​also include a third register configuration value. The third register configuration value is less than the second register configuration value. The multiple reserved depths also include a second reserved depth. When the first chip determines that the remaining available depth is less than or equal to the second register configuration value and greater than the third register configuration value, the first chip subtracts the initial reserved link margin from the remaining available depth and subtracts the second reserved depth to obtain the current credit value.

[0016] In the artificial intelligence chip according to an embodiment of the present invention, the second reserved depth is greater than the first reserved depth.

[0017] In an artificial intelligence chip according to an embodiment of the present invention, when the first chip determines that the remaining available depth is less than or equal to the configuration value of the third register, the first chip sets the current credit value to 0.

[0018] In an artificial intelligence chip according to an embodiment of the present invention, a first chip includes a plurality of receiver data buffers corresponding to different virtual channels. The first chip calculates a plurality of current credit values ​​corresponding to the plurality of receiver data buffers. The first chip provides the plurality of current credit values ​​to a second chip. The second chip individually decides whether to output data to the first chip via the corresponding virtual channel based on the plurality of current credit values.

[0019] According to an embodiment of the present invention, the computer-readable storage medium is used to store a computer program. The computer program is executed by a processor to implement the steps of the UCIe-based flow control method described above.

[0020] Based on the above, the UCIe-based flow control method, storage medium, and artificial intelligence chip of the present invention can compare the remaining available depth of the receiving end data buffer by designing multiple register configuration values, and then implement an adaptive flow control mechanism for buffer depths with different data storage levels.

[0021] To make the above features and advantages of the present invention more apparent and understandable, specific embodiments are described below in conjunction with the accompanying drawings. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of a data transmission system according to an embodiment of the present invention;

[0023] Figure 2 This is a flowchart of a flow control method according to an embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of a data transmission system according to another embodiment of the present invention;

[0025] Figure 4 This is a flowchart illustrating the calculation of credit scores according to an embodiment of the present invention;

[0026] Figure 5 This is a schematic diagram showing the depth of the receiving end data buffer in an embodiment of the present invention;

[0027] Figure 6 This is a schematic diagram of an artificial intelligence chip according to an embodiment of the present invention.

[0028] Explanation of icon numbers

[0029] 10: Data transmission system;

[0030] 100: First chip;

[0031] 110: First transmitting module;

[0032] 210: Second transmitting module;

[0033] 111_1: Data buffer for the first transmitting end;

[0034] 111_M: Data buffer at the Mth transmitting end;

[0035] 112: NOC transmitter interface;

[0036] 113: Data packet encoder;

[0037] 114: Transmitter FDI interface;

[0038] 120: First receiving module;

[0039] 220: Second receiver module;

[0040] 121_1: Data buffer at receiver 1_1;

[0041] 121_N: The first_Nth receiver data buffer;

[0042] 221_1: Data buffer at the 2_1st receiver;

[0043] 221_N: The 2_Nth receiver data buffer;

[0044] 122: NOC receiver interface;

[0045] 123: Packet decoder;

[0046] 124: Receiver FDI interface;

[0047] 130: Protocol layer register;

[0048] 140: Credit score controller;

[0049] 200: Second core;

[0050] 500: Buffer space configuration;

[0051] 600: Artificial intelligence chip;

[0052] 610: First chip;

[0053] 620: Second core. Detailed Implementation

[0054] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same element symbols are used in the drawings and description to denote the same or similar parts.

[0055] Figure 1 This is a schematic diagram of a data transmission system according to an embodiment of the present invention. (See reference) Figure 1The data transmission system 10 includes a first core 100 and a second core 200. The first core 100 and the second core 200 are used for die-to-die communication based on UCIe. The first core 100 can be coupled to the second core 200, for example, through multiple bump lanes. In one embodiment of the invention, the first core 100 includes a first transmitting module 110 and a first receiving module 120. The second core 200 includes a second transmitting module 210 and a second receiving module 220. Multiple virtual channels can be established at the link layer between the first core 100 and the second core 200 to transmit data.

[0056] In one embodiment of the present invention, the first receiving module 120 includes a plurality of receiving data buffers (e.g., Figure 1 The first-to-first (1_1) receiver data buffers 121_1 to the first-to-N (1_N) receiver data buffers shown are illustrated, where N is a positive integer. The first-to-first (1_1) receiver data buffers 121_1 to the first-to-N (1_N) receiver data buffers correspond to different virtual channels. The second receiver module 220 includes multiple receiver data buffers (e.g., ...). Figure 1 The data buffers shown are the second-to-first receiving data buffers 221_1 to the second-to-N receiving data buffers 221_N. Each of these buffers corresponds to a different virtual channel. In one embodiment of the invention, the first transmitting module 110 may also include multiple transmitting data buffers, and the second transmitting module 210 may also include multiple transmitting data buffers.

[0057] In one embodiment of the present invention, the first receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N are used for independent data buffering of different virtual channels. The second receiving end data buffer 221_1 to the second_N receiving end data buffer 221_N are used for independent data buffering of different virtual channels. The first chip 100 can calculate the credit value corresponding to different virtual channels based on the usage of the first receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N. The second chip 200 can calculate the credit value corresponding to different virtual channels based on the usage of the second receiving end data buffer 221_1 to the second_N receiving end data buffer 221_N. The first chip 100 can determine whether to send data to the second chip 200 through the first transmitting end module 110 based on the credit values ​​corresponding to multiple virtual channels sent by the second chip 200. Furthermore, the second chip 200 can also determine whether to send data to the first chip 100 through the second transmitting module 210 based on the credit value corresponding to the multiple virtual channels sent by the first chip 100.

[0058] In one embodiment of the present invention, the data transmission system 10 may be implemented in an electronic device, and the electronic device includes a storage unit and a processor. The storage unit is used to store a computer program. The processor is coupled to the storage unit and is used to execute the computer program stored in the storage unit to cause the electronic device to perform the UCIe-based data transmission method as described in the embodiments of the present invention.

[0059] Processors may include, for example, a central processing unit (CPU) or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), programmable logic device (PLD), other similar processing devices, or combinations thereof.

[0060] Storage units may include, for example, random access memory (RAM), non-volatile memory, hard disk drive (HDD), or solid state drive (SSD). Random access memory may include, for example, dynamic random access memory (DRAM) or static random access memory (SRAM). Non-volatile memory may include, for example, flash memory or read-only memory (ROM).

[0061] Figure 2 This is a flowchart of a flow control method according to an embodiment of the present invention. (See also...) Figure 1 as well as Figure 2 To update Figure 1 Taking the credit value of the first receiving end data buffer 121_1 as an example, Figure 1 The data transmission system 10 can execute the following steps S210 to S240. In step S210, the first chip 100 reads the initial reserved link margin corresponding to the first_1 receiving end data buffer 121_1. In step S220, the first chip 100 reads multiple register configuration values ​​and multiple reserved depths corresponding to the first_1 receiving end data buffer 121_1. In step S230, the first chip 100 obtains the remaining available depth of the first_1 receiving end data buffer 121_1. In step S240, the first chip 100 compares the remaining available depth of the first_1 receiving end data buffer 121_1 with the multiple register configuration values ​​to determine whether to select one of the multiple reserved depths to calculate the current credit value.

[0062] In this embodiment, multiple registers are configured with different values, and the first chip 100 can compare the remaining available depth of the first-1 receiving data buffer 121_1 with each of the multiple register configuration values ​​to determine the remaining available depth level of the first-1 receiving data buffer 121_1, and decide whether to select one of the multiple reserved depths to calculate the current credit value. Therefore, the first chip 100 can select one of the multiple reserved depths corresponding to the remaining available depth level of the first-1 receiving data buffer 121_1 for dynamically adjusting the current credit value of the first-1 receiving data buffer 121_1. Thus, the flow control method of the present invention can implement an adaptive flow control mechanism for the first-1 receiving data buffer 121_1.

[0063] In this embodiment, the first chip 100 can provide the current credit value of the first_1 receiving end data buffer 121_1 to the second chip 200, so that the second chip 200 can individually determine whether the corresponding virtual channel should transmit data based on the current credit value of the first_1 receiving end data buffer 121_1 received from the first chip 100.

[0064] Furthermore, the first chip 100 can perform the operations described in steps S210 to S240 above for the first_1 receiving data buffer 121_1 to the first_N receiving data buffer 121_N of different virtual channels, respectively calculating multiple current credit values ​​corresponding to the first_1 receiving data buffer 121_1 to the first_N receiving data buffer 121_N. The first chip 100 can provide the multiple current credit values ​​of the first_1 receiving data buffer 121_1 to the first_N receiving data buffer 121_N to the second chip 200, so that the second chip 200 can individually decide whether to output data to the first chip 100 via the corresponding virtual channel based on the multiple current credit values.

[0065] Figure 3 This is a schematic diagram of a data transmission system according to another embodiment of the present invention. (See reference) Figure 3 The specific protocol layer architecture of the first chip 100 can be as follows: Figure 3 As shown, the specific protocol layer architecture of the second chip 200 can be similar to that of the first chip 100. In this embodiment, the first chip 100 includes a first transmitting module 110, a first receiving module 120, a protocol layer register 130, and a credit value controller 140. The credit value controller 140 is coupled to the first transmitting module 110, the first receiving module 120, and the protocol layer register 130. In this embodiment, the first transmitting module 110 includes a first transmitting data buffer 111_1 to an Mth transmitting data buffer 111_M, a NOC (Network on Chip) transmitting interface 112, a packet encoder 113, and a transmitting FDI (Flit-aware Die-To-Die Interface) 114. In this embodiment, the first receiving module 120 includes a first receiving data buffer 121_1 to a first receiving data buffer 121_N, an NOC receiving interface 122, a data packet decoder 123, and a receiving FDI interface 124.

[0066] In this embodiment, the NOC transmitter interface 112 and the NOC receiver interface 122 are used to couple the circuitry on the NOC side and to manage the interface timing and logic between UCIe and the NOC. The NOC transmitter interface 112 can process received NOC data packets and store them in at least one of the first to Mth transmitter data buffers 111_1 corresponding to a specific virtual channel. Furthermore, when one of the first to Mth transmitter data buffers 111_M outputs data to the other end, the NOC transmitter interface 112 can release a corresponding credit value to the NOC side, allowing the NOC side circuitry to retransmit data to the NOC transmitter interface 112 of the first chip 100. It is worth noting that the number of the first to Mth transmitter data buffers 111_1 can be determined based on the actual design of the NOC and the entire System on a Chip (SoC).

[0067] In this embodiment, the credit value controller 140 can individually determine whether the data cached in the first to Mth transmitter data buffers 111_M can be transmitted to the second core 200 via the packet encoder 113 and the transmitter FDI interface 114 based on the multiple credit values ​​corresponding to multiple virtual channels received from the second core 200. In this embodiment, the packet encoder 113 can encode and assemble the data of the flow units (flits) cached in the first to Mth transmitter data buffers 111_M. In this embodiment, the transmitter FDI interface 114 can transmit the encoded and assembled data packets and the updated credit value information to the adapter layer, and then output them to the second core 200, so that the second core 200 can determine whether to output data to the first core 100 via the corresponding virtual channel based on the received updated credit value information, so that the receiver data buffer in the first core 100 corresponding to this virtual channel can cache data.

[0068] In this embodiment, the receiving end FDI interface 124 can receive data from traffic units transmitted by the second chip 200. The receiving end FDI interface 124 can send the received traffic unit data to the packet decoder 123 for decoding, and provide corresponding data to at least one of the first_1 receiving end data buffers 121_1 to the first_N receiving end data buffers 121_N according to the corresponding virtual channel. The packet decoder 123 can parse the corresponding NOC data and credit value information from the traffic unit data. NOC data can be stored in the first_1 receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N, and the credit value controller 140 can calculate the current credit value of the first_1 receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N according to the remaining available depth of the first_1 receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N, so as to update the multiple credit values ​​corresponding to the first_1 receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N and encapsulate them into the corresponding data packet (i.e. the corresponding traffic unit), which is provided to the second core 200 by the first transmitting end module 110.

[0069] Figure 4 This is a flowchart illustrating the calculation of credit scores according to an embodiment of the present invention. Figure 5 This is a schematic diagram illustrating the depth of the receiving end data buffer according to an embodiment of the present invention. (Reference) Figure 3 as well as Figure 5 Credit scores can be calculated in ways such as... Figure 4 As shown. The first-level receiver data buffer 121_1 may have the following characteristics: Figure 5 The buffer space configuration shown is 500. Taking the credit value update of the first_1 receiver data buffer 121_1 as an example, in this embodiment, the credit value controller 140 first reads the pre-configured initial reserved link margin, first register configuration value, second register configuration value, third register configuration value, first reserved depth, and second reserved depth corresponding to the first_1 receiver data buffer 121_1 from the protocol layer register 130. Furthermore, the credit value controller 140 determines the remaining available depth of the first_1 receiver data buffer 121_1 as follows: Figure 5 As shown. The buffer space of the first-level receiver data buffer 121_1 may include a pre-allocated initial reserved link margin. The credit value controller 140 can obtain the remaining available depth by subtracting the used depth from the total buffer space of the first-level receiver data buffer 121_1.

[0070] In step S410, the credit value controller 140 can determine whether the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the first register configuration value. When the credit value controller 140 determines that the remaining available depth is greater than the first register configuration value (no), in step S420, the credit value controller 140 subtracts the remaining available depth from the initial reserved link margin to obtain the current credit value.

[0071] When the credit value controller 140 determines that the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the first register configuration value (yes), in step S430, the credit value controller 140 determines whether the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the second register configuration value. The second register configuration value is less than the first register configuration value.

[0072] When the credit value controller 140 determines that the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the first register configuration value and greater than the second register configuration value (No), in step S440, the credit value controller 140 subtracts the initial reserved link margin from the remaining available depth of the first_1 receiving end data buffer 121_1 and subtracts the first reserved depth to obtain the current credit value.

[0073] When the credit value controller 140 determines that the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the second register configuration value (yes), in step S450, the credit value controller 140 determines whether the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the third register configuration value. The third register configuration value is less than the second register configuration value.

[0074] When the credit value controller 140 determines that the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the second register configuration value and greater than the third register configuration value (No), in step S460, the credit value controller 140 subtracts the initial reserved link margin from the remaining available depth of the first_1 receiving end data buffer 121_1 and subtracts the second reserved depth to obtain the current credit value. The second reserved depth is greater than the first reserved depth.

[0075] When the credit value controller 140 determines that the remaining available depth of the first_1 receiving end data buffer 121_1 is less than or equal to the third register configuration value (yes), in step S470, the credit value controller 140 sets the current credit value to 0.

[0076] For example, the data buffer 121_1 at the first receiving end has a depth of 64, an initial reserved link margin of 20, a first register configuration value of 50, a second register configuration value of 40, a third register configuration value of 28, a first reserved depth of 3, and a second reserved depth of 5.

[0077] When the buffer space of the first receiving end data buffer 121_1 is occupied by a request sent by the other end, reducing the remaining available depth to 60, the credit value controller 140 can calculate the current credit value of the first receiving end data buffer 121_1 to be 40 (i.e., 60-20=40).

[0078] When the remaining available depth of the first_1 receiver data buffer 121_1 decreases to 50, the credit value controller 140 can calculate the current credit value of the first_1 receiver data buffer 121_1 as 27 (i.e., 50-20-3=27).

[0079] When the remaining available depth of the first_1 receiver data buffer 121_1 decreases to 45, the credit value controller 140 can calculate the current credit value of the first_1 receiver data buffer 121_1 as 22 (i.e., 45-20-3=22).

[0080] When the remaining available depth of the first_1 receiver data buffer 121_1 decreases to 40, the credit value controller 140 can calculate the current credit value of the first_1 receiver data buffer 121_1 as 15 (i.e., 40-20-5=15).

[0081] When the remaining available depth of the first_1 receiver data buffer 121_1 decreases to 30, the credit value controller 140 can calculate the current credit value of the first_1 receiver data buffer 121_1 as 5 (i.e., 30-20-5=5).

[0082] When the remaining available depth of the first_1 receiver data buffer 121_1 decreases to 28, the credit value controller 140 can directly set the current credit value of the first_1 receiver data buffer 121_1 to 0.

[0083] Therefore, the flow control method of the present invention can compare the remaining available depth of the receiving end data buffer by designing different register configuration values, and further utilizes a pre-set different reserved depth to perform an adaptive flow control mechanism for the buffer depth with different buffer space usage levels.

[0084] It is worth noting that the register configuration values ​​and the number of reserved depths in this invention are not limited to those specified in the invention. Figure 4Example. This invention can set multiple depth levels according to different usage requirements, and thus correspondingly set the register configuration values ​​and the actual value and quantity of the reserved depth. Furthermore, the process of the above embodiment can be applied to the current credit value update of each of the first_1 receiving end data buffer 121_1 to the first_N receiving end data buffer 121_N.

[0085] Figure 6 This is a schematic diagram of an artificial intelligence chip according to an embodiment of the present invention. (See reference) Figure 6 In one embodiment, the artificial intelligence chip 600 may include a first chip 610 and a second chip 620. The first chip 610 and the second chip 620 are used for inter-chip communication based on UCIe. The first chip 610 and the second chip 620 can form a data transmission system. For related implementation methods and technical details regarding the first chip 610 and the second chip 620, please refer to the descriptions of the first chip 100 and the second chip 200 in the above embodiments. Furthermore, the data transmission method between the first chip 610 and the second chip 620 can also refer to the processes in the above embodiments, thus providing sufficient teaching, suggestions, and implementation instructions. In addition, in another embodiment, the number of chips in the artificial intelligence chip 600 is not limited to... Figure 6 The first core 610 and the second core 620 are shown.

[0086] In one embodiment, the artificial intelligence chip 600 may be any one of a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processing unit (NPU), a deep learning processing unit (DPU), an accelerated processing unit (APU), and a general-purpose graphics processing unit (GPGPU).

[0087] In summary, the UCIe-based flow control method, artificial intelligence chip, and computer-readable storage medium of the present invention can adopt different control strategies according to the actual usage changes (depth changes) of the buffer space of the receiving end data buffer, so as to effectively reduce the situation where the receiving end data buffer overflows the effective buffer space capacity due to changes in link latency during actual application.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A UCIe-based flow control method, applicable to a first corelet and a second corelet, the first corelet and the second corelet being in communication, the first corelet comprising a receive-end data buffer, characterized in that, The UCIe-based flow control method comprises: reading, by the first corelet, an initial reserved link margin corresponding to the receive-end data buffer; reading, by the first corelet, a plurality of register configuration values and a plurality of reserved depths corresponding to the receive-end data buffer; obtaining, by the first corelet, a remaining available depth of the receive-end data buffer; and comparing, by the first corelet, the remaining available depth of the receive-end data buffer with the plurality of register configuration values to determine whether to select one of the plurality of reserved depths to calculate a current credit value. The plurality of reserved depths comprises a first reserved depth and a second reserved depth, and the plurality of register configuration values comprises a first register configuration value and a second register configuration value, wherein the second register configuration value is smaller than the first register configuration value. The step of calculating the current credit value comprises: when the first corelet determines that the remaining available depth is smaller than or equal to the first register configuration value and greater than the second register configuration value, the first corelet subtracts the initial reserved link margin and the first reserved depth from the remaining available depth to obtain the current credit value. The second reserved depth is greater than the first reserved depth.

2. The UCIe-based flow control method of claim 1, wherein, The step of calculating the current credit value comprises: when the first corelet determines that the remaining available depth is greater than the first register configuration value, the first corelet subtracts the initial reserved link margin from the remaining available depth to obtain the current credit value.

3. The UCIe-based flow control method of claim 1, wherein, The plurality of register configuration values further comprises a third register configuration value, wherein the third register configuration value is smaller than the second register configuration value. The step of calculating the current credit value comprises: when the first corelet determines that the remaining available depth is smaller than or equal to the second register configuration value and greater than the third register configuration value, the first corelet subtracts the initial reserved link margin and the second reserved depth from the remaining available depth to obtain the current credit value.

4. The UCIe-based flow control method of claim 3, wherein, The step of calculating the current credit value comprises: when the first corelet determines that the remaining available depth is smaller than or equal to the third register configuration value, the first corelet sets the current credit value to 0.

5. The UCIe-based flow control method of claim 1, wherein, The first corelet comprises a plurality of receive-end data buffers corresponding to different virtual channels, and the UCIe-based flow control method further comprises: calculating, by the first corelet, a plurality of current credit values corresponding to the plurality of receive-end data buffers, respectively; providing, by the first corelet, the plurality of current credit values to the second corelet; and determining, by the second corelet, whether to output data to the first corelet via a corresponding virtual channel according to the plurality of current credit values, respectively.

6. A UCIe-based artificial intelligence chip, comprising: a first corelet comprising a plurality of receive-end data buffers; and a second corelet coupled to the first corelet and configured to communicate with the first corelet. ​ wherein the first kernel reads an initial reserved link margin corresponding to the receive-side data buffer, and reads a plurality of register configuration values and a plurality of reserved depths corresponding to the receive-side data buffer; wherein the first kernel obtains a remaining available depth of the receive-side data buffer, and compares the remaining available depth of the receive-side data buffer with the plurality of register configuration values to determine whether to select one of the plurality of reserved depths to calculate a current credit value; wherein the plurality of reserved depths comprises a first reserved depth and a second reserved depth, and the plurality of register configuration values comprises a first register configuration value and a second register configuration value, the second register configuration value being smaller than the first register configuration value, wherein when the first kernel determines that the remaining available depth is smaller than or equal to the first register configuration value and greater than the second register configuration value, the first kernel subtracts the initial reserved link margin and the first reserved depth from the remaining available depth to obtain the current credit value; wherein the second reserved depth is greater than the first reserved depth.

7. The UCIe-based artificial intelligence chip according to claim 6, wherein, When the first kernel determines that the remaining available depth is greater than the first register configuration value, the first kernel subtracts the initial reserved link margin from the remaining available depth to obtain the current credit value.

8. The UCIe-based artificial intelligence chip according to claim 6, wherein, The plurality of register configuration values further comprises a third register configuration value, the third register configuration value being smaller than the second register configuration value, wherein when the first kernel determines that the remaining available depth is smaller than or equal to the second register configuration value and greater than the third register configuration value, the first kernel subtracts the initial reserved link margin and the second reserved depth from the remaining available depth to obtain the current credit value.

9. The UCIe-based artificial intelligence chip according to claim 8, wherein, When the first kernel determines that the remaining available depth is smaller than or equal to the third register configuration value, the first kernel sets the current credit value to 0.

10. The UCIe-based artificial intelligence chip according to claim 6, wherein, The first kernel comprises a plurality of receive-side data buffers corresponding to different virtual lanes, wherein the first kernel calculates a plurality of current credit values corresponding to the plurality of receive-side data buffers respectively, and the first kernel provides the plurality of current credit values to the second kernel, wherein the second kernel individually determines whether to output data to the first kernel via a corresponding virtual lane according to the plurality of current credit values.

11. A computer readable storage medium for storing a computer program, characterized in that, The computer program is executed by a processor to implement the steps of the UCIe-based flow control method of any one of claims 1-5.

Citation Information

Patent Citations

  • Inter-core-particle communication control method, device, equipment, medium and product

    CN120017598A

  • Data transmission method based on UCIe, storage medium and artificial intelligence chip

    CN120934696A