Data transmission method, device, medium and program product

Through custom data formats and arbitration scheduling based on CHI C2C protocol and combined with credit flow control mechanism, the traditional interconnection protocol has been solved in terms of bandwidth, latency and consistency, and efficient and reliable communication and dynamic resource balance between processing units are achieved.

CN120353747BActive Publication Date: 2025-09-02LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510847285.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-02
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Traditional interconnect protocols (such as PCIE) have difficulty meeting high-performance computing needs in terms of bandwidth, latency and consistency support, which has become a core obstacle to advanced processor architectures achieving computing power expansion through heterogeneous integration technologies.

Method used

The data transmission method based on the CHI C2C protocol is adopted, and sharding and arbitration scheduling are performed through custom data formats, combined with the credit flow control mechanism, efficient and reliable communication between processing units is achieved.

Benefits of technology

It realizes flexible data splitting and adaptation, effectively deals with data volume differences in different channels, prevents data overflow or transmission blockage, and improves communication efficiency and resource balance between processing units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353747B_ABST
    Figure CN120353747B_ABST
Patent Text Reader

Abstract

The present application discloses a data transmission method, device, medium and program product, which relate to the field of computer technology. By acquiring bus protocol data and performing sharding processing based on a custom data format, flexible splitting and adaptation of data during transmission between processing units are achieved, which can effectively cope with the transmission requirements of different data volumes of different channels. By arbitrating and scheduling the sharded data according to the channel type, data transmission of different types of data is achieved. By dynamically updating the sending side credit value based on the authorized credit value returned by the receiving side, flow control is performed to prevent data overflow or transmission congestion, thereby achieving efficient scheduling, reliable transmission and dynamic balance of resources for communication between processing units, and improving overall transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data transmission method, device, medium, and program product. Background Art

[0002] As semiconductor processes approach physical limits and the marginal benefits of Moore's Law diminish, single-module fully integrated solutions face challenges such as surging costs and extended design cycles. Simultaneously, the surging demand for heterogeneous computing in scenarios such as artificial intelligence and data centers has driven the rise of multi-chip heterogeneous integration technology (chiplets). This technology integrates computing units of different process nodes and functionalities at the packaging level to create high-performance heterogeneous system architectures. This technology not only significantly improves computing density and energy efficiency, but also significantly reduces the difficulty of developing complex functional modules. However, efficient interconnection between heterogeneous units has become a key bottleneck. Traditional interconnect protocols, such as PCIE (Peripheral Component Interconnect Express), struggle to meet the bandwidth, latency, and coherency requirements of high-performance computing. To address this, Arm, building on its mature AMBA (Advanced Microcontroller Bus Architecture) ecosystem, has launched the CHI-C2C (Coherent Hub Interface Chip-to-Chip) protocol as an extension of the AMBA CHI architecture. This protocol is deeply optimized for high-speed data exchange in multi-chip heterogeneous integrated systems and hybrid packaging systems. Currently, research on inter-unit communication implementation methods for the CHI-C2C protocol is still in its early stages, which has become a core obstacle hindering the expansion of computing power of advanced processor architectures through heterogeneous integration technology. Summary of the Invention

[0003] The present application provides a data transmission method, device, medium and program product, which can realize communication between processing units based on the CHI C2C protocol.

[0004] In a first aspect, the present application provides a data transmission method, comprising:

[0005] Obtain bus protocol data, and fragment the bus protocol data according to preset fragmentation rules based on a custom data format to generate different fragmented data;

[0006] Arbitrate and schedule the sliced ​​data according to the channel type to which the bus protocol data belongs, so as to divide the sliced ​​data into different slice data carriers to transmit the sliced ​​data and output the complete protocol frame;

[0007] The complete protocol frame is sent to the receiving side based on the sending side credit value, and the sending side credit value is updated according to the authorized credit value returned by the receiving side to dynamically control the number of slice data carriers sent; wherein, the authorized credit value is the credit value that can be authorized to the sending side by the receiving side dynamically generated according to the released slice data carrier when each complete protocol frame is processed.

[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data transmission methods when executing the computer program.

[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data transmission methods are implemented.

[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned data transmission methods when executed by a processor.

[0011] Beneficial effects: This method obtains bus protocol data and performs sharding processing based on a custom data format, thereby achieving flexible splitting and adaptation of data during transmission between processing units, and can effectively cope with the transmission requirements of different data volumes of different channels; by arbitrating and scheduling the sharded data according to the channel type, data transmission of different types of data is achieved, and by dynamically updating the sending side credit value based on the authorized credit value returned by the receiving side to perform flow control and prevent data overflow or transmission congestion, thereby achieving efficient scheduling of communication between processing units, reliable transmission and dynamic balance of resources, and improving overall transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 This is a flow chart of a data transmission method disclosed in this application;

[0014] Figure 2 A flow control diagram for implementing credit management between a receiving side and a sending side disclosed in this application;

[0015] Figure 3 A schematic diagram of slot scheduling and framing disclosed in this application;

[0016] Figure 4 A data transmission architecture diagram disclosed in this application;

[0017] Figure 5 This is a structural schematic diagram of a data transmission device disclosed in this application. DETAILED DESCRIPTION

[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0020] Multi-chip heterogeneous integrated systems combine different functional modules (such as CPUs (Central Processing Units), GPUs (Graphics Processing Units), and accelerators) in a heterogeneous integration approach, breaking through the area limitations of a single processing unit and reducing design complexity. Common multi-chip heterogeneous integrated systems utilize chiplet technology. With the development of chiplet technology, the CHI protocol has been further expanded to the CHI C2C protocol to support high-speed data transmission and communication between chiplets. Its purpose is to:

[0021] 1. Achieve high-bandwidth, low-latency interconnection: Support multi-processing unit symmetric multiprocessing (SMP) and accelerator connection topologies;

[0022] 2. Maintain cache consistency: support cross-chip shared memory and atomic operations;

[0023] 3. Compatibility with the existing ecosystem: Inherits the architectural features of the AMBA CHI single processing unit protocol (such as the Confidential Computing Extension (RME)), avoiding the performance loss caused by protocol conversion;

[0024] 4. Standardized chiplet interface: By defining a layered protocol stack that is decoupled from physical layer standards such as UCIE (Unified Compute Interface Extension) and CXL (Compute Express Link), flexible deployment is achieved.

[0025] Currently, the CHI protocol is implemented based on containers for communication between processing units. The CHI C2C container implementation adopts a container, which is the smallest fixed unit of transmission, with a size of 256 bytes and consists of three parts:

[0026] Link Header: Contains link control information (such as CRC checksum and sequence number);

[0027] Protocol Header: defines message type, credit allocation, resource plane identification, etc.

[0028] Payload: encapsulates the actual CHI C2C message;

[0029] Dynamic credit update: Credit is dynamically granted through the protocol header or MiscU messages, allowing the sender to adjust the flow based on the receiver's buffer capacity.

[0030] The aforementioned implementation scheme, based on the container mechanism, can improve communication efficiency within processing units or chiplet interconnects through layered encapsulation and optimization. However, container-based implementation is more complex, challenging to design, and consumes more resources. The combinational logic is often excessive, resulting in poor timing.

[0031] To this end, the present application provides a data transmission solution that can implement communication between processing units based on the CHI C2C protocol. In order to enable those skilled in the art to better understand the present application solution, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the data transmission method depends, the specific application environment architecture or specific hardware architecture is described here.

[0033] In the embodiment of the present application, a data transmission method is provided, which is described by applying the CHI protocol as an example. CHI is a bus protocol for internal communication between processing units. Through the present invention, the CHI protocols within two processing units can be opened up. Similarly, the present invention can also be applied to other internal buses, and used to open up the protocols of the internal buses of two processing units. In conjunction with the execution process of the data transmission method, see Figure 1 As shown, the method is described in detail, and the method includes:

[0034] Step S11: Obtain bus protocol data, and slice the bus protocol data according to a preset splitting rule based on a custom data format to generate different sliced ​​data.

[0035] In the embodiment of the present application, data transmission between processing units is performed based on the CHI protocol to realize the communication process. Among them, the processing unit refers to an integrated electronic module with data processing, storage and communication functions. Its hardware architecture may include components such as a computing core, a cache unit, and a communication interface. It can also realize data sending, receiving and protocol processing. For example, the processing unit may include a CPU (Central Processing Unit) core, a GPU (Graphics Processing Unit) computing unit, a cache controller and a high-speed communication interface.

[0036] First, the user layer or transport layer generates bus protocol data. Taking CHI protocol data as an example, it includes Req (Request), Rsp (Response), Snp (Snoop), and Dat (Data) type data, which arrives at the sending side.

[0037] Each data type then enters its own independent TX_queue (transmit queue), where it is queued and fragmented. During the fragmentation process, bus protocol data is fragmented according to pre-set splitting rules (based on the number of slots occupied by the protocol frame and the bit width of each channel) based on a custom data format, generating different fragmented data. The details of this process will be discussed in detail in the following embodiments and will not be elaborated upon here.

[0038] It's important to note that after system power-up or a link reset, the Initializer module must first perform an initialization process. Specifically, it receives an initialization handshake signal from the transmitter to the receiver, establishes a link connection between the two sides based on the initialization handshake signal, receives a loopback test request from the transmitter based on the link connection, and performs a functional test on the link connection based on the loopback test request. Only after the test passes will the link enter the working state.

[0039] First, a special Init Slot data is sent to complete the link initialization handshake. This data does not contain user data. The purpose is to allow the receiving end to detect the existence of the link and prepare for communication. Example: Initialization handshake flag: When (slot_valid=0)&(grant_valid=0)&(slot1=1), it is considered that the peer end has initiated the initialization handshake flag and established the initial connection.

[0040] After a successful handshake, a loopback test request is sent. The data is sent back to the local end (looped back) intact by the peer end to test the link loopback functionality. The sending end receives and checks the looped-back data. If the data matches, the physical link's send and receive functions are functioning properly. Example: Loopback Request Flags: When (slot_valid=0) & (grant_valid=0) & (slot1=0), the peer end considers a loopback request and loops back all received data to the peer end.

[0041] After the test is complete, the link enters the working state. Once in the working state, user-layer data or transport-layer data can be sent, and the link is ready to transmit actual CHI protocol data. It is worth noting that after initialization, an initial credit value must be sent to the peer. This initial credit value serves as the "permission token" that allows the sender to subsequently begin sending data. For example, the loopback frame transmission flag is (slot0=user_tx), (slot1=1). When slot 1 is 1, user information is sent via slot 0.

[0042] Step S12: Arbitrate and schedule the sliced ​​data according to the channel type to which the bus protocol data belongs, so as to divide the sliced ​​data into different slice data carriers for transmission, and output a complete protocol frame.

[0043] Since bus protocol data corresponds to four different data types, different data types are transmitted through different channels. Therefore, in the embodiment of the present application, four-channel heterogeneous scheduling (Req / Rsp / Dat / Snp) is used. The dispatcher on the sending side extracts the sliced ​​data from each TX_queue, schedules and frames it according to the custom data format, divides the sliced ​​data carriers (slots) for transmission, and generates a complete protocol frame after arbitration.

[0044] Step S13: Send the complete protocol frame to the receiving side based on the sending side credit value, and update the sending side credit value according to the authorized credit value returned by the receiving side to dynamically control the number of slice data carriers sent.

[0045] In the embodiment of the present application, a grant-based credit system is used to implement flow control and ensure transmission reliability, which is the key to ensuring that the receiving queue (Rx_queue) at the receiving end does not overflow. The sending side transmits the complete protocol frame to the receiving side through the physical link based on the sending side credit value. The protocol frame received by the physical link layer will pass through the pipe cache to smooth out transmission delays and clock domain differences. Further, based on the channel ID in the complete protocol frame that can be used to determine the data type, each slot data in the complete protocol frame is distributed to the corresponding receiving queue (Rx_queue). The receiving queue completes the transport layer protocol parsing based on the number of slots occupied by each channel bit width (consistent with the sending side TX_queue), and reassembles the received slot data into complete CHI protocol data.

[0046] It's important to note that the sender's credit value is used to record how many slice data carriers (slots) the local end can still send to the peer end. Slots can only be sent when this value is greater than 0, and sending slots consumes this value. After processing each complete protocol frame, the receiver's Rx_queue releases space and restores one credit. The credit value dynamically generated based on the released slice data carriers and authorized to the sender is the authorized credit value. This credit value is sent to the peer sender via the receiver's Tx_queue and used in the peer's TX_queue.

[0047] Furthermore, after the initialization phase ends and the system enters the working state, the receiving side (Rx_queue is empty) will have free slots equal to the initial maximum credit value. This initial credit value should not exceed the maximum depth of the receive queue, which is the upper limit of flow control. Simultaneously, the receiving side sets an authorization credit value based on the initial credit value and sends it to the peer end via Tx_queue. After the peer end receives this authorization credit value, its sending side credit value is set to this value before it can begin sending data.

[0048] Therefore, in the flow control mechanism, the first thing to do is to ensure that the receiving queue depth >= initial credit value >= authorized credit value. This inequality ensures that the authorized credit does not exceed the actual queue capacity and is a prerequisite for system security.

[0049] In this embodiment of the present application, the sending side credit value = the sending side's current credit value - the number of slice data carriers sent (the number of sending slots) + the authorized credit value; the receiving side credit value = the receiving side's current credit value - the authorized credit value + the recycled credit value. The recycled credit value is the number of slice data carriers released by the receiving side after processing each complete protocol frame.

[0050] like Figure 2The figure shows the credit flow control mechanism and queue interaction architecture used by the CHI protocol to implement communication between the receiving and transmitting sides. The receiving side's Rx_queue is a receive buffer divided by channel, where Req / Rsp / Snp / Dat ​​are associated with credits grant0-grant3, respectively. When the receiving Rx_queue finishes processing data from a channel, it releases buffer space and generates the corresponding grant signal. The grant signal is directly connected to the dispatcher, the dispatcher's scheduling module. The sending side's tr_queue is a transport layer queue used to buffer data to be sent. The dispatcher checks the grant value for each channel and consumes credit for each physical slot sent.

[0051] It is understandable that this embodiment targets flow control for point-to-point communications, and multi-node collaborative flow control can be considered in the future. For example, multiple receiving-side processing units can jointly allocate credits to the sending side. The receiving sides negotiate the total credit allocation based on the status of their respective buffers, and dynamically adjust the priority through "credit allocation weights" (such as increasing the credit allocation ratio of the GPU when computing urgently). In this way, resource allocation in multi-receiving-side scenarios can be optimized, preventing one receiving side from blocking the overall transmission due to insufficient buffer space, and improving the system's parallel processing capabilities.

[0052] In this embodiment, by acquiring bus protocol data and performing slicing processing based on a custom data format, flexible splitting and adaptation of data during transmission between processing units is achieved, which can effectively cope with the transmission requirements of different data volumes of different channels; by arbitrating and scheduling the slicing data according to the channel type, data transmission of different types of data is achieved, and by dynamically updating the sending side credit value based on the authorized credit value returned by the receiving side, flow control is performed to prevent data overflow or transmission congestion, thereby achieving efficient scheduling, reliable transmission and dynamic balance of resources for communication between processing units, thereby improving overall transmission efficiency.

[0053] Based on the aforementioned embodiment, CHI currently supports data bus widths of 128-bit, 256-bit, and 512-bit. In the embodiment of the present application, during data transmission between processing units, the transmission method is described using a fixed 256-bit data stream as the exchange unit on the C2C interface. The custom data format includes a first preset number of slice data carriers and a second preset number of flow control tokens. The total frame length determined based on the first preset number and the bit width occupied by a single slice data carrier, and based on the second preset number and the bit width occupied by a single flow control token, is the data bus width supported by the preset bus protocol. The preset bus protocol is used to transmit bus protocol data.

[0054] In the custom data format, 256 bits are divided into three slots (slice data carriers) and four grants (flow control tokens). A slot is the smallest unit within a frame that carries sliced ​​data (flit slices). A grant is a flow control token used to manage channel resources on the sending side (such as preventing overflow). Each grant frame corresponds to a channel. The custom data format is shown in Table 1 below:

[0055] Table 1 Custom data format

[0056]

[0057] In a specific embodiment, Slot contains three fields: a field (valid) used to indicate whether the data of the slice data carrier is valid, a data field (data) used to carry the slice data, and a channel identification field (chID) used to identify the channel type to which it belongs. The Slot format is shown in Table 2 below:

[0058] Table 2 Slot format

[0059]

[0060] In another specific implementation, the Grant contains two fields: a field (valid) used to indicate whether the channel authorization is valid, and a data field (data) used to update the credit value of the sending side. The Grant format is shown in Table 3 below:

[0061] Table 3 Grant format

[0062]

[0063] Here, "<<" is a left shift operator. For example, if the data value is 5, 1<<5 = 32. This value can be used to update the sender's credit value. If the credit counter corresponding to the original sender's credit value is A, then 32 is added to the original credit count, resulting in A+32.

[0064] Furthermore, since bus protocol data corresponds to four different data types, during data transmission, the channel types to which the bus protocol data belongs include a request channel (req), a response channel (rsp), a data channel (dat), and a snoop channel (snp). Accordingly, the field formats for different channels are defined differently. Therefore, the data transmission method further includes: determining the field formats corresponding to different channels, and determining the channel type to which the bus protocol data belongs based on configuration parameters of the field formats.

[0065] In the first specific implementation, the request channel (req) can transmit control requests (such as memory read / write, cache consistency operations, etc.). The field format of the request channel includes a request address field (Addr), an operation code field (Opcode), a data size field (Size), a memory attribute field (MemAttr), a field used to indicate whether to trigger a sniffing operation (SnpAttr), a transaction order control field (Order), a source node identification field (SrcID), a target node identification field (TgtID), and a transaction identification field (TxnID). Table 4 shows the request channel (req) Flit fields and their meanings:

[0066] Table 4 Request channel (req) Flit fields and their meanings

[0067]

[0068] In a second specific implementation, the response channel can transmit the response result of the request (such as data return, status code). The field format of the response channel includes a response type field (Resp), a first error code field (RespErr), a forwarding state field (FwdState), a response operation code field (Opcode), a source node identification field (SrcID), a target node identification field (TgtID), and a transaction identification field (TxnID). Table 5 shows the response channel (rsp) Flit fields and their meanings:

[0069] Table 5 Response channel (rsp) Flit fields and their meanings

[0070]

[0071] In a third specific implementation, the data channel can transmit large amounts of data payloads (such as memory data and intermediate calculation results). The field format of the data channel includes the actual data content field (Data), the byte enable field (BE), the data source field (DataSource), the data response status field (Resp), the second error code field (RespErr), the data operation code field (Opcode), the home node identification field (HomeNID), the source node identification field (SrcID), the target node identification field (TgtID), and the transaction identification field (TxnID). Table 6 shows the data channel (dat) Flit fields and their meanings:

[0072] Table 6 Data channel (dat) Flit fields and their meanings

[0073]

[0074] In a fourth specific implementation, the listening channel can transmit listening / monitoring signals (such as sniffing requests in the cache consistency protocol, monitoring the cache status of other processing units). The field format of the listening channel includes a sniffing address field (Addr), a sniffing operation code field (Opcode), a field used to indicate whether the result is returned to the source node (RetToSrc), a field used to indicate whether the data is allowed to become a shareable dirty state (DoNotGoToSD), a forwarding transaction identification field (FwdTxnID), a forwarding target node identification field (FwdNID), a source node identification field (SrcID), a target node identification field (TgtID), and a transaction identification field (TxnID). Table 7 shows the listening channel (snp) Flit fields and their meanings:

[0075] Table 7 Listening channel (snp) Flit fields and their meanings

[0076]

[0077] Based on the above embodiment, the content of the custom data format is introduced. Therefore, this embodiment introduces the fragmentation process. Specifically, based on the custom data format, the bus protocol data is fragmented according to the preset splitting rules to generate different fragmented data, including the following steps:

[0078] Determine the number of slices according to the matching relationship between the protocol data bit width corresponding to the channel type to which the bus protocol data belongs and the bit width of the slice data carrier;

[0079] The bus protocol data is sliced ​​according to the number of slices, and the data that is less than the bit width of the slice data carrier is padded to generate different slice data.

[0080] In this embodiment, the Flit data of each channel is divided into multiple slots based on its bit width (slots with less than a slot width are padded with 0s). The data bit width of each slot is 77 bits. Therefore, based on the bit width occupied by the data of different channel types, the request channel (req) and the listen channel (snp) occupy 2 slots, the response (rsp) channel occupies 1 slot, and the data (dat) channel occupies 5 slots. For example, for the request channel (req), its bit width is 93 bits and occupies 2 slots. Then, Slot 1 corresponds to the first 77 bits, and Slot 2 corresponds to the remaining 16 bits, and then 61 bits are padded with 0s.

[0081] In one feasible implementation, to reduce invalid padding overhead and improve transmission efficiency, a dynamic sharding strategy based on semantics can be implemented. For example, instead of fixed sharding based on bit width, the sharding granularity can be dynamically adjusted based on data semantics (e.g., instruction type, data structure). For example, small shards can be used for short instructions (e.g., control instructions) and large shards for large data blocks. The sharding attributes can be identified using a "semantic tag" field, allowing the receiving side to optimize the reassembly logic accordingly.

[0082] Furthermore, after the fragmentation process, slot scheduling and framing are performed. Specifically, the following steps are included:

[0083] The request channel and the listening channel are combined and arbitrated to generate a first subframe, and the response channel is fixedly filled to the end of the first subframe;

[0084] The data channel independently generates a second subframe, and performs combined arbitration on the first subframe and the second subframe to output a complete protocol frame.

[0085] In the embodiment of this application, Figure 3 As shown, the data (dat) channel is independently sliced, the request channel (req) and the snoop channel (snp) are combined for arbitration, and the response (rsp) channel is fixedly padded to the frame tail. Specifically, the request channel and the snoop channel compete for resources corresponding to the slice data carrier, and the winning channel is filled into the slice data carrier to generate the first subframe.

[0086] Because the request channel (req) and the listen channel (snp) both occupy two slots, and the response (rsp) occupies one slot, req and snp first arbitrate and compete for slot 0 and slot 1, and rsp is fixed to slot 2. The resulting protocol frame is combArb.

[0087] The data (DAT) channel requires five slots, so the protocol frame of the DAT channel group is further arbitrated with combArb to complete the final output. Specifically, the first and second subframes containing the winning and response channels are combined for arbitration, and the winning subframe is used as the complete protocol frame.

[0088] like Figure 4The diagram below shows the overall data transmission process for inter-processing unit communication. The upper section corresponds to the sending side process. The transport layer protocol interface receives different types of data (Req / Rsp / Snp / Dat) and stores them in the corresponding Tx_queue. User layer interfaces (such as io_user_tx) can also inject data. Data is queued and fragmented in the TX_queue (based on the number of slots occupied by the protocol frame and the channel width). The dispatcher then dispatches the data for framing and transmission. The dispatcher module extracts data from each Tx_queue as needed and assembles frames according to the protocol rules. The Arbitrator (Arbitrator) performs arbitration and selects the assembled frames for transmission over the physical link. The lower section then corresponds to the receiving side process. Protocol frames received by the physical link layer are buffered by pipes and distributed to different receive queues (Rx_queues) based on the channel IDs within the pipes. The receive queues then perform transport layer protocol parsing based on the number of slots occupied by their respective channel widths. It can be seen that by decoupling data from different channels through multiple queues, orderly framing and transmission are achieved with the help of scheduling and arbitration mechanisms, while supporting two-way interaction between the user layer and the transport layer, ensuring efficient and reliable communication between processing units.

[0089] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0090] The embodiment of the present application provides a data transmission device, see Figure 5 As shown, the device includes:

[0091] The slicing module 11 is used to obtain bus protocol data and slice the bus protocol data according to a preset splitting rule based on a custom data format to generate different slicing data;

[0092] The framing module 12 is used to arbitrate and schedule the sliced ​​data according to the channel type to which the bus protocol data belongs, so as to divide the sliced ​​data into different slice data carriers for transmission and output a complete protocol frame;

[0093] The flow control module 13 is used to send the complete protocol frame to the receiving side based on the sending side credit value, and update the sending side credit value according to the authorization credit value returned by the receiving side to dynamically control the number of slice data carriers sent; wherein the authorization credit value is the credit value that can be authorized to the sending side by the receiving side dynamically generated according to the released slice data carriers when each complete protocol frame is processed.

[0094] For the description of the features in the embodiment corresponding to the data transmission device, please refer to the relevant description of the embodiment corresponding to the data transmission method, and will not be repeated here.

[0095] In a feasible implementation manner, the data transmission device further includes an initialization module, which is configured to:

[0096] Acquire an initialization handshake signal sent by the sending side to the receiving side, and establish a link connection between the sending side and the receiving side based on the initialization handshake signal;

[0097] A loopback test request sent by the sending side is obtained based on the link connection, and a functional test is performed on the link connection according to the loopback test request, so as to enter a working state after the test passes.

[0098] The initial credit value sending module is used to send the initial credit value to the sending side through the receiving side based on the receiving queue depth of the receiving side when entering the working state for the first time.

[0099] The channel type judgment module is used to determine the field formats corresponding to different channels and determine the channel type to which the bus protocol data belongs based on the configuration parameters under the field format.

[0100] In a feasible implementation, the sharding module 11 is specifically configured to:

[0101] Determine the number of slices according to the matching relationship between the protocol data bit width corresponding to the channel type to which the bus protocol data belongs and the bit width of the slice data carrier;

[0102] The bus protocol data is sliced ​​according to the number of slices, and the data that is less than the bit width of the slice data carrier is padded to generate different slice data.

[0103] In a feasible implementation manner, the framing module 12 is specifically configured to:

[0104] The request channel and the listening channel are combined and arbitrated to generate a first subframe, and the response channel is fixedly filled to the end of the first subframe;

[0105] The data channel independently generates a second subframe, and performs combined arbitration on the first subframe and the second subframe to output a complete protocol frame.

[0106] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data transmission method embodiments.

[0107] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned data transmission method embodiments when running.

[0108] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0109] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned data transmission method embodiments are implemented.

[0110] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data transmission method embodiments are implemented.

[0111] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0112] The above is a detailed introduction to a data transmission method, device, medium, and program product provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data transmission method, characterized in that: include: Obtain bus protocol data, and slice the bus protocol data according to a preset splitting rule based on a custom data format to generate different sliced ​​data; Determine the number of slices according to a matching relationship between the protocol data bit width corresponding to the channel type to which the bus protocol data belongs and the bit width of the slice data carrier; The bus protocol data is sliced ​​according to the number of slices, and data that is less than the bit width of the slice data carrier is padded to generate different slice data; the channel types include a request channel, a response channel, a data channel, and a listening channel; the custom data format includes a first preset number of slice data carriers and a second preset number of flow control tokens; wherein the total frame length determined based on the first preset number and the bit width occupied by a single slice data carrier, and based on the second preset number and the bit width occupied by a single flow control token, is the data bus width supported by a preset bus protocol; the preset bus protocol is used to transmit the bus protocol data; Arbitrating and scheduling the sliced ​​data according to the channel type to which the bus protocol data belongs, so as to divide different sliced ​​data carriers to transmit the sliced ​​data, and output a complete protocol frame; the complete protocol frame is the winning subframe after merging and arbitrating the first subframe and the second subframe containing the winning channel and the response channel; the winning channel is the channel that wins when the request channel and the listening channel compete for the resources corresponding to the sliced ​​data carrier; The complete protocol frame is sent to the receiving side based on the sending side credit value, and the sending side credit value is updated according to the authorized credit value returned by the receiving side to dynamically control the number of slice data carriers sent; wherein the authorized credit value is a credit value that can be authorized to the sending side dynamically generated by the receiving side according to the released slice data carrier when each complete protocol frame is processed.

2. The data transmission method according to claim 1, wherein: Before obtaining the bus protocol data, the method further includes: Acquire an initialization handshake signal sent by the sending side to the receiving side, and establish a link connection between the sending side and the receiving side based on the initialization handshake signal; A loopback test request sent by the sending side is obtained based on the link connection, and a functional test is performed on the link connection according to the loopback test request, so as to enter a working state after the test passes.

3. The data transmission method according to claim 2, wherein: Also includes: When entering the working state for the first time, an initial credit value is sent to the sending side through the receiving side based on the receiving queue depth of the receiving side.

4. The data transmission method according to claim 1, wherein: The data format of the slice data carrier includes a field for indicating whether the data of the slice data carrier is valid, a data field for carrying the slice data, and a channel identification field for identifying the channel type to which it belongs.

5. The data transmission method according to claim 1, wherein: The data format of the flow control token includes a field for indicating whether the channel authorization is valid, and a data field for updating the sending side credit value.

6. The data transmission method according to claim 1, wherein: The data transmission method further includes: Determine the field formats corresponding to different channels, and determine the channel type to which the bus protocol data belongs according to the configuration parameters under the field format.

7. The data transmission method according to claim 6, characterized in that: Determining the field formats corresponding to different channels includes: The field format of the request channel includes a request address field, an operation code field, a data size field, a memory attribute field, a field for indicating whether a sniffing operation is triggered, a transaction sequence control field, a source node identification field, a target node identification field, and a transaction identification field; The field format of the response channel includes a response type field, a first error code field, a forwarding status field, a response operation code field, the source node identification field, the target node identification field, and the transaction identification field; The field format of the data channel includes an actual data content field, a byte enable field, a data source field, a data response status field, a second error code field, a data operation code field, a home node identification field, the source node identification field, the target node identification field, and the transaction identification field; The field format of the listening channel includes a sniffing address field, a sniffing opcode field, a field for indicating whether the result is returned to the source node, a field for indicating whether the data is allowed to become a shareable dirty state, a forwarding transaction identification field, a forwarding target node identification field, the source node identification field, the target node identification field, and the transaction identification field.

8. The data transmission method according to claim 6, wherein: The slice data is arbitrated and scheduled according to the channel type to which the bus protocol data belongs, so as to divide the slice data into different slice data carriers to transmit the slice data, and output a complete protocol frame, including: Combining and arbitrating the request channel and the listening channel to generate a first subframe, and filling the response channel to the end of the first subframe; The data channel is independently used to generate a second subframe, and the first subframe and the second subframe are combined and arbitrated to output a complete protocol frame.

9. The data transmission method according to claim 8, characterized in that: The combining and arbitrating the request channel and the listening channel to generate a first subframe includes: Competing between the request channel and the listening channel for resources corresponding to the slice data carrier, and filling the winning channel into the slice data carrier to generate a first subframe; Accordingly, performing combined arbitration on the first subframe and the second subframe to output a complete protocol frame includes: Combining arbitration is performed on the first subframe and the second subframe including the winning channel and the response channel, and the winning subframe is used as the complete protocol frame.

10. The data transmission method according to claim 1, wherein: The updating formula of the sending side credit value is the current credit value of the sending side - the number of slice data carriers sent + the authorized credit value; wherein, the authorized credit value is determined based on the receiving side credit value, and the updating formula of the receiving side credit value is the current credit value of the receiving side - the authorized credit value + the recovery credit value; the recovery credit value is the number of slice data carriers released by the receiving side when a complete protocol frame is processed.

11. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data transmission method according to any one of claims 1 to 10 when executing the computer program.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data transmission method according to any one of claims 1 to 10.

13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data transmission method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Ordered delivery of data packets based on type of path information in each packet

    CN116324747A

  • Data transmission method, device, equipment and medium

    CN118612076A