Interface circuit, processor comprising interface circuit and method of processing packets
By distributing message fields in the interface circuit and merging packets using extended information, the performance limitation of fixed message length in the system is solved, achieving efficient system communication and performance improvement.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2021-08-25
- Publication Date
- 2026-07-24
Smart Images

Figure CN114443516B_ABST
Abstract
Description
[0001] This application claims priority to Korean Patent Application No. 10-2020-0146196, filed on November 4, 2020, with the Korean Intellectual Property Office, the entire disclosure of which is incorporated herein by reference. Technical Field
[0002] The inventive concept relates to packet management, and more specifically, to interface circuitry, a processor, and a method for processing packets for generating extension packets. Background Technology
[0003] Devices configured to process data can perform operations by accessing memory. For example, a device can process data read from memory and write processed data back to memory. Due to desired system performance and functional requirements, various devices that provide high bandwidth and low latency and communicate with each other via links can be included in the system. The memory included in the system can be shared and accessed by two or more devices. Therefore, system performance can depend on the operating speed of individual devices, the efficiency of communication between devices, and the duration of memory access. Summary of the Invention
[0004] Embodiments of this disclosure may provide interface circuitry, a processor, and a method for processing packets that include messages with an increasing number of fields.
[0005] According to embodiments of this disclosure, an interface circuit is provided, the interface circuit comprising: a packet transmitter configured to: generate a plurality of transport packets based on a request output from a core circuit, and output the plurality of transport packets, the plurality of transport packets including information indicating packets to be merged; and a packet receiver configured to: generate a merged packet by merging a plurality of extended packets among a plurality of received packets received from outside the interface circuit, the plurality of extended packets including information indicating packets to be merged.
[0006] According to embodiments of the present disclosure, a processor is provided, the processor comprising: a controller configured to generate a message to be transmitted to an external device; and an interface circuit configured to generate packets to be output to a bus based on the message, wherein the interface circuit is configured to: generate multiple packets in different formats such that multiple fields included in the message are distributed in the multiple packets; and output the multiple packets to the bus.
[0007] According to embodiments of this disclosure, a method for processing packets is provided, the method comprising: receiving a plurality of received packets from an external source; identifying a plurality of extended packets among the plurality of received packets, the plurality of extended packets including information indicating packets to be merged; generating a merged packet by arranging a plurality of fields included in the plurality of extended packets; and outputting the merged packet. Attached Figure Description
[0008] Embodiments of this disclosure will become clearer from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0009] Figure 1 This is a block diagram illustrating a system according to an embodiment of the present disclosure;
[0010] Figure 2 This is a data diagram illustrating a package according to an embodiment of the present disclosure;
[0011] Figure 3 This is a data diagram illustrating messages according to embodiments of this disclosure;
[0012] Figure 4 This is a data diagram illustrating messages according to embodiments of this disclosure;
[0013] Figure 5 This is a data diagram illustrating messages according to embodiments of this disclosure;
[0014] Figure 6 This is a block diagram illustrating a system according to an embodiment of the present disclosure;
[0015] Figure 7 This is a block diagram illustrating an interface circuit according to an embodiment of the present disclosure;
[0016] Figure 8 This is a flowchart illustrating a method for processing a package according to an embodiment of the present disclosure;
[0017] Figure 9 This is a hybrid diagram illustrating a method for processing a package according to an embodiment of the present disclosure;
[0018] Figure 10A and Figure 10B This is a block diagram illustrating an example of a system according to an embodiment of the present disclosure; and
[0019] Figure 11 This is a block diagram illustrating a data center to which an embodiment of the system according to this disclosure is applied. Detailed Implementation
[0020] In the following description, embodiments of the present disclosure will be described with reference to the accompanying drawings.
[0021] Figure 1A system according to an embodiment of the present disclosure is illustrated. System 100 may include any computing system or components included in a computing system, the computing system including means 110 and a host processor 120 communicating with each other. For example, system 100 may be included in a fixed computing system (such as a desktop computer, server, self-service terminal, etc.) or system 100 may be included in a portable computing system (such as a laptop computer, mobile phone, wearable device, etc.). Furthermore, in some embodiments, system 100 may be included in a system-on-a-chip (SoC) or system-in-package (SiP) in which means 110 and host processor 120 are implemented in a single chip or package. Figure 1 As shown, system 100 may include device 110, host processor 120, device-attached memory 130, and host memory 140. In some embodiments, device-attached memory 130 may be omitted from system 100.
[0022] Reference Figure 1 Device 110 and host processor 120 can communicate with each other via link 150 and can perform message sending or receiving between them via link 150. Although embodiments of this disclosure will be described with reference to link 150 based on the CXL specification supporting the compute expresslink (CXL) protocol, device 110 and host processor 120 can communicate with each other based on coherent interconnect technologies such as, but not limited to, the XBus protocol, the NVLink protocol, the Infinity Fabric protocol, the Cache Coherent Interconnect for Accelerators (CCIX) protocol, or the Coherent Accelerator Processor Interface (CAPI).
[0023] In some embodiments, link 150 may support multiple protocols, and messages may be transmitted via multiple protocols. For example, link 150 may support the CXL protocol, which includes non-conformance protocols (such as CXL.io), conformance protocols (such as CXL.cache), and memory access protocols or memory protocols (such as CXL.mem). In some embodiments, link 150 may support protocols (such as, but not limited to, Peripheral Component Interconnect (PCI), PCIe, Universal Serial Bus (USB), or Serial Advanced Technology Attachment (SATA)). Here, the protocols supported by link 150 may also be referred to as interconnect protocols.
[0024] Device 110 may represent any means for providing useful functionality to host processor 120, and in some embodiments, device 110 may correspond to an accelerator conforming to the CXL specification. For example, software running on host processor 120 may offload at least some of computational operations and / or input / output (I / O) operations to device 110. In some embodiments, device 110 may include at least one of programmable components (such as graphics processing unit (GPU) or neural processing unit (NPU)), components providing fixed functionality (such as semiconductor intellectual property (IP) cores), and reconfigurable components (such as application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA) logic systems (such as, but not limited to, systems that use IP cores as building blocks)). Figure 1 As shown, device 110 may include physical layer 111, multiprotocol multiplexer 112, interface circuitry (e.g., first interface circuitry) 113 and accelerator circuitry 114, and device 110 may communicate with device attached memory 130.
[0025] The accelerator circuit 114 executes the device 110 to provide useful functions to the host processor 120, and may also be referred to as accelerator logic. When the device-attached memory 130 is included in, for example... Figure 1 In the system 100 shown, the accelerator circuit 114 can communicate with the device-attached memory 130 based on a protocol independent of link 150 (i.e., a device-specific protocol). Furthermore, as... Figure 1 As shown, accelerator circuit 114 can communicate with host processor 120 via interface circuit 113 using multiple protocols. Accelerator circuit 114 can generate requests and responses and transmit them to host processor 120 via interface circuit 113, thereby performing useful functions provided to host processor 120.
[0026] Interface circuitry 113 can determine one of a plurality of protocols based on messages used for communication between accelerator circuitry 114 and host processor 120. Interface circuitry 113 can be connected to at least one protocol queue included in multiprotocol multiplexer 112 and can send and receive messages to and from host processor 120 via at least one protocol queue. In some embodiments, interface circuitry 113 and multiprotocol multiplexer 112 can be integrated into a single component. In some embodiments, multiprotocol multiplexer 112 can include multiple protocol queues, each corresponding to a plurality of protocols supported by link 150. Furthermore, in some embodiments, multiprotocol multiplexer 112 can perform arbitration between communications using different protocols and can provide selected communications to physical layer 111. In some embodiments, physical layer 111 can be connected to physical layer 121 of host processor 120 via a single interconnect, bus, trace, etc. (but not limited to).
[0027] The host processor 120 may be the main processor of the system 100 (e.g., a central processing unit (CPU)), and in some embodiments, the host processor 120 may correspond to a host processor or host conforming to the CXL specification. Figure 1 As shown, the host processor 120 may be connected to the host memory 140 and may include a physical layer 121, a multiprotocol multiplexer 122, an interface circuit (e.g., a second interface circuit) 123, a coherence / cache circuit 124, a bus circuit 125, at least one core 126, and an I / O device 127.
[0028] At least one core 126 is executable and can be connected to coherent / caching circuitry 124. At least one core 126 can provide a request corresponding to the instruction to device 110 via interface circuitry 123. Coherent / caching circuitry 124 may include a cache hierarchy and may also be referred to as coherent / caching logic. Figure 1As shown, the coherence / caching circuit 124 can communicate with at least one core 126 and interface circuit 123. For example, the coherence / caching circuit 124 can allow communication via two or more protocols, including a coherence protocol and a memory access protocol. In some embodiments, the coherence / caching circuit 124 may include direct memory access (DMA) circuitry. The coherence / caching circuit 124 can generate requests and responses such that cache coherence between the device-attached memory 130 and the host memory 140 is maintained, and the coherence / caching circuit 124 can provide requests and responses to the device 110 via interface circuit 123. I / O device 127 can be used to communicate with bus circuit 125. For example, bus circuit 125 may be PCIe logic, and I / O device 127 may be a PCIe I / O device. Here, the components that generate requests or responses (e.g., accelerator circuit 114, at least one core 126, coherence / caching circuit 124, or bus circuit 125) may be referred to as core circuitry.
[0029] Interface circuitry 123 allows communication between device 110 and components of host processor 120 (e.g., coherence / caching circuitry 124 and bus circuitry 125). In some embodiments, interface circuitry 123 allows messages to be transferred between device 110 and components of host processor 120 according to multiple protocols (e.g., non-coherence protocols, coherence protocols, and / or memory protocols). For example, interface circuitry 123 may determine one of multiple protocols based on messages used for communication between components of device 110 and host processor 120.
[0030] The multiprotocol multiplexer 122 may include at least one protocol queue. Interface circuitry 123 may be connected to at least one protocol queue and may send messages to and receive messages from device 110 via at least one protocol queue. In some embodiments, interface circuitry 123 and multiprotocol multiplexer 122 may be integrated into a single component. In some embodiments, multiprotocol multiplexer 122 may include multiple protocol queues, each corresponding to a plurality of protocols supported by link 150. Furthermore, in some embodiments, multiprotocol multiplexer 122 may perform arbitration between communications using different protocols and may provide selected communications to physical layer 121.
[0031] In this embodiment, device 110 and host processor 120 can perform message sending and receiving between them. Messages provided by host processor 120 to device 110 may be referred to as host-to-device (H2D) requests, H2D responses, master-to-slave (M2S) requests, or M2S responses. Messages provided by device 110 to host processor 120 may be referred to as device-to-host (D2H) requests, D2H responses, slave-to-master (S2M) requests, or S2M responses. Because the information provided by device 110 and host processor 120 to each other varies depending on the type of protocol, the number, size, and type of fields included in the messages may vary.
[0032] The data unit (or data unit) transmitted via link 150 in each clock cycle may be referred to as a packet. Based on the CXL specification, a packet may also be referred to as a flow control unit or flow control number (flit) that forms a network packet or flow (such as a link-level atomic piece). In one embodiment, a packet may include multiple messages. Therefore, host processor 120 and device 110 can improve communication speed by including request messages and response messages simultaneously in a single packet. However, because the packet length may be fixed (e.g., the packet lengths may be equal to each other) to satisfy or maintain latency on link 150, the information that can be included in a single packet may be limited.
[0033] Because of the diversity of functionalities supported by the protocol, the number of fields included in a message can increase, or the types of fields in a message can be diversified. In this case, because although the message length inevitably increases, the packet length is fixed for the purpose of latency, a trade-off may occur between providing various functionalities and maintaining latency.
[0034] Each of the interface circuits 113 and 123 according to embodiments of the present disclosure can distribute the fields constituting a message into separate packets, thereby having the effect of increasing the number of fields constituting a message without modifying the length of the packets.
[0035] Figure 2 A package according to an embodiment of this disclosure is shown. (Refer to...) Figure 2The transmission packet 1000 may include, for example, a protocol identifier (ID) field, four time slots, and a cyclic redundancy check (CRC) field. Although the transmission packet 1000 is described as including four time slots, the number of time slots is not limited to this. The transmission packet 1000 may be a data unit transmitted via link 150. The protocol ID field may be information used to identify at least one of a plurality of protocols supported by link 150. A time slot may be a region including at least one message. In embodiments of this disclosure, a time slot may be a region including at least one extended message. The extended message may include information indicating a message to be merged with at least one other message (i.e., extended information). References are provided below. Figures 3 to 5 The extended message is described in more detail. The CRC field may include bits used for transmission error detection. A packet that includes a timeslot may be called a transaction packet. A packet that includes both a transaction field and a CRC field may be called a link packet. A packet that includes both a link packet and a protocol ID field may be called a physical layer packet.
[0036] In one embodiment, the message may include valid fields ( Figure 2 "valid" in the middle), opcode field ( Figure 2 "opcode" in the address field Figure 2 "ADDR" and reserved fields Figure 2 (as indicated by "RSVD" in the original text). Messages may also include additional fields. The number, size, and type of fields included in a message may vary depending on the protocol. Each field included in a message may include at least one bit. For example, a validity field may include one bit indicating that the message is valid. An opcode field may include multiple bits defining the operation corresponding to the message. For example, an opcode field may represent an opcode corresponding to a read or write command to memory. An address field may include multiple bits indicating the address associated with the opcode field. For example, when the opcode corresponds to a read command, the address field may indicate the address of the memory region where the read data is stored. Furthermore, when the opcode corresponds to a write command, the address field may indicate the address of the memory region where data will be written. Reserved fields may be areas that can include additional information. Therefore, information newly added to the message via the protocol may be included in reserved fields.
[0037] like Figure 2 As shown, because the packet length, in addition to the message length, may also be fixed to maintain latency on link 150, the information that can be included in a message may be limited. However, since the functionality supported by the protocol can be diverse, the number of fields included in a message can be increased, or the types of fields in the message can be diversified.
[0038] As described below, the interface circuit according to embodiments of this disclosure can distribute multiple fields included in a message across multiple packets, thereby increasing the number of fields included in the message without modifying the length of the packets.
[0039] Figure 3 The message illustrates an embodiment according to this disclosure. (See also...) Figure 3 According to embodiments of the present disclosure, extended messages may include extended information. Extended information may be information indicating whether a packet is to be merged with another packet. That is, a packet including extended information may be merged with another packet. In one embodiment, extended information may be represented by bit values recorded in an opcode field. That is, one of the bit values that can be recorded in the opcode field can be used to represent extended information, thereby representing extended information without consuming or increasing packet size. However, embodiments of the present disclosure are not limited to this, and one of the bit values that can be recorded in a validity field or an address field can be used to represent extended information. Although extended information is described as being recorded in the opcode field, opcodes other than extended information may also be represented by bit values. For example, when "111" is recorded in the opcode field, the corresponding packet may be an extended message and may be related to a write operation. However, embodiments of the present disclosure are not limited to this, and extended information may be represented by various bit values. Bit values recorded in the opcode field can indicate both opcode information and extended information. That is, bit values recorded in the opcode field can be decoded, thereby identifying that the corresponding packet is the packet to be merged, and identifying what operation will be performed on the merged packet. Reference Figure 3 Fields included in a single message can be distributed across a first extended message and a second extended message. In other words, a single message can be generated by merging the first and second extended messages. Each of the first and second extended messages can include extended information. In the first extended message, additional information can be recorded in the first field ( Figure 3 The additional information can be, for example, information used to maintain cache coherence between the host processor and the device. The additional information can be defined in various ways according to the relevant protocol. Because the message length can be fixed, the amount of additional information that can be included in the first extended message can be limited. Therefore, additional information not included in the first extended message can be included in the second extended message. That is, additional information not included in the first extended message can be recorded in the second field (field 1) of the second extended message. Figure 3 "Field 2" and / or the third field ( Figure 3In "Field 3" of the first extended message. The second extended message includes fields that do not overlap with the fields included in the first extended message, thereby allowing the host processor and device to exchange additional information with each other. For example, by changing the address field included in the first extended message to a second field that records additional information, the host processor and device can exchange additional information with each other. However, embodiments of this disclosure are not limited to this, and fields other than the address field can be changed to include fields that include additional information. Because the first extended message and the second extended message include different fields from each other, the first extended message and the second extended message can be described as being encoded in different formats from each other.
[0040] The first and second extended messages can be merged together. Specifically, fields included in the first and second extended messages can be integrated into a single stream to generate a merged packet (i.e., a single message). The merged packet can be decoded on a field-by-field basis, and the information recorded in each field can be extracted.
[0041] The length of the merged packet can be greater than the length of the first extended message or the length of the second extended message. By distributing multiple fields included in a single message across multiple packets and exchanging multiple packets, the interface circuitry according to embodiments of this disclosure can have the effect of increasing the number of fields included in a message, even when packets of finite length are used.
[0042] Figure 4 The message illustrates an embodiment according to this disclosure. (See also...) Figure 4 Extended messages may include extended fields specifically used to record extended information. Figure 4 (ext in the text). Figure 3 Unlike extended messages, the opcode or other fields do not need to include extended information. For example, an extended field can be implemented with 1 bit, where "1" can be recorded in the extended field when the corresponding packet is an extended message, and "0" can be recorded in the extended field when the corresponding packet is not an extended message.
[0043] Fields other than the extended fields included in the first and second extended messages can be integrated into a single stream to generate a merged packet (i.e., a single message).
[0044] Figure 5 The message illustrates an embodiment according to this disclosure. (See also...) Figure 5 The extended message may also include order information indicating the arrangement order of the fields. In one embodiment, the extended message may include an order field specifically for order information. In another embodiment, the order information may be recorded in a separate field; for example, the order information may be recorded in the opcode field along with the opcode itself. Although Figure 3This shows that extended information can be recorded in the opcode field, but extended information can optionally be recorded in other fields such as... Figure 4 In the dedicated fields shown.
[0045] In one embodiment, information indicating the first order may be included in the first extended message, and information indicating the second order may be included in the second extended message. When the first and second extended messages are merged, the fields included in the first and second extended messages can be arranged with reference to the order information. That is, the first field included in the first extended message ( Figure 5 The "Field 1" in the second message is arranged to include the second and third fields in the second extended message. Figure 5 Fields “2” and “3” in the document can be arranged in the order stated. Although the fields constituting a message are described as being distributed across two extended messages, embodiments of this disclosure are not limited thereto, and the fields constituting a message can be distributed across multiple extended messages. Because the number of cases where multiple extended messages are arranged can be reduced by including order information in each of the multiple extended messages, resources used for decoding the merged message can be saved. In one embodiment, the order field may include information indicating the last of the multiple extended messages corresponding to a message. In one embodiment, an extended message may include a field specifically for indicating whether an extended message is the last extended message. See also Figure 5 Since the first field included in the first extended message and the second and third fields, both included in the second extended message, are arranged in the order of presentation, it can be understood that the last extended message is the second extended message, but is not limited thereto.
[0046] Figure 6 A system according to an embodiment of the present disclosure is shown. (Refer to...) Figure 6 System 100a may include device 200 and host processor 300. Device 200 and host processor 300 may be respectively Figure 1 Examples of device 110 and host processor 120.
[0047] The device 200 may include an accelerator circuit 210, a first controller 220, and a first interface circuit 230. The accelerator circuit 210 may be... Figure 1An example of accelerator circuit 114. A first controller 220 may be connected to accelerator circuit 210, and may generate multiple messages for processing requests received from accelerator circuit 210, and may provide the multiple messages to first interface circuit 230. The first controller 220 may be connected to first interface circuit 230, and may generate responses corresponding to requests based on the multiple messages received from first interface circuit 230, and may provide the responses to accelerator circuit 210. First interface circuit 230 may include a first packet transmitter 231 and a first packet receiver 232. First packet transmitter 231 may generate packets for transmitting multiple messages to host processor 300. First packet transmitter 231 may generate multiple extended messages corresponding to a message. Multiple extended messages may be transmitted to second interface circuit 330. Referring to the above... Figures 3 to 5 The extended messages may include extended information indicating packets to be merged, and multiple fields constituting a single message may be distributed across multiple extended messages. The first packet receiver 232 may receive multiple extended messages from the second interface circuit 330 and may generate a single message by merging the multiple extended messages.
[0048] The host processor 300 may include at least one core 310, a second controller 320, and a second interface circuit 330. The at least one core 310 can execute multiple instructions. The second controller 320 may include... Figure 1 A consistency / caching circuit 124 is included. A second controller 320 may be connected to at least one core 310, and can generate multiple messages for processing requests received from at least one core 310, and can provide the multiple messages to a second interface circuit 330. The second controller 320 may be connected to the second interface circuit 330, and can generate a response corresponding to the request based on the multiple messages received from the second interface circuit 330, and can provide the response to at least one core 310. The second interface circuit 330 may include a second packet transmitter 331 and a second packet receiver 332. The second packet transmitter 331 may be configured substantially the same as the first packet transmitter 231, and the second packet receiver 332 may be configured substantially the same as the first packet receiver 232. Therefore, the description of the first packet transmitter 231 may be applied to the second packet transmitter 331, and the description of the first packet receiver 232 may be applied to the second packet receiver 332, so repeated descriptions may be omitted.
[0049] Figure 7 An interface circuit according to an embodiment of the present disclosure is shown. (Refer to...) Figure 7 The first packet transmitter 231 can provide multiple extended messages e_m1 and e_m2 to the second packet receiver 332. Although reference... Figure 7The packet transmission between the first packet transmitter 231 and the second packet receiver 332 is described, but the packet transmission between the second packet transmitter 331 and the first packet receiver 232 can also be performed in substantially the same way, so the repeated description can be omitted.
[0050] The first packet transmitter 231 may include a transaction packet generator 410, a link packet generator 420, and an extension capability register 430. The transaction packet generator 410 may receive messages from the first controller 220 and may generate multiple extended messages e_m1 and e_m2, each containing multiple fields constituting a message. In one embodiment, the transaction packet generator 410 may include the generated extended messages e_m1 and e_m2 in different transaction packets. In another embodiment, the transaction packet generator 410 may include the generated extended messages e_m1 and e_m2 in different time slots contained within a single transaction packet. The transaction packet generator 410 may provide the transaction packet including the extended messages to the link packet generator 420. The link packet generator 420 may generate a link packet by recording bits used for transmission error detection in a CRC field. The link packet may be sent to the physical layer, and the physical layer may generate a transport packet by adding a protocol ID field to the link packet. The link packet generator 420 may be connected via a bus (e.g., Figure 1 The link 150 output includes at least one transport packet containing multiple extended messages e_m1 and e_m2. The extension capability register 430 can store information about whether a peer device that will receive the transaction packet can recognize the extended information. The transaction packet generator 410 can refer to the extension capability register 430 to determine whether to generate an extended message. In one example, the transaction packet generator 410 can refer to the extension capability register 430 to provide the transport packet to an external device capable of recognizing information indicating that the packet will be merged.
[0051] The second packet receiver 332 may include an extended packet detector 510, a packet merger 520, and a buffer 530. The extended packet detector 510 can identify extended messages included in a time slot by examining extended information. The extended packet detector 510 can transmit the extended messages to the packet merger 520. The packet merger 520 can generate a merged packet (i.e., a message) by merging multiple extended messages. The packet merger 520 can temporarily store the extended messages in the buffer 530 and can generate a single message by arranging fields included in multiple extended messages stored in the buffer 530. The packet merger 520 can provide the message to the second controller 320.
[0052] Figure 8 A method for processing a package according to an embodiment of this disclosure is illustrated. Specifically, Figure 8 The method can be derived from Figure 1 The second interface circuit 123 or Figure 6An example of a method for processing packets executed by the second interface circuit 330. The method for processing packets may include multiple operations S801 to S807.
[0053] Reference Figure 8 In operation S801, the second interface circuit 123 or 330 can receive messages. The message can be an extended message including extended information, or a normal message without extended information.
[0054] In operation S802, the second interface circuit 123 or 330 can determine whether the received message includes extended information. Specifically, the second interface circuit 123 or 330 can determine whether the received message includes extended information by reading bits in a field recorded at a preset position. The field at the preset position can be a field specifically for extended information. Optionally, the field at the preset position can be an opcode field. When the received message includes extended information (i.e., when the received message is an extended message), operation S804 can be executed, and when the received message does not include extended information (i.e., when the received message is a normal message), operation S803 can be executed.
[0055] In operation of S803, the second interface circuit 123 or 330 can provide ordinary messages. Figure 6 The second controller 320. Ordinary messages can represent messages that are not merged. The second controller 320 can generate messages provided based on ordinary messages. Figure 6 The response of at least one core 310.
[0056] In operation S804, the second interface circuit 123 or 330 can determine whether the received extended message is the last extended message. (Refer to the above.) Figure 5 The extended message may include information indicating whether it is the last extended message. Therefore, the second interface circuit 123 or 330 can determine whether the received extended message is the last extended message by identifying a sequence field or a field specifically used to indicate whether the extended message is the last extended message. When the received extended message is the last extended message, operation S806 can be executed. Otherwise, when the received extended message is not the last extended message, operation S805 can be executed.
[0057] In operation S805, the second interface circuit 123 or 330 can store the received extended messages in a buffer. That is, the second interface circuit 123 can temporarily store the received extended messages in the buffer until all extended messages corresponding to a message have been received.
[0058] In operation S806, the second interface circuit 123 or 330 can generate a message by merging extended messages stored in the buffer. Specifically, the second interface circuit 123 or 330 can generate a message by integrating the fields included in the extended messages into a sequence in a manner that prevents the fields from overlapping. The second interface circuit 123 or 330 can determine the order in which the fields are arranged into a sequence based on the order information included in the extended messages.
[0059] In operation S807, the second interface circuit 123 or 330 can provide a merged message to the second controller 320. The second controller 320 can generate a response to be provided to at least one core 310 based on the merged message.
[0060] Figure 9 A method for processing a package according to an embodiment of this disclosure is shown. For example... Figure 9 As shown, the method for processing a packet may include multiple operations S901 to S905. Referring below... Figure 1 , Figure 6 and Figure 7 Conducting discussions on Figure 9 The description.
[0061] In operating S901, Figure 1 First interface circuit 113 or Figure 6 The first interface circuit 230 can generate and from Figure 6 The first controller 220 receives multiple extended messages corresponding to the message it receives. The extended messages may include extended information indicating packets to be merged. The first interface circuitry 113 or 230 may generate multiple extended messages, whereby multiple fields included in one message may be distributed across multiple extended messages. The length of the extended messages may be less than the length of the message. In one embodiment, the extended messages may include fields specifically for extended information. In one embodiment, the extended information may be represented by bit values recorded in an opcode field. The first controller 220 may be based on... Figure 6 The accelerator circuit 210 receives the request to generate a message. Although Figure 8 A method for processing packets transmitted from a first interface circuit 113 or 230 to a second interface circuit 123 or 330 is shown, but Figure 8 The method can also be applied to methods of processing packets transmitted from the second interface circuit 123 or 330 to the first interface circuit 113 or 230. In this case, the second controller 320 may generate a message based on a request received from at least one core 310, and the second interface circuit 123 or 330 may generate multiple extended messages corresponding to the message received from the second controller 320. In one example, the sum of the lengths of the multiple fields used to process the request is greater than the length of each extended message.
[0062] In operation S902, the first interface circuit 113 or 230 can provide multiple extended messages to the second interface circuit 123 or 330. In one example, the multiple extended messages may be included in different transport packets. In another example, the multiple extended messages may be included in different time slots within the same transport packet.
[0063] In operation S903, the second interface circuit 123 or 330 can detect extended information included in the extended message. By detecting the extended information, the second interface circuit 123 or 330 can identify the extended message that will be merged into a single message and the ordinary message that will not be merged.
[0064] In operation S904, the second interface circuit 123 or 330 can temporarily store extended messages in a buffer. The buffer can be... Figure 7 The buffer 530 is shown. Specifically, in order to merge sequentially received extended messages, the second interface circuit 123 or 330 can store the received extended messages in the buffer. The extended messages stored in the buffer can be kept in the buffer until all extended messages corresponding to a message have been fully received.
[0065] In operation S905, the second interface circuit 123 or 330 can generate a message by merging extended messages. Specifically, the second interface circuit 123 or 330 can generate a message by arranging multiple fields included in the extended messages and combining the multiple fields into a sequence. In one example, the second interface circuit 123 or 330 can generate a message by merging extended messages temporarily stored in a buffer. In one example, the second interface circuit 123 or 330 can determine the arrangement order of fields included in multiple extended messages based on the order information included in the extended messages. The second interface circuit 123 or 330 can provide the generated message to the second controller 320.
[0066] The processing packet method according to embodiments of the present disclosure can send multiple fields simultaneously, which include multiple fields in one message, distributed across multiple extended messages, thereby having the effect of increasing the number of fields constituting the message while maintaining latency on the link.
[0067] Figure 10A and Figure 10B Examples of systems according to embodiments of the present disclosure are shown. Specifically, Figure 10A and Figure 10B The block diagrams below show systems 5a and 5b, each comprising multiple CPUs. Figure 10A and Figure 10B In the description, repeated descriptions can be omitted.
[0068] Figure 10A Examples of systems according to embodiments of the present disclosure are shown. (Refer to...) Figure 10A System 5a may include a first CPU 11a and a second CPU 21a, and may include a first Double Data Rate (DDR) memory 12a and a second Double Data Rate (DDR) memory 22a respectively connected to the first CPU 11a and the second CPU 21a. The first CPU 11a and the second CPU 21a may be interconnected via an interconnect system 30a based on processor interconnect technology. Figure 10A As shown, interconnect system 30a can provide at least one CPU-to-CPU coherent link.
[0069] System 5a may include a first I / O device 13a and a first accelerator 14a communicating with a first CPU 11a, and may include a first device memory 15a connected to the first accelerator 14a. The first CPU 11a may communicate with the first I / O device 13a via bus 16a, and may communicate with the first accelerator 14a via bus 17a. Furthermore, system 5a may include a second I / O device 23a and a second accelerator 24a communicating with a second CPU 21a, and may include a second device memory 25a connected to the second accelerator 24a. The second CPU 21a may communicate with the second I / O device 23a via bus 26a, and may communicate with the second accelerator 24a via bus 27a.
[0070] The first CPU 11a, the first accelerator 14a, the second CPU 21a, and the second accelerator 24a can support the extended messages described above with reference to the foregoing figures. Therefore, while maintaining the latency of buses 17a and 27a and the interconnect system 30a, it is possible to increase the number of fields included in a single message.
[0071] Figure 10B Examples of systems according to embodiments of this disclosure are shown. Repeated descriptions may be omitted. See also... Figure 10B Similar to Figure 10ASystem 5a and system 5b may include a first CPU 11b and a second CPU 21b, a first DDR memory 12b and a second DDR memory 22b, a first I / O device 13b and a second I / O device 23b, and a first accelerator 14b and a second accelerator 24b, and may also include remote far memory 40. The first CPU 11b and the second CPU 21b may communicate with each other via interconnect system 30b. The first CPU 11b may be connected to the first I / O device 13b via bus 16b and to the first accelerator 14b via bus 17b. The second CPU 21b may be connected to the second I / O device 23b via bus 26b and to the second accelerator 24b via bus 27b.
[0072] The first CPU 11b and the second CPU 21b can be connected to the remote memory 40 via the first bus 18 and the second bus 28, respectively. The remote memory 40 can be used for memory expansion in system 5b, and the first bus 18 and the second bus 28 can be used as memory expansion ports. (Refer to the above...) Figure 10A As described above, the first CPU 11b, the first accelerator 14b, the second CPU 21b, and the second accelerator 24b can support the extended messages described with reference to the foregoing figures. Therefore, while maintaining the latency of buses 17b and 27b and the interconnect system 30b, it is possible to increase the number of fields included in a single message.
[0073] Figure 11 A data center is shown that employs an embodiment of the system according to this disclosure.
[0074] Reference Figure 11 Data center 3000 can refer to a facility for storing and providing services for various types of data, and may also be referred to as a data storage center. Data center 3000 can be a system for operating search engines and databases, and can be a computing system used in a company (such as a bank) or government agency. Data center 3000 may include application servers 3100 to 3100n and storage servers 3200 to 3200m. The number of application servers 3100 to 3100n and the number of storage servers 3200 to 3200m may be selected differently according to embodiments, and the number of application servers 3100 to 3100n may differ from the number of storage servers 3200 to 3200m.
[0075] Application server 3100 or storage server 3200 may respectively include at least one of processors 3110 and 3210 and at least one of memories 3120 and 3220. When storage server 3200 is described as an example, processor 3210 may control all operations of storage server 3200 and may access memory 3220 to execute instructions and / or data loaded into memory 3220. Memory 3220 may include double data rate synchronous DRAM (DDR SDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), dual in-line memory module (DIMM), Optane DIMM, or non-volatile DIMM (NVDIMM). According to embodiments, the corresponding number of processors 3210 and memory 3220 included in storage server 3200 may be selected differently. In some embodiments, processors 3110 and 3210 may be... Figure 1 The host processor 120 or device 110 is implemented. In some embodiments, Figure 1 Both the host processor 120 and the device 110 may be included in the application server 3100 and / or the storage server 3200.
[0076] In one embodiment, processor 3210 and memory 3220 may provide a processor-memory pair. In one embodiment, the number of processors 3210 may differ from the number of memories 3220. Processor 3210 may include a single-core processor or a multi-core processor. The above description of storage server 3200 can also be applied similarly to application server 3100. According to an embodiment, application server 3100 does not need to include storage device 3150. Storage server 3200 may include at least one storage device 3250. The number of storage devices 3250 included in storage server 3200 may be selected differently depending on the embodiment.
[0077] Application servers 3100 to 3100n and storage servers 3200 to 3200m can communicate with each other via network 3300. Network 3300 can be implemented using Fibre Channel (FC), Ethernet, etc. Here, FC is a medium used for relatively high-speed data transmission, and optical switches providing high performance / high availability can be used. Depending on the access method of network 3300, storage servers 3200 to 3200m can be provided as file storage devices, block storage devices, or object storage devices.
[0078] In one embodiment, network 3300 may include a storage-specific network (such as a storage area network (SAN)). For example, the SAN may include an FC-SAN that uses an FC network and is implemented according to the FC protocol (FCP). As another example, the SAN may include an Internet Protocol (IP)-SAN that uses a Transmission Control Protocol (TCP / IP) network and is implemented according to the Internet Small Computer System Interface (iSCSI or SCSI over TCP / IP) protocol. In another embodiment, network 3300 may include a general-purpose network (such as a TCP / IP network). For example, network 3300 may be implemented according to protocols such as FC over Ethernet (FCoE), Network Attached Storage (NAS), or NVMe over a network architecture (NVMe-oF).
[0079] The following text will primarily describe application server 3100 and storage server 3200. The description of application server 3100 can also be applied to another application server 3100n, and the description of storage server 3200 can also be applied to another storage server 3200m.
[0080] Application server 3100 can store data requested by users, clients, and / or applications in one of storage servers 3200 to 3200m via network 3300. Furthermore, application server 3100 can retrieve data requested by users, clients, and / or applications from one of storage servers 3200 to 3200m via network 3300. For example, application server 3100 can be implemented as a web server, database management system (DBMS), etc.
[0081] Application server 3100 can access memory 3120n or storage device 3150n included in another application server 3100n via network 3300, or can access memory 3220 to 3220m or storage device 3250 to 3250m included in storage servers 3200 to 3200m respectively via network 3300. Therefore, application server 3100 can perform various operations on data stored in application servers 3100 to 3100n and / or storage servers 3200 to 3200m. For example, application server 3100 can execute instructions for moving or copying data between application servers 3100 to 3100n and / or storage servers 3200 to 3200m. Here, data can be moved directly or via storage devices 3250 to 3250m of storage servers 3200 to 3200m to storage devices 3220 to 3220m of application servers 3100 to 3100n. Data moved via network 3300 may be encrypted for security or privacy purposes.
[0082] When the storage server 3200 is described as an example, the interface (I / F) 3254 provides a physical connection between the processor 3210 and the controller 3251, and a physical connection between the network interface controller (NIC) 3240 and the controller 3251. For example, the interface 3254 may be implemented as a Direct Attached Storage (DAS) interface, in which the storage device 3250 is directly connected via a dedicated cable. Furthermore, the interface 3254 may be implemented as various interface types, such as Advanced Technology Attachment (ATA), Serial ATA (SATA), External SATA (e-SATA), Small Computer Small Interface (SCSI), Serial Attached SCSI (SAS), Peripheral Component Interconnect (PCI), PCI Fast (PCIe), NVM Fast (NVMe), IEEE 1394, Universal Serial Bus (USB), Secure Digital (SD) card, Multimedia Card (MMC), Embedded Multimedia Card (eMMC), Universal Flash (UFS), Embedded Universal Flash (eUFS), and / or Compact Flash (CF) card interface.
[0083] Storage server 3200 may further include switch 3230 and NIC 3240. Under the control of processor 3210, switch 3230 may selectively connect processor 3210 to storage device 3250 or selectively connect NIC 3240 to storage device 3250. Application server 3100 may further include switch 3130 and NIC 3140. Under the control of processor 3110, switch 3130 may selectively connect processor 3110 to storage device 3150 or selectively connect NIC 3140 to storage device 3150.
[0084] In one embodiment, NIC 3240 may include a network interface card, network adapter, etc. NIC 3240 can connect to network 3300 via a wired interface, wireless interface, Bluetooth interface, optical interface, etc. NIC 3240 may include internal memory, digital signal processor (DSP), host bus interface, etc., and can connect to processor 3210 and / or switch 3230. The host bus interface may be implemented using one of the examples of interface 3254 described above. In one embodiment, NIC 3240 may be integrated with at least one of processor 3210, switch 3230, and storage device 3250.
[0085] In storage servers 3200 to 3200m or application servers 3100 to 3100n, the processor can send commands to storage devices 3150 to 3150n or 3250 to 3250m or memory 3120 to 3120n or 3220 to 3220m, thus programming data to or reading data from storage devices 3150 to 3150n or 3250 to 3250m or memory 3120 to 3120n or 3220 to 3220m. Here, the data can be data corrected by an error correction code (ECC) engine. The data can be data that has undergone data bus inversion (DBI) or data masking (DM), and the data may include CRC information. The data can be data encrypted for security or privacy.
[0086] Storage devices 3150 to 3150n or 3250 to 3250m can send control signals and command / address signals to NAND flash memory devices 3252 to 3252m in response to a read command received from the processor. Therefore, when reading data from NAND flash memory devices 3252 to 3252m, the read enable (RE) signal can be input as a data output control signal, thus enabling data to be output to the data queue (DQ) bus. A data strobe (DQS) can be generated using the RE signal. Command and address signals can be latched onto the page buffer based on the rising or falling edge of the write enable (WE) signal.
[0087] Controller 3251 provides overall control over the operation of storage device 3250. In one embodiment, controller 3251 may include static random access memory (SRAM). Controller 3251 may write data to NAND flash memory device 3252 in response to a write command, or read data from NAND flash memory device 3252 in response to a read command. For example, write and / or read commands may be provided by processor 3210 in storage server 3200, processor 3210m in another storage server 3200m, or processor 3110 or 3110n in application server 3100 or 3100n. DRAM 3253 may temporarily store or buffer data to be written to or already read from NAND flash memory device 3252. Furthermore, DRAM 3253 may store metadata. Here, metadata is generated by controller 3251 for managing user data or data in NAND flash memory device 3252. Storage device 3250 may include a security element (SE) for security or privacy purposes.
[0088] Although the inventive concept has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the appended claims.
Claims
1. An interface circuit, comprising: A packet transmitter is configured to generate a plurality of transport packets based on a request output from a core circuit, and to output the plurality of transport packets, each of the plurality of transport packets including information indicating packets to be merged; as well as The packet receiver is configured to generate a merged packet by merging multiple extended packets from a plurality of received packets received from outside the interface circuitry, each of the multiple extended packets including information indicating the packets to be merged. Wherein, the sum of the lengths of the multiple fields used to process the request is greater than the length of each of the multiple transmission packets, and The packet sender is further configured to encode the plurality of transport packets in different formats, such that the plurality of fields for the request are distributed across the plurality of transport packets.
2. The interface circuit according to claim 1, wherein, The multiple extension packs are of equal length to each other, and The length of the merged package is greater than the length of each of the multiple expansion packages.
3. The interface circuit according to claim 1, wherein, The packet transmitter is also configured to include order information indicating the order in which the fields of the plurality of transmission packets are arranged.
4. The interface circuit according to claim 3, wherein, The packet receiver is also configured to generate a merged packet by arranging fields included in the plurality of extended packets based on order information included in the plurality of extended packets.
5. The interface circuit according to claim 1, wherein, The packet receiver includes a buffer that temporarily stores the identified plurality of extended packets, and The packet receiver is also configured to generate a merged packet by merging the multiple extended packets stored in the buffer.
6. The interface circuit according to claim 1, wherein, The packet transmitter also includes registers that store information about whether an external device is capable of recognizing information indicating that packets will be merged. The packet transmitter is also configured to: refer to a register to provide the plurality of transport packets to an external device capable of identifying information indicating that packets to be merged.
7. The interface circuit according to claim 1, wherein, Each of the plurality of transport packets includes an opcode field, and information regarding the operation corresponding to the request is allowed to be written into the opcode field. The packet sender is also configured to write at least one of the information about the operation and information indicating the packets to be merged into the opcode field.
8. The interface circuit according to any one of claims 1 to 7, wherein, The packet transmitter is also configured to generate the plurality of transport packets such that a field specifically used to record information indicating packets to be merged is included in each of the plurality of transport packets.
9. A processor connected to an external device via a bus, the processor comprising: The controller is configured to generate messages that will be transmitted to external devices. as well as The interface circuitry is configured to generate packets that will be output to the bus based on message generation. The interface circuit is configured as follows: Generate multiple packets, each packet including multiple fields contained in the message; and Output the multiple packets to the bus. Wherein, the sum of the lengths of the plurality of fields is greater than the length of each of the plurality of packets, wherein the formats of the plurality of packets are different from each other, and the plurality of fields are distributed among the plurality of packets.
10. The processor according to claim 9, wherein, The interface circuitry is also configured to generate the plurality of packets such that each of the plurality of packets includes information indicating the packets to be merged.
11. The processor according to claim 10, wherein, The interface circuit is also configured as follows: Identify extended packets among multiple received packets from the bus; the extended packets include information indicating that the packets will be merged; and Messages are generated by merging extension packages.
12. The processor according to claim 11, wherein, The interface circuitry is also configured to generate a message by arranging multiple fields included in the extension package such that the multiple fields do not overlap with each other.
13. The processor according to claim 12, wherein, Each expansion pack includes order information that indicates the order in which the multiple fields of the expansion pack are arranged when merging expansion packs, and The interface circuitry is also configured to arrange the multiple fields included in the extension package based on sequence information.
14. A method for processing packets, the method comprising: Receive multiple packets; Identify multiple extended packets among the plurality of received packets, each of the plurality of extended packets including information indicating the packets to be merged; A merged package is generated by arranging multiple fields included in the multiple extension packages; as well as Output merged package Wherein, the sum of the lengths of the plurality of fields is greater than the length of each of the plurality of received packets. The multiple received packets have different formats, and the multiple fields are distributed across the multiple received packets.
15. The method of claim 14, further comprising: Multiple transport packets are generated in different formats, such that multiple fields included in the message used to process the request output from the core are distributed across the multiple transport packets; as well as Output the multiple transmission packets.
16. The method according to claim 15, wherein, The step of generating the plurality of transport packets includes adding information indicating the packets to be merged to each of the plurality of transport packets.
17. The method according to claim 16, wherein, The step of generating the plurality of transport packets further includes: adding order information to the plurality of transport packets, wherein the order information indicates the order in which the fields of the plurality of transport packets are arranged when merging the plurality of transport packets.
18. The method according to any one of claims 14 to 17, further comprising: The identified multiple extension packages are temporarily stored in a buffer. After receiving the extended package containing the last arranged field among the multiple fields, the step of generating a merge package is performed.
19. The method according to any one of claims 14 to 17, wherein, The multiple extension packs are of equal length to each other, and The length of the merged package is greater than the length of each of the multiple expansion packages.
Citation Information
Patent Citations
Communication system
US20070286077A1
Data processing systems and methods
US4445171A