Multi-core processor

By adopting internal data transmission circuits and polling mechanisms in multi-core processors, the data transmission delay problem caused by external bus and external memory is solved, and the rapid data transmission between the processor cores is realized, and the effectiveness of network packet forwarding is improved.

CN120144525APending Publication Date: 2025-06-13AIROHA TECH (SUZHOU) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311699347.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the data transmission between different processor cores, existing multi-core processors are limited by the high access delay of external bus and external memory, which affects the processing efficiency of network packet forwarding.

Method used

The internal data transmission circuit and polling mechanism are adopted to realize fast data transmission between multiple processor cores, avoiding the passage of external bus and external memory.

Benefits of technology

It greatly improves the data transmission efficiency between processor cores and improves the processing efficiency of network packet forwarding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144525A_ABST
    Figure CN120144525A_ABST
Patent Text Reader

Abstract

The invention provides a multi-core processor. The multi-core processor comprises a plurality of processor cores and a data transmission circuit, the plurality of processor cores includes a first processor core and a second processor core. The first processor core is provided with a first buffer and is used for writing first data into the first buffer. The second processor core has a second buffer. The data transmission circuit is used for polling whether the first buffer has data waiting for transmission and transmitting the first data from the first buffer to the second buffer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a processor, and more particularly to a multi-core processor that uses an internal data transfer circuit and a polling mechanism to achieve fast data transfer between multiple processor cores. Background Art

[0002] A network processing unit (NPU) is a processor specifically applied to network packet processing, and particularly has some characteristics and architectures to accelerate the processing efficiency of network packets. When the network processor has multiple processor cores, the processing efficiency of network packets can be further improved through parallel processing. For example, the characteristics of multiple processor cores can be used for pipeline parallel processing to improve the forwarding efficiency of network packets. However, parallel processing using multiple processor cores requires data transfer between different processor cores. Generally speaking, limited by power consumption and cost requirements, a network processor generally needs to use an external bus to access an external memory to achieve data transfer between different processor cores in the same multi-core processor. However, the external memory itself has a high access latency, that is, the access efficiency of the external memory is low. Accessing the external memory through the external bus to complete data transfer between different processor cores will directly affect the overall packet forwarding processing efficiency. Summary of the Invention

[0003] One object of the present invention is to provide a multi-core processor that uses an internal data transfer circuit and a polling mechanism to achieve fast data transfer between multiple processor cores.

[0004] In an embodiment of the present invention, a multi-core processor is disclosed. The multi-core processor includes multiple processor cores and a data transfer circuit. The multiple processor cores include a first processor core and a second processor core. The first processor core has a first buffer, and the first processor core is used to write a first data into the first buffer. The second processor core has a second buffer. The data transfer circuit is used to poll whether there is data in the first buffer waiting to be transferred, and transfer the first data from the first buffer to the second buffer.

[0005] In another embodiment of the present invention, a multi-core processor is disclosed. The multi-core processor includes a plurality of processor cores and a data transmission circuit. The plurality of processor cores includes a first processor core and a second processor core. The first processor core has a first buffer. The second processor core has a second buffer, wherein the second processor core is used to poll whether there is data waiting to be read in the second buffer and read a data from the second buffer. The data transmission circuit is used to transmit the data from the first buffer to the second buffer.

[0006] With the assistance of the internal data transmission circuit and the polling mechanism, the data transmission between multiple processor cores in the multi-core processor does not need to pass through an external bus and an external memory. In this way, the data transmission efficiency between processor cores can be greatly improved. When the multi-core processor adopts a pipeline parallel computing architecture to process network packet forwarding tasks, the performance of network packet forwarding can be greatly improved. Brief Description of the Drawings

[0007] Figure 1 It is a schematic diagram of a multi-core processor according to an embodiment of the present invention.

[0008] Figure 2 It is a schematic diagram of the operation of data transmission from one processor core to another processor core according to an embodiment of the present invention.

[0009] Figure 3 It is a flowchart of the data transmission circuit executing internal data transmission of a multi-core processor according to an embodiment of the present invention.

[0010] Figure 4 It is a schematic diagram of the operation of data transmission from one processor core to multiple processor cores according to an embodiment of the present invention.

[0011]

Symbol Description

[0012] 100: Multi-core processor

[0013] 102_1, 102_2, 102_3, 102_N: Processor core

[0014] 104_1, 104_2, 104_3, 104_N, 106_1, 106_2, 106_3, 106_N: Buffer

[0015] 108: Data transmission circuit

[0016] 202, 402: Header

[0017] 204, 404: Payload

[0018] PS_1, PS_2, PS_3: Pipeline stage

[0019] D_TX, D_TX’: Transmitted data

[0020] D_C2, D_C3: Processed data

[0021] B0, B1: Bytes

[0022] TP, LEN, DST, SRC: Fields

[0023] fifo_write(): Write instruction

[0024] fifo_read(): Read instruction

[0025] fifo_empty(): Polling instruction

[0026] S302, S304, S306, S308: Steps Detailed implementation manners

[0027] In the specification and claims, certain terms are used to refer to specific elements. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same element. The specification and claims do not use the difference in names as a way to distinguish elements, but use the difference in functions of elements as the criterion for distinction. The terms "comprise" and "comprising" mentioned throughout the specification and claims are open-ended terms, and should be interpreted as "including but not limited to". In addition, the term "coupled" or "coupling" herein includes any direct and indirect electrical connection means. Therefore, if a first device is described as being coupled to a second device in the text, it means that the first device can be directly electrically connected to the second device, or indirectly electrically connected to the second device through other devices and connection means.

[0028] Figure 1 It is a schematic diagram of a multi-core processor according to an embodiment of the present invention. For example, the multi-core processor 100 may be a multi-core network processor without a shared cache, and may be applied to network devices such as gateways. The multi-core processor 100 includes multiple processor cores 102_1, 102_2,..., 102_N (N≥2) and a data transmission circuit 108. The processor cores 102_1 to 102_N may have the same or similar processor architectures, such as Figure 1As shown, inside the processor core 102_1, there are multiple buffers 104_1, 106_1; inside the processor core 102_2, there are multiple buffers 104_2, 106_2; and inside the processor core 102_N, there are multiple buffers 104_N, 106_N. In one example, the buffers 104_1 to 104_N, 106_1 to 106_N can be implemented by the internal static random access memory (SRAM) of the processor cores 102_1 to 102_N. In another example, the buffers 104_1 to 104_N, 106_1 to 106_N can be first-in first-out (FIFO) buffers. Note that Figure 1 only the elements related to the present invention are shown. In fact, the multi-core processor 100 may include other elements to implement the specified functions.

[0029] In this embodiment, buffers 104_1 to 104_N serve as transmit (TX) buffers, and buffers 106_1 to 106_N serve as receive (RX) buffers. In other words, the processor core 102_i (1 ≤ i ≤ N) can write the data to be transmitted to another processor core 102_j (1 ≤ j ≤ N, j ≠ i) into its own transmit buffer (i.e., buffer 104_i), and the processor core 102_j can read the data from processor core 102_i from its own receive buffer (i.e., buffer 106_j). In addition, the processor core 102_i (1 ≤ i ≤ N) can also write another data to be transmitted to another processor core 102_k (1 ≤ k ≤ N, k ≠ j, k ≠ i) into its own transmit buffer (i.e., buffer 104_i), and the processor core 102_k can read the data from processor core 102_i from its own receive buffer (i.e., buffer 106_k). In this embodiment, the data transmission circuit 108 is designed to be responsible for data transmission between the processor cores 102_1 to 102_N. For example, the data transmission circuit 108 will transmit the data in buffer 104_i to buffer 106_j, and transmit another data in the same buffer 104_i to buffer 106_k. Since the data transmission circuit 108 is an internal circuit of the multi-core processor 100, the data transmission between the processor cores 102_1 to 102_N does not need to pass through an external bus and an external memory. In other words, the data transmission between buffer 104_i and buffer 106_j is only carried out inside the multi-core processor 100, and the data transmission between buffer 104_i and buffer 106_k is only carried out inside the multi-core processor 100. In this way, the data transmission efficiency between the processor cores 102_1 to 102_N can be greatly improved. In addition, different from the traditional multi-level cache hierarchy, the present invention realizes fast data transmission between different processor cores through the internal buffer of the processor core and the polling mechanism. When the multi-core processor 100 adopts a pipeline parallel computing architecture (i.e., the processor cores 102_1 to 102_N are different pipeline stages of pipeline processing respectively) to process the network packet forwarding work, the performance of network packet forwarding can be greatly improved.For example, it takes about 14 clock cycles for a traditional multi-core processor to perform one read operation on an external static random access memory, and it takes about 7 clock cycles for a traditional multi-core processor to perform one write operation on an external static random access memory. However, for the multi-core processor 100 of the present invention, it only takes 1 clock cycle to perform one read operation on the internal buffers 104_1 to 104_N and 106_1 to 106_N, and it only takes 1 clock cycle for the multi-core processor 100 of the present invention to perform one write operation on the internal buffers 104_1 to 104_N and 106_1 to 106_N.

[0030] Figure 2Schematic diagram of the operation of data transfer from one processor core to another in an embodiment of the present invention. Assume that processor cores 102_1 and 102_2 are respectively responsible for different pipeline stages PS_1 and PS_2 in pipeline processing, and processor core 102_1 needs to transfer the processed data D_C2 of pipeline stage PS_1 to pipeline stage PS_2 for further processing. In this embodiment, processor core 102_1 generates transfer data D_TX, where the transfer data D_TX includes a header 202 and a payload 204, and the processed data D_C2 to be transferred to processor core 102_2 is carried in the payload 204. Processor core 102_1 writes the transfer data D_TX into buffer (such as a first-in-first-out buffer) 104_1 through the write instruction fifo_write(). The header 202 includes multiple fields for recording information about the transfer operation. For example, the header 202 may include a DST field, an SRC field, a TP field, and a LEN field. The DST field is used to indicate the destination processor core of the transfer data D_TX (i.e., to which processor core the transfer data D_TX is to be transferred. In this embodiment, the DST field is set to Core 2 to indicate that the transfer data D_TX is to be transferred to processor core 102_2), the SRC field is used to indicate the source processor core of the transfer data D_TX (i.e., from which processor core the transfer data D_TX is generated. In this embodiment, the SRC field is set to Core 1 to indicate that the transfer data D_TX is sent by processor core 102_1), the LEN field is used to indicate the data length of the payload 204 immediately following the header 202 (i.e., the amount of the processed data D_C2 to be transferred to processor core 102_2), and the TP field is used to indicate that the data carried in the payload 204 is for a specified function (i.e., by which function the processed data D_C2 is to be processed. In this embodiment, the TP field indicates a function in pipeline stage PS_2). Note that the above fields are only for illustrative purposes and not for limiting the present invention. That is, the fields included in the header 202 can be set according to actual needs.

[0031] As described above, the data transfer circuit 108 is responsible for data transfer between processor cores 102_1 to 102_N. Please refer to Figure 2 and Figure 3 , Figure 3 Flowchart of the data transfer circuit 108 in an embodiment of the present invention for performing internal data transfer of a multi-core processor. If the same result can be obtained generally, the steps do not necessarily have to be exactly in accordance with Figure 3Execute in the order shown. In step S302, the data transfer circuit 108 executes the polling instruction fifo_empty() to poll whether there is data waiting to be transmitted in the buffer of the processor core. In this embodiment, after the processor core 102_1 writes the transfer data D_TX into the buffer 104_1 through the write instruction fifo_write(), the polling instruction fifo_empty() of the data transfer circuit 108 gets a response that there is data to be transmitted in the buffer 104_1. Therefore, the data transfer circuit 108 then executes the read instruction fifo_read() to read the header 202 of the transfer data D_TX and parse the fields in the header 202, such as the LEN field and the DST field (step S304). In step S306, the data transfer circuit 108 reads the data carried by the payload 204 (such as two bytes B0, B1) immediately following the header 202 according to the data length indicated by the LEN field (for example, LEN = 2). In step S308, the data transfer circuit 108 executes the write instruction fifo_write() according to the destination processor core indicated by the DST field (for example, DST = Core 2) to transfer the transfer data D_TX (including the header 202 and the payload 204) to the buffer (such as a first-in-first-out buffer) 106_2 in the processor core 102_2.

[0032] The buffer 106_2 serves as the receive buffer of the processor core 102_2. As shown in Figure 2, the processor core 102_2 executes the polling instruction fifo_empty() to poll whether there is data waiting to be processed in its own buffer 106_2. In this embodiment, after the data transfer circuit 108 writes the transfer data D_TX into the buffer 106_2 through the write instruction fifo_write(), the polling instruction fifo_empty() of the processor core 102_2 gets a response that there is data to be processed in the buffer 106_2. Therefore, the processor core 102_2 then executes the read instruction fifo_read() to read the transfer data D_TX (including the header 202 and the payload 204) in the buffer 106_2 and perform relevant processing of the pipeline stage PS_2 on the data carried by the payload 204 (that is, the processed data D_C2 of the pipeline stage PS_1) based on the fields (such as the TP field) in the header 202.

[0033] The buffers 104_1 to 104_N serve as transmission buffers. In this embodiment, the same transmission buffer (such as a first-in-first-out buffer) can be reused to sequentially store the data to be transmitted to different processor cores. In other words, each processor core does not need to configure multiple transmission buffers for different processor cores, so the configuration and maintenance of the buffers can be simplified.Figure 4 This is an operation schematic diagram of a data transfer from one processor core to multiple processor cores according to an embodiment of the present invention. Assume that processor cores 102_1, 102_2, and 102_3 are respectively responsible for different pipeline stages PS_1, PS_2, and PS_3 in pipeline processing, and processor core 102_1 needs to transfer the processed data D_C2 (including two bytes B0, B1) of pipeline stage PS_1 to pipeline stage PS_2 for further processing, and transfer the processed data D_C3 (including one byte B0) of pipeline stage PS_1 to pipeline stage PS_3 for further processing. In this embodiment, processor core 102_1 generates transfer data D_TX and transfer data D_TX', where transfer data D_TX includes header 202 and payload 204, and transfer data D_TX' includes header 402 and payload 404. In addition, the processed data D_C2 to be transferred to processor core 102_2 is carried in payload 204, and the processed data D_C3 to be transferred to processor core 102_3 is carried in payload 404. Processor core 102_1 writes the transfer data D_TX and D_TX' into buffer 104_1 in sequence through the write instruction fifo_write(). Then, data transfer circuit 108 will Figure 3 process the transfer data D_TX and transfer data D_TX' stored in sequence in the same buffer 104_1 according to the process shown, that is, data transfer circuit 108 first transfers the transfer data D_TX from the buffer (such as a first-in-first-out buffer) 104_1 of processor core 102_1 to the buffer (such as a first-in-first-out buffer) 106_2 of processor core 102_2. Then, data transfer circuit 108 transfers the transfer data D_TX' from the buffer (such as a first-in-first-out buffer) 104_1 of processor core 102_1 to the buffer (such as a first-in-first-out buffer) 106_3 of processor core 102_3.

[0034] Buffer 106_2 serves as the receive buffer for processor core 102_2. As shown in Figure 4, processor core 102_2 executes the polling instruction fifo_empty() to poll whether buffer 106_2 has data waiting to be processed. In this embodiment, after data transfer circuit 108 writes the transfer data D_TX to buffer 106_2 via the write instruction fifo_write(), the polling instruction fifo_empty() of processor core 102_2 receives a response indicating that there is data to be processed in buffer 106_2. Therefore, processor core 102_2 then executes the read instruction fifo_read() to read the transfer data D_TX (including header 202 and payload 204) in buffer 106_2, and performs relevant processing for pipeline stage PS_2 on the data carried by payload 204 (i.e., the processed data D_C2 of pipeline stage PS_1) based on the fields in header 202 (such as the TP field).

[0035] Buffer 106_3 serves as the receive buffer for processor core 102_3. As shown in Figure 4, processor core 102_3 executes the polling instruction fifo_empty() to poll whether buffer 106_3 has data waiting to be processed. In this embodiment, after data transfer circuit 108 writes the transfer data D_TX' to buffer 106_3 via the write instruction fifo_write(), the polling instruction fifo_empty() of processor core 102_3 receives a response indicating that there is data to be processed in buffer 106_3. Therefore, processor core 102_3 then executes the read instruction fifo_read() to read the transfer data D_TX' (including header 402 and payload 404) in buffer 106_3, and performs relevant processing for pipeline stage PS_3 on the data carried by payload 404 (i.e., the processed data D_C3 of pipeline stage PS_1) based on the fields in header 402 (such as the TP field).

[0036] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the claims of the present invention shall fall within the scope of the present invention.

Claims

1. A multi-core processor, comprising: A plurality of processor cores, comprising: A first processor core having a first buffer, wherein the first processor core is configured to write first data into the first buffer ; And A second processor core having a second buffer; And A data transfer circuit configured to poll whether there is data in the first buffer waiting to be transferred and transfer the first data from the first buffer to the second buffer.

2. The multi-core processor according to claim 1, wherein the plurality of processor cores further comprises: A third processor core having a third buffer; The first processor core is further configured to write second data into the first buffer; and the data transfer circuit is further configured to transfer the second data from the first buffer to the third buffer.

3. The multi-core processor according to claim 1, wherein the first data comprises a header and a payload.

4. The multi-core processor according to claim 3, wherein the header indicates that the destination processor core of the first data is the second processor core.

5. The multi-core processor according to claim 3, wherein the header indicates the data length of the payload.

6. The multi-core processor according to claim 3, wherein the header indicates that the payload is to be processed by a specified function.

7. The multi-core processor according to claim 3, wherein the header indicates that the source processor core of the first data is the first processor core.

8. The multi-core processor according to claim 1, wherein the first and second processor cores are respectively different pipeline stages of pipeline processing.

9. The multi-core processor according to claim 1, wherein the multi-core processor is a multi-core network processor.

10. A multi-core processor, comprising: A plurality of processor cores, comprising: A first processor core having a first buffer; and A second processor core having a second buffer, wherein the second processor core is configured to poll whether there is data in the second buffer waiting to be read and read data from the second buffer; and A data transfer circuit configured to transfer the data from the first buffer to the second buffer.

11. The multi-core processor according to claim 10, wherein the data transfer circuit is further configured to poll whether there is data in the first buffer waiting to be transferred.

12. The multi-core processor according to claim 10, wherein the data comprises a header and a payload.

13. The multi-core processor according to claim 12, wherein the header indicates that the destination processor core of the data is the second processor core.

14. The multi-core processor according to claim 12, wherein the header indicates the data length of the payload.

15. The multi-core processor according to claim 12, wherein the header indicates that the payload is to be processed by a specified function.

16. The multi-core processor according to claim 12, wherein the header indicates that the source processor core of the data is the first processor core.

17. The multi-core processor according to claim 10, wherein the first and second processor cores are respectively different pipeline stages of pipeline processing.

18. The multi-core processor according to claim 10, wherein the multi-core processor is a multi-core network processor.