A packet processing method, device, medium, and product based on network-on-chip

By setting up caches on the slave node and performing status query processing, the problem of increasing complexity and load in maintaining cache consistency by master nodes is solved, and packet processing is efficiently processed and processor temperature is reduced.

CN119557264BActive Publication Date: 2025-05-30SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510059226.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

In an on-chip network, the master node needs to maintain cache consistency between multiple processor cores. Adding a shared cache will increase the functional complexity and load of the master node, causing the processor to be locally overheated, affecting performance and reliability.

Method used

The first cache is set on the slave node to store the data that needs to be written into memory and the corresponding address. Through cache status query and address status query, it is determined to read or write data from the cache or memory, and generate a response message without setting a shared cache in the master node.

Benefits of technology

It improves the efficiency of processing packets in the on-chip network, reduces the load on the master node, avoids the problem of excessive local temperature of the processor, and improves the operating efficiency and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557264B_ABST
    Figure CN119557264B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method and device, medium, and product for message processing based on a network-on-chip, including: receiving a request message for a processor core or a master node to access memory, where the request message includes a target address for accessing memory, querying the address status according to the target address to obtain an address status query result, and in the case where the address status query result indicates that the target address is not occupied, querying the cache status according to the target address to obtain a cache status query result in a first cache, and reading data from the first cache or memory, or writing data to the first cache or memory according to the cache status query result and the type of the request message, and generating a response message for the request message. Through the embodiment of the present invention, the efficiency of processing messages in the network-on-chip and the overall operation efficiency of the network-on-chip are improved, the load on the master node can be reduced, and the local temperature of the processor during the operation of the network-on-chip can be prevented from being too high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of processor technology, and in particular to a message processing method and device, medium and product based on an on-chip network. Background Art

[0002] System-on-Chip is a communication architecture applied to processors. It consists of basic units such as processor cores, home nodes (HN), slave nodes (SN), and memory. Each processor core can act as a request node (RN) to send a request message to the home node to request the home node to process the request transaction. When the request transaction involves a processing flow that requires access to memory, the home node interacts with the memory through the slave node and generates a corresponding response message to the request node.

[0003] In the related art, a method of adding a shared cache to the master node is adopted to improve the efficiency of processing request messages in the on-chip network. However, the master node needs to be responsible for maintaining the cache consistency between multiple processor cores in the process of processing request transactions. Adding a shared cache will greatly increase the functional complexity of the master node, and the load of the master node will be further increased, which will lead to the problem of local overheating of the processor during the operation of the on-chip network, affecting the performance and reliability of the processor. Summary of the invention

[0004] In view of the above problems, a method and device, medium, and product for processing messages based on a network on chip are proposed to overcome the above problems or at least partially solve the above problems, including:

[0005] A message processing method based on a network on chip, the network on chip comprising a memory, a plurality of processor cores, at least one master node and at least one slave node, the slave node being provided with a first cache, the method comprising:

[0006] receiving a request message from the processor core or the master node to access the memory; wherein the request message includes a target address of the memory to be accessed;

[0007] Performing an address status query according to the target address to obtain an address status query result;

[0008] When the address status query result indicates that the target address is not occupied, performing a cache status query according to the target address to obtain a cache status query result in the first cache;

[0009] Read data from the first cache or the memory, or write data to the first cache or the memory, according to the cache status query result and the type of the request message;

[0010] Generate a response message for the request message.

[0011] Optionally, the type of the request message includes a read data request message or a write data request message, and the first cache stores a plurality of first data to be written to the memory and a first address corresponding to each of the first data and the memory; wherein, when the request message is the write data request message, the request message further includes second data; the performing a cache status query according to the target address to obtain a cache status query result in the first cache includes:

[0012] If the target address exists among the plurality of first addresses, the cache status query result indicates a cache hit;

[0013] If the target address does not exist among the plurality of first addresses, the cache status query result indicates a cache miss.

[0014] Optionally, the reading data from the first cache or the memory, or writing data to the first cache or the memory, according to the cache status query result and the type of the request message, includes:

[0015] If the cache status query result indicates a cache hit and the type of the request message is a write data request message, determine a first target data corresponding to the target address in the first cache, and update the first target data to the second data.

[0016] Optionally, the reading data from the first cache or the memory, or writing data to the first cache or the memory, according to the cache status query result and the type of the request message, further includes:

[0017] If the cache status query result indicates a cache hit and the type of the request message is a read data request message, determine a first target data corresponding to the target address in the first cache, and read the first target data;

[0018] The generating a response message for the request message includes:

[0019] Generate a response message for the request message according to the first target data.

[0020] Optionally, reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0021] If the cache status query result indicates cache miss and the type of the request message is a write data request message, determine whether the first data reaches the storage threshold in the first cache;

[0022] If the first data reaches the storage threshold in the first cache, determine second target data among the multiple pieces of first data; wherein, the second target data is the piece of first data that has not been used for the longest time;

[0023] Write the second target data to the memory, and delete the second target data from the first cache;

[0024] Write the second data to the first cache.

[0025] Optionally, after determining whether the first data reaches the storage threshold in the first cache, further includes:

[0026] If the first data does not reach the storage threshold in the first cache, write an empty cache line to the first cache according to the target address.

[0027] Optionally, reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0028] If the cache status query result indicates cache miss and the type of the request message is a read data request message, read third target data from the memory according to the target address;

[0029] Generating a response message for the request message includes:

[0030] Generate a response message for the request message according to the third target data.

[0031] Optionally, after receiving a request message for the processor core or the master node to access the memory, further includes:

[0032] In the case where the address status query result indicates that the target address is occupied, store the request message in a buffer stream.

[0033] Optionally, the input channels corresponding to the slave nodes include a request channel and a data channel. Both the request channel and the data channel are used to receive the request message, and the priority of receiving the request message through the data channel is higher than that of receiving the request message through the request channel.

[0034] Optionally, the network-on-chip further includes a memory control unit. The slave node is communicatively connected to the memory through the memory control unit, and the master node interacts with the memory through the slave node.

[0035] Optionally, the network-on-chip further includes a crossbar switch, which is communicatively connected to the master node, the slave node, and the multiple processor cores.

[0036] Optionally, each of the processor cores is provided with a second cache.

[0037] An electronic device includes a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the method for processing messages based on the network-on-chip as described above is implemented.

[0038] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for processing messages based on the network-on-chip as described above is implemented.

[0039] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the method for processing messages based on the network-on-chip as described above is implemented.

[0040] Embodiments of the present invention have the following advantages: By receiving a request message for a memory access from a processor core or a master node, where the request message includes a target address for accessing the memory, performing an address status query based on the target address to obtain an address status query result, and when the address status query result indicates that the target address is not occupied, performing a cache status query based on the target address to obtain a cache status query result in a first cache, reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, and generating a response message for the request message. Instead of setting up a shared cache in the master node, a first cache is set up on a slave node with a relatively lower functional complexity than the master node. This first cache is used to store first data (i.e., dirty data) to be written to the memory and a first address corresponding to the first data. When the target address is hit, the request transaction for the memory can be processed through the first target data without accessing the memory, which can not only improve the efficiency of processing messages in the on-chip network and the overall operation efficiency of the on-chip network, but also reduce the load of the master node and avoid the problem of excessive local temperature of the processor during the operation of the on-chip network. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0042] Figure 1 is a flowchart of the steps of a method for processing messages based on an on-chip network provided by an embodiment of the present invention;

[0043] Figure 2 is a schematic diagram of the architecture of an on-chip network provided by an embodiment of the present invention;

[0044] Figure 3 is a schematic diagram for comparing the processing process of a write data request message between the present invention and related technologies provided by an embodiment of the present invention;

[0045] Figure 4 is a schematic flowchart of performing an address status query and a cache status query when processing a write data request message provided by an embodiment of the present invention;

[0046] Figure 5 is a schematic flowchart of performing different processes according to the cache status query result when processing a write data request message provided by an embodiment of the present invention;

[0047] Figure 6It is a schematic diagram showing the process of comparing the present invention with the related art for processing read data request messages provided by an embodiment of the present invention;

[0048] Figure 7 It is a schematic flowchart showing the process when processing a read data request message provided by an embodiment of the present invention;

[0049] Figure 8 It is another schematic diagram showing the process of comparing the present invention with the related art for processing read data request messages provided by an embodiment of the present invention;

[0050] Figure 9 It is another schematic diagram showing the process of comparing the present invention with the related art for processing write data request messages provided by an embodiment of the present invention;

[0051] Figure 10 It is a schematic diagram of a message processing pipeline architecture established based on different clock cycles provided by an embodiment of the present invention;

[0052] Figure 11 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0053] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] Refer to Figure 1 , which shows a flowchart of the steps of a message processing method based on a network-on-chip provided by an embodiment of the present invention. The network-on-chip includes a memory, multiple processor cores, at least one master node, and at least one slave node. The slave node is provided with a first cache. The message processing method may specifically include the following steps:

[0055] Step 101, receive a request message for the processor core or the master node to access the memory; wherein, the request message includes a target address for accessing the memory;

[0056] In this embodiment, the processor can be a multi-core processor, and any core of the multi-core processor can send a request message. The request message carries the cache address corresponding to the cache data that the core needs to request for processing, that is, the target address. Through the request message, the core can request to perform transaction processing on the cache data at the target address. These transactions can include changing the cache state of the core for the cache data, reading the cache data, or writing the cache data to the memory, etc.

[0057] The master node (HN node) is responsible for receiving the request messages from the multi-core processor and then sending a request message to access the memory to the slave node.

[0058] The slave node (SN node) is provided with a first cache, which can be a private cache and does not share data with other cores or nodes. Therefore, it does not need to maintain cache coherence. When the slave node receives a request message from the core or the master node, it then interacts with the memory.

[0059] In some embodiments of the present invention, the on-chip network further includes a memory control unit. The slave node is communicatively connected to the memory through the memory control unit, and the master node interacts with the memory through the slave node.

[0060] The memory control unit is used to manage memory access. When the master node needs to access the memory, it sends a request message to the slave node, and the slave node interacts with the memory through the memory control unit.

[0061] In some embodiments of the present invention, the on-chip network further includes a crossbar, and the crossbar is communicatively connected to the master node, the slave node, and the multiple processor cores.

[0062] A crossbar is a network topology structure used to establish direct connections between multiple inputs and outputs. In an on-chip network, the crossbar is responsible for task scheduling within a cluster composed of multiple processing cores, ensuring that only one message is sent to the HN node per clock cycle to prevent task conflicts.

[0063] Exemplarily, as Figure 2 shown, the crossbar is communicatively connected to the master node, the slave node, and multiple processor cores (core 1, core 2, core 3, core 4) respectively, and the slave node is responsible for interacting with the memory (DDR).

[0064] In some embodiments of the present invention, each of the processor cores is provided with a second cache.

[0065] The second cache refers to the shared cache of each processing core, specifically, it can be a level 2 cache (Cache). When processing the request message and the corresponding transaction initiated by the core, it is necessary to maintain cache consistency among multiple cores.

[0066] Exemplarily, as Figure 2 shown, Core 1, Core 2, Core 3, and Core 4 are all provided with a second cache.

[0067] In some embodiments of the present invention, the input channels corresponding to the slave nodes include a request channel and a data channel. Both the request channel and the data channel are used to receive the request message, and the priority of receiving the request message through the data channel is higher than that of receiving the request message through the request channel.

[0068] In this embodiment, the input channels of the slave node include a request channel and a data channel. Both the request channel and the data channel are used to receive the request message. To ensure processing efficiency, the priority of receiving the request message through the data channel is higher than that of receiving the request message through the request channel.

[0069] Step 102, perform an address status query according to the target address to obtain an address status query result;

[0070] For the address status query, it is to query whether the target address is occupied by other request messages or transactions. If it is occupied, in order to avoid transaction conflicts, the request message needs to be suspended for processing; if it is not occupied, the request message can be processed in the next step.

[0071] Exemplarily, the address status of each cache address can be maintained through a pre-set address entry, and then the address status of the target address can be queried through this address entry.

[0072] Step 103, in the case where the address status query result indicates that the target address is not occupied, perform a cache status query according to the target address to obtain a cache status query result in the first cache;

[0073] The cache status query is to query whether there is cache data corresponding to the target address in the first cache of the slave node. According to different cache status query results, different processing is performed on the request message.

[0074] In some embodiments of the present invention, the type of the request message includes a read data request message or a write data request message. The first cache stores a plurality of first data that need to be written into the memory and a first address corresponding to each of the first data and the memory. Wherein, when the request message is the write data request message, the request message further includes second data. The caching status query according to the target address to obtain a caching status query result in the first cache includes:

[0075] If the target address exists among the plurality of first addresses, the caching status query result indicates a cache hit;

[0076] If the target address does not exist among the plurality of first addresses, the caching status query result indicates a cache miss.

[0077] In this embodiment, the first cache stores a plurality of first data that need to be written into the memory, that is, the first data is dirty data (data inconsistent with the data stored in the memory), and the first address, which can be the cache address corresponding to the first data and the memory.

[0078] A read data request message, that is, a request message for a transaction involving reading data from the memory; a write data request message, that is, a request message for a transaction involving writing data into the memory. The second data is the data that needs to be written into the memory.

[0079] Specifically, by matching the target address with the plurality of first addresses, if the match is successful, the caching status query result indicates a cache hit, indicating that data corresponding to the target address is stored in the first cache; if the match is successful, the caching status query result indicates a cache miss, indicating that data corresponding to the target address is not stored in the first cache.

[0080] In some embodiments of the present invention, after receiving the request message for the processor core or the master node to access the memory, it further includes:

[0081] When the address status query result indicates that the target address is occupied, store the request message in the buffer stream.

[0082] Specifically, if the address status query result indicates that the target address is occupied, it means that other transactions are occupying the target address. To avoid transaction conflicts, it is necessary to defer the processing of the request message. Therefore, the request message is placed in the buffer stream (buffer), and after the target address is released from occupation, the request message and the corresponding transaction are taken out from the buffer stream for processing.

[0083] Step 104: Read data from the first cache or the memory, or write data to the first cache or the memory, according to the cache status query result and the type of the request message.

[0084] In this embodiment, different processing for writing or reading data is required according to different types of request messages; and different processing based on the first cache or the memory is required according to different cache status query results.

[0085] In this embodiment, by setting a first cache as a private cache in the slave node, when reading data from the first cache or writing data to the first cache, since the slave node does not need to interact with the memory, the processing efficiency of the request message is improved; at the same time, there is no need to set a shared cache in the master node to improve the message processing efficiency, which greatly reduces the load pressure on the master node.

[0086] In some embodiments of the present invention, the reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message includes:

[0087] If the cache status query result indicates cache hit and the type of the request message is a write data request message, then determine first target data corresponding to the target address in the first cache, and update the first target data to the second data.

[0088] In this embodiment, when the type of the request message is a write data request message and the cache is hit, it means that the second data is the latest data (dirty data) relative to the first data stored in the first cache, so the first data in the first cache can be updated to the second data.

[0089] In some embodiments, as Figure 3 shown, there are core 0, core 1, and core 2, where core 0 and core 2 are in the I state (invalid state), and core 1 is in the SD state (Share Dirty, shared dirty, that is Figure 3 the second shared state in Figure 3 ). When core 0 initiates a ReadShared (shared read request) transaction for the target address to the master node, since there is dirty data in core 1, the master node sends a SnpShared (shared resource listen) request to core 1, and core 1 responds to the master node with SnpRespData_SC_PD (shared resource listen response with the status of SC state (Share Clean, shared clean, that is

[0090] Further, the master node initiates a WriteNoSnpFull (write request transaction) to the slave node, and the slave node replies with a CompDBIDResp response (response data), declaring that data can be written into the memory; then the master node initiates an NCBWrData (write data request message) to the slave node.

[0091] At this time, what the slave node receives is a write data request message (NCBWrData). In the related art, since the slave node does not set a cache, it needs to interact with the memory, and the slave node initiates a Write_REQ (write operation) to the memory to write the dirty data into the memory.

[0092] In the example of the present invention ( Figure 3 in the direction indicated by the arrow), since the slave node is provided with a first cache and assuming a cache hit, the cache data corresponding to the target address in the first cache is updated, and there is no need to write the dirty data into the memory (no write operation is performed), that is, the slave node does not need to interact with the memory, greatly improving the processing efficiency of the request message.

[0093] Then, the master node responds to Core 0 with CompData_SC (response data with the status of the first shared state), and after Core 0 replies with a CompAck (complete response), the cache states of Core 0 and Core 1 become the first shared state, and the ReadShared transaction ends.

[0094] Step 105, generating a response message for the request message.

[0095] In some embodiments of the present invention, the reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0096] If the cache status query result indicates a cache miss and the type of the request message is a write data request message, then determine whether the first data reaches a storage threshold in the first cache;

[0097] If the first data reaches the storage threshold in the first cache, then determine second target data among the multiple pieces of first data; wherein, the second target data is the data that has not been used for the longest time among the multiple pieces of first data;

[0098] Write the second target data into the memory, and delete the second target data from the first cache;

[0099] Write the second data into the first cache.

[0100] In this embodiment, when the type of the request message is a write data request message and the cache misses, it is necessary to determine whether the first data stored in the first cache reaches the storage threshold, that is, whether the first cache is full.

[0101] Exemplarily, if the first cache is full, based on the LRU (Least Recently Used) algorithm, determine the second target data, which is the data that has not been used for the longest time. Then, write the second target data into the memory, delete the second target data, and release the storage space of the first cache. After the storage space of the first cache is released, write the second data into the first cache.

[0102] In some embodiments of the present invention, after determining whether the first data in the first cache reaches the storage threshold, it further includes:

[0103] If the first data in the first cache does not reach the storage threshold, write an empty cache line to the first cache according to the target address.

[0104] In this embodiment, when the first cache is not full, write an empty cache line corresponding to the target address to the first cache to play a placeholder role.

[0105] Exemplarily, after writing the empty cache line, when the slave node receives a write data request for the target address again, the cache hits, and the empty cache line is updated to the corresponding memory data to be written.

[0106] In some examples, the write data request message can be processed based on a message processing pipeline. A message processing pipeline is a technology used in communication systems for efficient data processing and transmission. By dividing the data processing process into multiple stages, each stage can process different messages in parallel, thereby improving the throughput and efficiency of the system. The specific pipeline stages for processing the write data request message are as follows:

[0107] In the message processing pipeline, as Figure 4 shown, after the slave node receives a write data request message (write request) for the target address, it enters the processing stage of address query and cache status query. In this stage, determine whether the target address is occupied. If it is occupied, wait for the target address to be released; if it is not occupied, query and record whether the first cache in the slave node hits. After completion, take one beat on the message processing pipeline (pipeline bubble).

[0108] Further, enter the stage of performing different processing according to the cache status query result, as Figure 5 shown:

[0109] If the cache status query result indicates a cache hit, update the cache data stored in the slave node, and the transaction ends; if not, continue to determine whether the cache is full:

[0110] If the cache is full, perform cache replacement based on the LRU algorithm, and the transaction ends; if not, write to an empty cache line, and the transaction ends.

[0111] In some embodiments of the present invention, the step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0112] If the cache status query result indicates a cache hit and the type of the request message is a read data request message, determine the first target data corresponding to the target address in the first cache, and read the first target data;

[0113] Generating the response message of the request message includes:

[0114] Generate the response message of the request message according to the first target data.

[0115] In this embodiment, if the cache status query result indicates a cache hit, it means that the first target data corresponding to the target address is stored in the first cache of the slave node. The first target data is directly obtained from the first cache, and a response message with the first target data attached is generated and returned to the initiator of the request message.

[0116] In some examples, as Figure 6 shown, there are core 0, core 1, and core 2, where core 0, core 1, and core 2 are all in the I state. Core 0 initiates a request message for a ReadOnce transaction (the first read request transaction) for the target address. After receiving it, the master node initiates a request message for a ReadNoSnp transaction (the second read request transaction) to the slave node.

[0117] At this time, the request message received by the slave node is a read data request message (ReadNoSnp). In the related art, the slave node needs to initiate a Read_REQ request (read request) to the memory, and the memory then returns Data (data) to the slave node. After receiving the Data, the slave node then returns CompData_I (that is, Figure 6 the response data_invalid state in, which refers to the response message with the attached status being the invalid state) to the master node. The process of the slave node initiating a request to the memory and reading data from the memory usually takes dozens of clock cycles, resulting in an extended processing time of the request message.

[0118] In the example of the present invention ( Figure 6in the direction indicated by the arrow), since the first cache is set at the slave node and the cache is hit, the first target data corresponding to the target address is directly read from the first cache (i.e., Figure 6 the response data_invalid state in), the slave node can directly return CompData_I to the master node without interacting with the memory. By setting the first cache, not only the memory access frequency is reduced, but also the efficiency of the system for processing request messages is improved.

[0119] Finally, the master node returns the CompData_I response message to core 0, and core 0 replies with CompAck (transaction completion response), and the processing of the message and the transaction ends.

[0120] In some embodiments of the present invention, the reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0121] If the cache status query result indicates cache miss and the type of the request message is a read data request message, the third target data is read from the memory according to the target address;

[0122] The generating the response message of the request message includes:

[0123] Generating the response message of the request message according to the third target data.

[0124] In this embodiment, if the cache miss occurs, it means that the data corresponding to the target address is not stored in the first cache, then the third target data corresponding to the target address needs to be read from the memory, and a response message with the third target data attached is generated to be returned to the initiator of the request message.

[0125] In some examples, as Figure 7 shown, the read data request message can be processed based on the message processing pipeline, specifically as follows:

[0126] In the message processing pipeline, when the slave node receives a read data request message (read request), it enters the processing stage of address query and cache status query. In this stage, if the target address is occupied, it waits for the target address to be released; if not occupied, it enters the processing stage of performing different processing according to the cache status query result:

[0127] If the cache status query result indicates cache hit, the cached data is read from the first cache; if the cache status query result indicates cache miss, the data is read from the memory.

[0128] Then, process the request message, generate a response message based on the cached data read from the first cache or the data read from the memory, and return it, ending the transaction.

[0129] In some embodiments of the present invention, the slave node can also receive a request message from the processor core without going through the master node, further improving the efficiency of message processing, reducing the workload of the master node, achieving load balancing, and avoiding local overheating during system operation.

[0130] In some examples, such as Figure 8 shown, there are core 0, core 1, and core 2, where core 0, core 1, and core 2 are all in the I state, and core 0 initiates a request message for a ReadNoSnp transaction (the second read request transaction) for the target address.

[0131] In the related art, after the master node receives ReadNoSnp, it needs to initiate a request message for a ReadNoSnp transaction (the second read request transaction) to the slave node. The slave node needs to initiate a Read_REQ request (read request) to the memory, and the memory then returns Data (data) to the slave node. After the slave node receives the Data, it then returns CompData_I to the master node (that is, Figure 8 the response data in, which refers to a response message with response data whose attached state is the invalid state).

[0132] In the example provided by the present invention (such as Figure 8 the direction indicated by the arrow), core 0 can directly initiate a request message for a ReadNoSnp transaction to the slave node without going through the master node. If the cache hits at this time, the slave node directly returns CompData_I (that is, the first target data required by core 0) to core 0, and core 0 replies with CompAck (transaction completion response), and the message and transaction processing end.

[0133] In some examples, such as Figure 9 shown, there are core 0, core 1, and core 2, where core 0, core 1, and core 2 are all in the I state, and core 0 initiates WriteNoSnp (write request transaction) for the target address.

[0134] In the related art, the master node needs to first reply with CompDBIDResp (write request transaction response) to declare that data can be sent. Then core 0 initiates NCBWrData (write data request message) to the master node, and the master node then sends WriteNoSnp (write request transaction) to the slave node. The slave node replies with CompDBIDResp (response data), and the master node then sends NCBWrData (write data request message). The slave node performs a Write_REQ (write operation) from the memory.

[0135] In the example provided by the present invention (such as Figure 9 the direction indicated by the arrow), directly receive the WriteNoSnp (write request transaction) of core 0 from the node without passing through the master node, return the CompDBIDResp (write request transaction response) to core 0, core 0 initiates the NCBWrData (write data request message), and at this time, the cache hits, and the slave node updates the first cache without interacting with the memory.

[0136] In the above example, not only is the interaction process between the slave node and the memory saved, but also the interaction process related to the master node is saved. Not only is the processing efficiency of the message improved, but also the load pressure on the master node is reduced, avoiding local overheating of the communication system of the on-chip network.

[0137] In some examples of the present invention, a message processing pipeline can also be established on the slave node based on different clock cycles to process the request message, such as Figure 10 shown as follows:

[0138] In clock cycle 0, the slave node receives the request message and loads the request message.

[0139] In clock cycle 1, the slave node simultaneously performs address status query and cache status query.

[0140] In clock cycle 2, the slave node performs different processing according to the result of the cache status query: if the cache hits, read or write data from the first cache; if the cache misses, read or write data from the memory.

[0141] In clock cycle 3, the slave node processes the request message to generate a corresponding response message or data. Among them, if the request message is a read data request message and the cache hits, a corresponding data packet is generated according to the data read from the first cache without accessing the memory. If it is a write data request message, a corresponding response message is directly generated.

[0142] The embodiments of the present invention have the following advantages: By receiving a request message for the processor core or the master node to access the memory, where the request message includes the target address for accessing the memory, performing an address status query based on the target address to obtain an address status query result, and when the address status query result indicates that the target address is not occupied, performing a cache status query based on the target address to obtain a cache status query result in the first cache, reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, and generating a response message for the request message. There is no need to set up a shared cache in the master node, but instead, a first cache is set up on a slave node with a relatively lower functional complexity than the master node. The first cache is used to store the first data (i.e., dirty data) to be written to the memory and the first address corresponding to the first data. When the target address is hit, the request transaction for the memory can be processed through the first target data without accessing the memory, which can not only improve the efficiency of processing messages in the on-chip network, improve the overall operating efficiency of the on-chip network, but also reduce the load of the master node and avoid the problem of excessive local temperature of the processor during the operation of the on-chip network.

[0143] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0144] The embodiments of the present invention also provide an electronic device, as Figure 11 shown, including a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104.

[0145] The memory 1103 is used to store a computer program.

[0146] When the processor 1101 is used to execute the program stored on the memory 1103, it can implement the above-mentioned message processing method based on the on-chip network. The on-chip network includes a memory, multiple processor cores, at least one master node, and at least one slave node. The slave node is provided with a first cache. The specific steps are as follows:

[0147] Receiving a request message for the processor core or the master node to access the memory; where the request message includes the target address for accessing the memory.

[0148] Perform an address status query based on the target address to obtain an address status query result;

[0149] When the address status query result indicates that the target address is not occupied, perform a cache status query based on the target address to obtain a cache status query result in the first cache;

[0150] Read data from the first cache or the memory, or write data to the first cache or the memory, according to the cache status query result and the type of the request message;

[0151] Generate a response message for the request message.

[0152] In some embodiments of the present invention, the type of the request message includes a read data request message or a write data request message, and the first cache stores a plurality of first data to be written to the memory and a first address corresponding to each first data and the memory; wherein, when the request message is the write data request message, the request message further includes second data; the performing a cache status query based on the target address to obtain a cache status query result in the first cache includes:

[0153] If the target address exists among the plurality of first addresses, the cache status query result indicates a cache hit;

[0154] If the target address does not exist among the plurality of first addresses, the cache status query result indicates a cache miss.

[0155] In some embodiments of the present invention, the reading data from the first cache or the memory, or writing data to the first cache or the memory, according to the cache status query result and the type of the request message, includes:

[0156] If the cache status query result indicates a cache hit and the type of the request message is a write data request message, determine a first target data corresponding to the target address in the first cache, and update the first target data to the second data.

[0157] In some embodiments of the present invention, the reading data from the first cache or the memory, or writing data to the first cache or the memory, according to the cache status query result and the type of the request message, further includes:

[0158] If the cache status query result indicates a cache hit and the type of the request message is a read data request message, determine a first target data corresponding to the target address in the first cache, and read the first target data;

[0159] Generating a response message for the request message includes:

[0160] Generating a response message for the request message according to the first target data.

[0161] In some embodiments of the present invention, the step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0162] If the cache status query result indicates cache miss and the type of the request message is a write data request message, then determine whether the first data reaches the storage threshold in the first cache;

[0163] If the first data reaches the storage threshold in the first cache, then determine a second target data among the multiple first data; wherein, the second target data is the data that has not been used for the longest time among the multiple first data;

[0164] Write the second target data into the memory, and delete the second target data from the first cache;

[0165] Write the second data into the first cache.

[0166] In some embodiments of the present invention, after determining whether the first data reaches the storage threshold in the first cache, it further includes:

[0167] If the first data does not reach the storage threshold in the first cache, then write an empty cache line to the first cache according to the target address.

[0168] In some embodiments of the present invention, the step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message further includes:

[0169] If the cache status query result indicates cache miss and the type of the request message is a read data request message, then read a third target data from the memory according to the target address;

[0170] Generating a response message for the request message includes:

[0171] Generating a response message for the request message according to the third target data.

[0172] In some embodiments of the present invention, after receiving a request message for the processor core or the master node to access the memory, it further includes:

[0173] When the query result of the address status indicates that the target address is occupied, store the request message in the buffer stream.

[0174] In some embodiments of the present invention, the input channels corresponding to the slave nodes include a request channel and a data channel. Both the request channel and the data channel are used to receive the request message, and the priority of receiving the request message through the data channel is higher than that of receiving the request message through the request channel.

[0175] In some embodiments of the present invention, the network-on-chip further includes a memory control unit. The slave node is communicatively connected to the memory through the memory control unit, and the master node interacts with the memory through the slave node.

[0176] In some embodiments of the present invention, the network-on-chip further includes a crossbar switch, which is communicatively connected to the master node, the slave node, and the multiple processor cores.

[0177] In some embodiments of the present invention, each of the processor cores is provided with a second cache.

[0178] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0179] The communication interface is used for communication between the above terminal and other devices.

[0180] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0181] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU for short), a Network Processor (NP for short), etc.; it may also be a Digital Signal Processor (DSP for short), an Application Specific Integrated Circuit (ASIC for short), a Field-Programmable Gate Array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0182] Some embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned message processing method based on the network-on-chip is implemented.

[0183] Some embodiments of the present invention also provide a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned message processing method based on the network-on-chip is implemented.

[0184] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments.

[0185] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0186] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, please refer to each other.

[0187] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0188] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a means for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0189] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0190] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0191] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0192] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the above element.

[0193] The above provides a detailed introduction to the network-on-chip-based message processing method, device, medium, and product. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A message processing method based on a network on chip, characterized in that: The network on chip includes a memory, a plurality of processor cores, at least one master node and at least one slave node, the slave node is provided with a first cache, the type of the request message includes a read data request message or a write data request message, the first cache stores a plurality of first data to be written into the memory and a first address corresponding to each of the first data and the memory; wherein, when the request message is the write data request message, the request message also includes second data; the method includes: receiving a request message from the processor core or the master node to access the memory; wherein the request message includes a target address of the memory to be accessed; Performing an address status query according to the target address to obtain an address status query result; When the address status query result indicates that the target address is not occupied, performing a cache status query according to the target address to obtain a cache status query result in the first cache; Reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message; Generating a response message to the request message; The querying the cache status according to the target address to obtain the cache status query result in the first cache includes: If the target address exists in the plurality of first addresses, the cache status query result indicates a cache hit; If the target address does not exist in the plurality of first addresses, the cache status query result indicates a cache miss.

2. The method for processing messages based on network on chip according to claim 1, characterized in that: The reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, includes: If the cache status query result indicates a cache hit, and the type of the request message is a write data request message, first target data corresponding to the target address is determined in the first cache, and the first target data is updated to the second data.

3. The method for processing messages based on network on chip according to claim 1, characterized in that: The step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, further includes: If the cache status query result indicates a cache hit, and the type of the request message is a read data request message, determining first target data corresponding to the target address in the first cache, and reading the first target data; The generating of a response message to the request message comprises: A response message to the request message is generated according to the first target data.

4. The method for processing messages based on network on chip according to claim 1, characterized in that: The step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, further includes: If the cache status query result indicates a cache miss, and the type of the request message is a write data request message, determining whether the first data reaches a storage threshold in the first cache; If the first data reaches the storage threshold in the first cache, second target data is determined from the plurality of first data; wherein the second target data is the data which has not been used for the longest time from the plurality of first data; Writing the second target data into the memory, and deleting the second target data in the first cache; The second data is written into the first cache.

5. The method for processing messages based on network on chip according to claim 4, characterized in that: After determining whether the first data reaches the storage threshold in the first cache, the method further includes: If the first data does not reach the storage threshold in the first cache, an empty cache line is written to the first cache according to the target address.

6. The method for processing messages based on network on chip according to claim 1, characterized in that: The step of reading data from the first cache or the memory, or writing data to the first cache or the memory according to the cache status query result and the type of the request message, further includes: If the cache status query result indicates a cache miss, and the type of the request message is a read data request message, reading third target data from the memory according to the target address; The generating of a response message to the request message comprises: A response message to the request message is generated according to the third target data.

7. The method for processing messages based on a network on chip according to any one of claims 1 to 6, characterized in that: After receiving the request message for the processor core or the master node to access the memory, the method further includes: When the address status query result indicates that the target address is occupied, the request message is stored in a buffer flow.

8. The method for processing messages based on a network on chip according to any one of claims 1 to 6, characterized in that: The input channel corresponding to the slave node includes a request channel and a data channel, both of which are used to receive the request message, and the priority of receiving the request message through the data channel is higher than the priority of receiving the request message through the request channel.

9. The method for processing messages based on a network on chip according to any one of claims 1 to 6, characterized in that: The on-chip network also includes a memory control unit, the slave node is communicatively connected to the memory via the memory control unit, and the master node interacts with the memory via the slave node.

10. The method for processing messages based on a network on chip according to any one of claims 1 to 6, characterized in that: The on-chip network further includes a cross switch, which is communicatively connected with the master node, the slave node, and the plurality of processor cores.

11. The method for processing messages based on a network on chip according to any one of claims 1 to 6, characterized in that: Each of the processor cores is provided with a second cache.

12. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the method for processing messages based on a network on a chip as claimed in any one of claims 1 to 11 is implemented.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for processing messages based on a network on a chip according to any one of claims 1 to 11 is implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by the processor, the method for processing messages based on a network on a chip as described in any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Data processing method of multi-core processor, server, product and medium

    CN118838863A