Multi-core on-chip network transaction processing system, method, device, medium and product

By setting a first cache on the master node of the multi-core on-chip network transaction processing system to store dirty data of the processing core and process it according to the request type, the problem of increased communication and synchronization overhead between cores in the multi-core processor is solved, and the efficiency of cache consistency maintenance is improved.

CN119544643BActive Publication Date: 2025-06-06SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510081134.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

As the number of processor cores increases, so does communication and synchronization overhead between cores, limiting the performance and efficiency of multi-core processors, especially with challenges in cache management and cache consistency maintenance.

Method used

A multi-core on-chip network transaction processing system is designed, wherein the master node is provided with a first cache for storing dirty data of the processing core, and when receiving a request transaction, it is processed according to the request type (read or write). For read requests, the master node queries whether the local cache directory hits the target address. If it hits, it will directly read data from the cache for reply response; for write requests, the master node will directly write data to the first cache.

Benefits of technology

The main node stores the latest dirty data, which improves the cache hit rate, reduces the frequency of memory access, and significantly improves the maintenance efficiency of cache consistency in the cache cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119544643B_ABST
    Figure CN119544643B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a multi-core on-chip network transaction processing system, method, device, medium and product, wherein the multi-core on-chip network transaction processing system comprises: a plurality of clusters, wherein the clusters comprise four processing cores, a slave node and a master node, wherein each of the processing cores is provided with a cache, the slave node is provided with a private cache, and the slave node comprises a four-stage pipeline architecture; the slave node is used to process the request transaction through the four-stage pipeline architecture upon receiving a request transaction, and generate a target response message or target response message data, wherein the four-stage pipeline architecture is constructed based on different clock cycles to perform periodic processing on the request transaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cluster technology, and in particular to a multi-core on-chip network transaction processing system, method, device, medium and product. Background Art

[0002] The MCSoC (Multi-Core System-on-Chip) multi-core on-chip network architecture consists of multiple clusters. In addition to multiple processing cores, each cluster also includes HN (Home Node), SN (Subordinate Node) and Crossbar. The HN (Home node) plays an important role in coordinating and managing the multi-core system in the cluster. The master node is committed to maintaining cache consistency within the cluster, controlling memory access and scheduling key tasks to improve system efficiency and performance.

[0003] In a cluster, as the number of processor cores increases, the communication and synchronization overhead between cores also increases, greatly limiting the performance and efficiency of multi-core processors. Therefore, when multiple cores access the cache, how to efficiently manage the communication and synchronization operations between cores and maintain the cache consistency of multiple cores becomes an urgent problem to be solved. Summary of the invention

[0004] The purpose of the embodiments of the present invention is to provide a multi-core on-chip network transaction processing system, method, device, medium and product. The specific technical solution is as follows:

[0005] In a first aspect of the present invention, a multi-core on-chip network transaction processing system is provided, the multi-core on-chip network transaction processing system comprising: at least one cluster, the cluster comprising a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and the cache data stored in the first cache is dirty data of the processing core;

[0006] The master node is used to receive a request transaction sent by any processing core; when it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or, when it is detected that the request transaction is a write request transaction, directly write the data carried by the write request transaction into the first cache.

[0007] Optionally, the cluster further includes at least one slave node, a memory control unit and a memory, the slave node is communicatively connected to the memory via the memory control unit, the master node interacts with the memory via the slave node, and each processing core is provided with a second cache.

[0008] Optionally, the cluster further comprises a cross switch, wherein the cross switch is communicatively connected to a plurality of the processing cores, at least one slave node, and at least one master node;

[0009] The crossbar switch is used for scheduling cluster tasks, and the scheduling cluster tasks includes sending a message to the master node in each preset clock cycle.

[0010] Optionally, the cache data stored in the first cache includes at least one of data actively written by the processing core, dirty data evicted and written back to the memory by the processing core, and dirty data written back to the memory after receiving monitoring by the processing core.

[0011] In a second aspect of the implementation of the present invention, a multi-core on-chip network transaction processing method is further provided, which is applied to a master node in any multi-core on-chip network transaction processing system according to the first aspect, and the method comprises:

[0012] Receive request transactions sent by any processing core;

[0013] In the case where it is detected that the request transaction is a read request transaction, determining whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, replying to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or,

[0014] When it is detected that the request transaction is a write request transaction, the data carried by the write request transaction is written into the first cache based on the preset cache consistency maintenance table.

[0015] Optionally, the request transaction includes a read request transaction, and when the request transaction is received, processing the request transaction by the second cache includes:

[0016] When a read request transaction is received, it is determined whether the target address carried by the read request transaction exists in the local cache directory, and if so, the read request transaction is processed.

[0017] Optionally, the preset cache consistency maintenance table includes a cache line status corresponding to a processing core, a cache line status corresponding to a monitored node, a cache line status corresponding to a master node, and whether to write back to memory.

[0018] Optionally, the cache line state includes an invalid state, a shared clean state, a shared dirty state, an exclusive clean state, and an exclusive dirty state.

[0019] Optionally, when detecting that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0020] When it is detected that the write request transaction is a WriteCleanFull write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the shared clean state, the data corresponding to the WriteCleanFull write request transaction is stored in the first cache, and the node state corresponding to the master node is converted from the invalid state to the shared dirty state.

[0021] Optionally, when detecting that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0022] When it is detected that the write request transaction is a WriteBackFull write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the invalid state, the data corresponding to the WriteBackFull write request transaction is stored in the first cache, and the node state corresponding to the master node is determined based on the node state corresponding to the processing core.

[0023] Optionally, when detecting that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0024] When it is detected that the write request transaction is a WriteEvictFull write request transaction, the node state corresponding to the processing core is converted from the exclusive clean state to the invalid state, and the data carried by the WriteEvictFull write request transaction is prohibited from being written to the first cache.

[0025] Optionally, after the step of receiving a request transaction sent by any processing core, the method further includes:

[0026] When it is detected that the request transaction sent by the processing core is any one of a ReadUnique transaction, a CleanUnique transaction, and a MakeUnique transaction, the node state corresponding to the processing core is converted from the invalid state to the exclusive clean state, the dirty state data stored in the first cache is used as cache data to be evicted, the cache data to be evicted is written back to the memory, and new cache data is received again.

[0027] Optionally, the method further comprises:

[0028] When it is detected that the first cache is in a full state, cache data to be evicted is determined in the cache data according to the usage frequency of the processing core, the cache data to be evicted is sent to the memory, and new cache data is received again.

[0029] In a third aspect of the implementation of the present invention, there is also provided a communication device, comprising: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor;

[0030] The processor is used to read the program in the memory to implement the multi-core on-chip network transaction processing method as described in any one of the first aspect or the second aspect.

[0031] In the fourth aspect of the implementation of the present invention, a computer-readable storage medium is also provided, in which instructions are stored. When the computer-readable storage medium is executed on a computer, the computer implements the multi-core on-chip network transaction processing method as described in any one of the first aspect or the second aspect.

[0032] In the fifth aspect of the implementation of the present invention, a computer program product is also provided, including a computer program / instruction, which, when executed by a processor, implements the multi-core on-chip network transaction processing method as described in any one of the first aspect or the second aspect.

[0033] The multi-core on-chip network transaction processing system provided by an embodiment of the present invention includes: at least one cluster, the cluster includes a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and the cache data stored in the first cache is the dirty state data of the processing core; the master node is used to receive a request transaction sent by any processing core; in the case where it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or, in the case where it is detected that the request transaction is a write request transaction, directly write the data carried by the write request transaction to the local cache directory corresponding to the first cache. Enter the first cache, that is, in the embodiment of the present invention, a first cache is set on the master node, and the first cache can store the latest dirty data from multiple processing cores. The dirty data is not written back to the memory immediately after passing through the master node but is stored in the HN cache. Therefore, when the processing core initiates a request transaction as a requesting core node, if it is a read transaction, it can be directly read in the first cache of the master node after hitting, and the cache consistency between multiple cores can be maintained based on the preset cache consistency maintenance table. If it is a write transaction, it can be determined based on the preset cache consistency maintenance table whether to write the carried data into the first cache. Therefore, storing the latest dirty data through the master node can reduce the frequency of memory access while improving the cache hit rate, thereby greatly improving the efficiency of maintaining cache consistency within the cache cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0035] Figure 1 A system schematic diagram of a multi-core on-chip network transaction processing system provided by an embodiment of the present invention;

[0036] Figure 2 A flowchart of a multi-core on-chip network transaction processing method provided by an embodiment of the present invention;

[0037] Figure 3 is a schematic diagram of a communication device provided by an embodiment of the present invention;

[0038] Figure 4 is a schematic diagram of an exemplary transaction processing provided by an embodiment of the present invention;

[0039] Figure 5 is another exemplary transaction processing schematic diagram provided by an embodiment of the present invention;

[0040] Figure 6is another exemplary transaction processing schematic diagram provided by an embodiment of the present invention;

[0041] Figure 7 is another exemplary transaction processing schematic diagram provided by an embodiment of the present invention;

[0042] Figure 8 It is another exemplary transaction processing diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. However, it can be understood by those skilled in the art that in the embodiments of the present invention, many technical details are proposed in order to enable readers to better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present invention can also be implemented. The division of the following embodiments is for the convenience of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined and referenced with each other without contradiction.

[0044] Reference Figure 1 , showing a system schematic diagram of a multi-core on-chip network transaction processing system provided by an embodiment of the present invention.

[0045] A multi-core on-chip network transaction processing system, the multi-core on-chip network transaction processing system comprising: at least one cluster, the cluster comprising a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and cache data stored in the first cache is dirty data of the processing core;

[0046] The master node is used to receive a request transaction sent by any processing core; when it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or, when it is detected that the request transaction is a write request transaction, write the data carried by the write request transaction into the first cache.

[0047] Furthermore, the cluster also includes at least one slave node, a memory control unit and a memory, the slave node is communicatively connected to the memory through the memory control unit, the master node interacts with the memory through the slave node, and each processing core is provided with a second cache.

[0048] Furthermore, the cluster further comprises a cross switch, wherein the cross switch is communicatively connected to a plurality of the processing cores, at least one slave node, and at least one master node;

[0049] The crossbar switch is used for scheduling cluster tasks, and the scheduling cluster tasks includes sending a message to the master node in each preset clock cycle.

[0050] It should be noted that, in the embodiments of the present invention, referring to Figure 1 The processing core is Core, a second cache is arranged on the processing core, a first cache is arranged on HN (Home Node), a SN (Subordinate Node) is connected to the memory through the memory control unit, and the master node interacts with the memory through the slave node, for example, writing data back to the memory.

[0051] Specifically, Figure 1 It is an exemplary multi-core on-chip network transaction processing system. It can be seen that it consists of eight processing cores, a master node and a slave node, wherein the processing cores all have a second cache, and the master node is designed to have a first cache. The data cached in the first cache are all dirty data and are not written back to the memory temporarily. The data in the cache will only be written to the memory when the cache is replaced and a specific transaction is initiated.

[0052] When receiving a data read transaction request, the master node first queries the local cache directory. If the directory hits, it can directly reply without reading data from other nodes or memory.

[0053] When the master node receives a write transaction, the data may be directly written into the cache without writing into the memory.

[0054] The slave node has no cache and is responsible for communicating with the DDR through the MC, that is, writing data into the memory and reading data from the memory, wherein the MC is a memory control unit and the DDR is a memory module.

[0055] In the architecture of the example of the present invention, in order to maintain cache consistency between the master node cache and other multiple processing cores, all processing must pass through the master node, and the slave node cannot communicate directly with the processing core.

[0056] It should be noted that, in the embodiment of the present invention, the crossbar is a crossbar, and the crossbar is responsible for task scheduling within the cluster. The same node can only receive one message within one clock.

[0057] Furthermore, the cache data stored in the first cache includes at least one of data actively written by the processing core, dirty data evicted and written back to the memory by the processing core, and dirty data written back to the memory after receiving monitoring by the processing core.

[0058] The multi-core on-chip network transaction processing system provided by an embodiment of the present invention includes: at least one cluster, the cluster includes a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and the cache data stored in the first cache is the dirty state data of the processing core; the master node is used to receive a request transaction sent by any processing core; in the case where it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or, in the case where it is detected that the request transaction is a write request transaction, directly write the data carried by the write request transaction to the local cache directory corresponding to the first cache. Enter the first cache, that is, in the embodiment of the present invention, a first cache is set on the master node, and the first cache can store the latest dirty data from multiple processing cores. The dirty data is not written back to the memory immediately after passing through the master node but is stored in the HN cache. Therefore, when the processing core initiates a request transaction as a requesting core node, if it is a read transaction, it can be directly read in the first cache of the master node after hitting, and the cache consistency between multiple cores can be maintained based on the preset cache consistency maintenance table. If it is a write transaction, it can be determined based on the preset cache consistency maintenance table whether to write the carried data into the first cache. Therefore, storing the latest dirty data through the master node can reduce the frequency of memory access while improving the cache hit rate, thereby greatly improving the efficiency of maintaining cache consistency within the cache cluster.

[0059] Reference Figure 2 , shows a flowchart of the steps of a multi-core on-chip network transaction processing method provided by an embodiment of the present invention, which is applied to a master node in any multi-core on-chip network transaction processing system described in the first aspect, and the method may include:

[0060] Step 101: Receive a request transaction sent by any processing core.

[0061] It should be noted that, in the embodiment of the present invention, it is applied to the master node. As described in the foregoing, a first cache is set on the master node. This first cache can store the latest dirty data from multiple processing cores, wherein the dirty data is data that has been modified but has not yet been written back to the memory or persistent storage.

[0062] Therefore, when receiving a request transaction sent by any processing core in the multi-core on-chip network transaction processing system, the corresponding processing core may be the request core that sends the request transaction.

[0063] The request transaction may include a transaction type defined based on the CHI (Coherent Hub Interface) protocol. For specific transaction types, please refer to the subsequent elaboration.

[0064] Step 102, when it is detected that the request transaction is a read request transaction, determining whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, replying to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or,

[0065] Step 103: When it is detected that the request transaction is a write request transaction, write the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table.

[0066] It should be noted that, in the embodiment of the present invention, the above steps 102 - 103 are respectively performed for processing a read request transaction and a write request transaction.

[0067] Specifically, when it is detected that the request transaction is a read request transaction, the master node first queries whether there is a hit in the local cache directory corresponding to the first cache set in the master node. Specifically, it can be confirmed in the cache directory based on the target address carried by the request transaction. If it hits, it can directly reply to the request transaction sent by the processing core without reading the data required for the request transaction from other nodes (such as other processing cores) or in the memory.

[0068] In the case where the detection request transaction is a write request transaction, similarly, when the master node receives the write transaction, the data carried by the write request transaction can be directly written into the first cache corresponding to the master node without writing into the memory.

[0069] It should be noted that in the embodiment of the present invention, the cache data sources of the designed master node include three types: one is the latest data that needs to be written back to the memory after the CPU calculates the data that the processing core actively writes. Another is the dirty data that needs to be written back to the memory because the cache in the core is full and is evicted by the core cache. The last is the dirty data that needs to be written back to the memory after the core receives the monitoring.

[0070] It can be seen that the cache data stored in the first cache of the master node designed by the present invention are all dirty data (dirty data has been explained in the foregoing). Therefore, by storing the latest dirty data in the first cache of the master node, the dirty data is not written back to the memory immediately after passing through the master node but is stored in the first cache of the master node. When the master node receives a read request transaction cache hit or receives a write request transaction, the master node can directly reply or write without sending a listening request to other nodes (processing cores) or performing corresponding read or write operations in the memory, thereby reducing the frequency of memory access while improving the cache hit rate.

[0071] Furthermore, in one embodiment, the method further comprises:

[0072] When it is detected that the first cache is in a full state, cache data to be evicted is determined in the cache data according to the usage frequency of the processing core, the cache data to be evicted is sent to the memory, and new cache data is received again.

[0073] It should be noted that when the cache of the master node is full, the cache needs to be replaced, and the replaced dirty cache data needs to be written back to the memory to maintain cache consistency. If the master node cache is full, the master node cache data is directly overwritten without writing it back to the memory. In addition, according to the CHI protocol, the core that initiates the ReadUnique transaction, CleanUnique transaction, and MakeUnique transaction must eventually obtain the UC state, and the dirty data must also be written back to the memory to maintain cache consistency for the master node and other cores.

[0074] If the master node cache data hits, it can be directly replied without reading data from other cores and memory, because the master node cache is designed to cache only the latest dirty data.

[0075] The data cached in the HN will be evicted and written back to the memory only when the master node cache data is full or the requesting node needs to become exclusive.

[0076] Therefore, in the embodiment of the present invention, a preset cache consistency maintenance table is designed to maintain cache consistency between the HN and other processing cores, and ensures that the processing efficiency of the HN node is greatly improved while maintaining cache consistency.

[0077] Furthermore, the preset cache consistency maintenance table includes the cache line status corresponding to the processing core, the cache line status corresponding to the monitored node, the cache line status corresponding to the master node, and whether to write back to the memory.

[0078] Specifically, the cluster maintains cache consistency using the CHI protocol, and defines five states for describing cache line states, as shown in the following Table 1.

[0079] Table 1 Example cache line state definition table

[0080]

[0081] Furthermore, the cache line states include an invalid state, a shared clean state, a shared dirty state, an exclusive clean state, and an exclusive dirty state.

[0082] It should be noted that, in an embodiment of the present invention, the cache line state may include five types, that is, the cache line state rules include an invalid state, a shared clean state, a shared dirty state, an exclusive clean state and an exclusive dirty state. Specifically, the I state is an invalid state; the SC state is a shared clean state, the cache line data is clean, and there may be cache line copies in other cores; the SD state is a shared dirty state, the cache line data is dirty, and there may be cache line copies in other cores; the UC state is an exclusive clean state, the cache line data is clean, and there are no cache line copies in other cores; the UD state is an exclusive dirty state, the cache line data is dirty, and there are no cache line copies in other cores.

[0083] Further, after the step of receiving a request transaction sent by any processing core, the method further includes:

[0084] When it is detected that the request transaction sent by the processing core is any one of an exclusive read request transaction, a unique exclusive request transaction, and an exclusive upgrade request transaction, the node state corresponding to the processing core is converted from the invalid state to the exclusive clean state, the dirty state data stored in the first cache is used as cache data to be evicted, the cache data to be evicted is written back to the memory, and new cache data is received again.

[0085] It should be noted that, in the embodiment of the present invention, the cluster supports 19 transactions specified by the CHI protocol, as shown in Table 2. Table 2 records the transactions supported by the master node, and the cache consistency maintenance scheme between the master node cache and other processing cores.

[0086] Therefore, as shown in Table 2, the exclusive read request transaction is the ReadUnique transaction, the unique exclusive request transaction is the CleanUnique transaction, and the exclusive upgrade request transaction is the MakeUnique transaction. Even if the memory is not full, the dirty data in the HN cache needs to be evicted from the SD or UD state to the I state. The reason is that the ReadUnique transaction, CleanUnique transaction, and MakeUnique transaction need to change the state of the node initiating the request to the exclusive state. In order to maintain the cache consistency between the master node cache and the core, the data cached in the HN needs to be written back to the memory.

[0087] The master node cache consistency maintenance solution can maximize the cache hit rate and memory access frequency, reduce the cache consistency maintenance overhead within the cluster, and improve processor efficiency.

[0088] Table 2 Exemplary master node cache maintenance scheme status table

[0089]

[0090] It should be noted that the CHI (Coherent Hub Interface) protocol is a complex cache coherent interconnect protocol that defines multiple transaction types to handle different types of data transfer and cache coherence operations.

[0091] Table 2 above shows some main transaction types specified in the CHI protocol. Specifically, the transaction types may include the following:

[0092] ReadOnce transaction type: ReadOnce transaction type is used for a single read operation. The operation does not change the state of the data and does not affect the consistency of the data in the cache.

[0093] ReadNoSnp transaction type: The ReadNoSnp transaction type is used for read operations and does not require cache consistency monitoring.

[0094] ReadOnceCleanInvalid transaction type: ReadOnceCleanInvalid transaction type is used for read operations. During the read operation, the data block in the cache can be cleaned up and set to an invalid state, thereby ensuring that the read data is consistent with the latest data in other caches or main memory.

[0095] ReadOnceMakeInvalid transaction type: The ReadOnceMakeInvalid transaction type is used to invalidate the requested data and set the corresponding status of the copies of the data block in all other caches to an invalid state, thereby ensuring the consistency of the read data and preventing the data block from being reused in other caches.

[0096] ReadShared transaction type: ReadShared transaction type is used for read operations. During the reading process, ReadShared transaction type allows multiple caches to hold copies of the data block and perform read-only access without changing the owner status of the data block. Therefore, this operation will not cause the data copies in other caches to become invalid.

[0097] ReadUnique transaction type: The ReadUnique transaction type is used to request an exclusive copy of a data block and remove the data block from all other caches (making it invalid), ensuring that the requester is the only cache holding a modifiable copy. The ReadUnique transaction type operation can ensure that the requester can write data without worrying about data inconsistency in other caches, ensuring that the returned data is not shared in other caches.

[0098] CleanUnique transaction type: The CleanUnique transaction type is used to request an exclusive copy of a data block and write the modified content of the data block back to the main memory (if there is dirty data) before removing other cached copies (making them invalid). The operation of the CleanUnique transaction type can ensure that the data in the main memory is the latest and set the requester as the sole holder of the data.

[0099] MakeUnique transaction type, MakeUnique transaction type is used to upgrade the existing shared or read-only data block in the requester's cache to an exclusive state and invalidate the copies of the data block in all other caches. Among them, the operation of MakeUnique transaction type allows the requester to write to the data block without re-reading the data from the memory or worrying about data inconsistency in other caches.

[0100] CleanInvalid transaction type: The CleanInvalid transaction type is used to write data blocks in the cache that may be modified (dirty data) back to the main memory and remove the data blocks from all caches (making them invalid). The command corresponding to the CleanInvalid transaction type can ensure that the data in the main memory is the latest and clear the corresponding copies in all caches.

[0101] MakeInvalid transaction type: The MakeInvalid transaction type is used to invalidate all copies of a specified data block held in all caches without writing the dirty data back to the main memory. When other caches are about to have exclusive access to the data block, the MakeInvalid transaction type clears all old copies to avoid data inconsistency.

[0102] CleanShared transaction type: The CleanShared transaction type is used to write data blocks that may be modified, i.e., dirty data, back to the main memory, but retain the shared copy in the cache so that the data block can continue to be read-only accessed by multiple caches. The CleanShared transaction type operation can ensure that the data in the main memory is up to date and maintain the shared state in the cache.

[0103] WriteUniqueFull transaction type: WriteUniqueFull transaction type is used to write a complete data block to the main memory and invalidate the copies of the data block in all other caches. The command corresponding to the WriteUniqueFull transaction type is used to ensure the consistency of data in the main memory and ensure that only one cache holds the latest copy of the data block.

[0104] WriteUniquePtl transaction type, WriteUniquePtl transaction type is used to write partial bytes of a data block, rather than the entire data block, to the main memory, and invalidate all other cache copies of the data block. That is, WriteUniquePtl transaction type is used to modify part of the content in a data block while maintaining data consistency.

[0105] WriteNoSnpFull transaction type,The WriteNoSnpFull transaction type is used to write a complete data block to the main memory without performing consistency checks or invalidation operations on copies in other caches.

[0106] WriteNoSnpPtl transaction type,The WriteNoSnpPtl transaction type is used to write some bytes of a data block to main memory without performing consistency checks or invalidation operations on copies in other caches.

[0107] WriteCleanFull transaction type, WriteCleanFull transaction type is used to write a complete data block, for example, dirty data, back to the main memory, and keep the copy of the data block in the cache in a clean state, that is, no longer dirty data. Among them, this command is used to ensure that the data in the main memory is up to date without affecting the cache status.

[0108] WriteBackFull transaction type: WriteBackFull transaction type is used to write a complete dirty data block back to the main memory and remove the data block from the cache, that is, invalidate it. This command is usually used to safely write modified data back to the main memory when the cache is replaced or invalidated.

[0109] WriteEvictFull transaction type, WriteEvictFull transaction type is used to write a complete data block, for example, dirty data, back to the main memory, and at the same time remove the data block from the cache, that is, invalidate it, and inform the main memory that the cache line is evicted. Among them, this command is usually used for cache replacement or eviction operations to ensure data consistency in the main memory.

[0110] Further, when it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0111] When it is detected that the write request transaction is a WriteCleanFull write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the shared clean state, the data corresponding to the WriteCleanFull write request transaction is stored in the first cache, and the node state corresponding to the master node is converted from the invalid state to the shared dirty state.

[0112] It should be noted that, in an embodiment of the present invention, the first write request transaction is the WriteCleanFull transaction. When the transaction is initiated, the state of the requesting node is UD state, and when the transaction is completed, the state of the requesting node is SC state. Therefore, the data can be saved in the master node cache and not written back to the memory immediately. The master node cache state changes from I state to SD state.

[0113] Further, when it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0114] When it is detected that the write request transaction is a WriteBackFull write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the invalid state, the data corresponding to the WriteBackFull write request transaction is stored in the first cache, and the node state corresponding to the master node is determined based on the node state corresponding to the processing core.

[0115] It should be noted that, in the embodiment of the present invention, the second write request transaction is the WriteBackFull transaction. When the transaction is initiated, the data is directly cached in the master node cache and is not immediately written back to the memory. The master node cache status is determined by the requesting node status.

[0116] Further, when it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes:

[0117] When it is detected that the write request transaction is the third write request transaction, the node state corresponding to the processing core is converted from the exclusive clean state to the invalid state, and the data carried by the third write request transaction is prohibited from being written into the first cache.

[0118] It should be noted that, in the embodiment of the present invention, the third write request transaction is the WriteEvictFull transaction. When the transaction is initiated, the state of the requesting node is UC state, and the data to be written is clean and does not need to be written to the master node cache.

[0119] In addition, it should be noted that other request transaction types in Table 2 above do not change the cache line state of the master node cache when the corresponding request transaction is initiated.

[0120] The multi-core on-chip network transaction processing system provided in an embodiment of the present invention includes: a plurality of clusters, wherein the cluster includes at least one processing core, at least one slave node, at least one master node and a memory control unit, wherein each of the processing cores is provided with a first cache, the master node is provided with a second cache, and the slave node and the memory control unit are communicatively connected; the master node is used to process the request transaction through the second cache when receiving the request transaction, that is, in an embodiment of the present invention, the master node, that is, the master node, has a cache and can store the latest non-persistent data, that is, dirty data. The dirty data is not written back to the memory immediately after passing through the master node but is stored in the HN cache. When the master node cache hits, the master node can directly reply without sending a listening request to other cores, which reduces the frequency of memory access while improving the cache hit rate, and greatly improves the efficiency of maintaining cache consistency within the cache cluster.

[0121] In addition, in order to facilitate those skilled in the art to understand the present invention, the present invention is described below through embodiments.

[0122] In one embodiment, referring to Figure 4 , Figure 4This is a flow chart of a ReadShared transaction. A ReadShared transaction is a read request transaction. A ReadShared transaction is essentially a read operation transaction request. For details, please refer to the previous description. The cache state of processing core 0 is I state, the cache state of processing core 1 is SD state, and the cache state of processing core 7 is I state.

[0123] The initial unoptimized ReadShared transaction flow chart is: Figure 4 As shown in the left figure, specifically, the master node receives a ReadShared transaction request, queries the node directory and finds that the directory does not hit, and sends a listen request to other cores. There is dirty data in processing core 1 that needs to be written back to the memory. The master node receives the data sent by processing core 1, sends a data message in the first exclusive state, namely CompData_SC, to processing core 0, and sends a write request WriteNoSnpFull to the slave node SN (the slave node and the memory module are connected in communication). Therefore, the dirty data is written back to the memory at this time. Processing core 0 replies with a transaction completion response CompAck to the master node to indicate the end of the transaction.

[0124] Therefore, in the embodiment of the present invention, by Figure 1 The architecture can optimize the request transaction processing process. For details, please refer to Figure 4 The right side, where the arrow points, is the optimized ReadShared transaction flow chart. The initial state of the master node's cache is I. The master node receives the data sent by processing core 1, keeps a copy of the data in the master node cache, updates the master node cache state to SD, and sends CompData_SC to processing core 0. Processing core 0 replies CompAck to HN to indicate the end of the transaction.

[0125] In summary, through Figure 4 Corresponding embodiments show that the embodiments of the present invention delay the writing of dirty data back to the memory, reduce the time of transaction flow processing, and greatly improve the efficiency of maintaining cache consistency. At the same time, when the address data is requested again, the master node can directly reply.

[0126] In one embodiment, referring to Figure 5 , Figure 5 The flowchart of the ReadUnique transaction is shown in FIG. 1 . Similarly, the processing core is the processing core, the ReadUnique transaction is a read request transaction, and the ReadUnique transaction is essentially a read operation transaction request. For details, please refer to the previous description. The cache state of the processing core 0 is I state, the cache state of the processing core 1 is SC state, and the cache state of the processing core 7 is I state.

[0127] Likewise, Figure 5The left figure is the unoptimized ReadUnique transaction flow chart. HN receives the ReadUnique transaction, queries the node directory and finds a directory hit. Since there is no dirty data in other cores, it sends a listen request to processing core 1. Processing core 1 is in SC state, so the data replied at this time is clean. The master node receives the data sent by processing core 1 and sends CompData_UC to processing core 0. Processing core 0 replies CompAck to HN to indicate the end of the transaction. No dirty data is written back to the memory in the entire process.

[0128] Therefore, in the embodiment of the present invention, by Figure 1 The architecture can optimize the request transaction processing flow. Specifically, the direction indicated by the arrow is the optimized ReadUnique transaction flow chart. The initial state of the master node cache is SD state. The master node receives the data sent by processing core 1. Since the data cached in the master node is non-persistent data, that is, the data cached in the master node is dirty, the data sent by processing core 1 is discarded, and the master node directly reads the data from the cache to reply CompData_UC. Processing core 0 replies CompAck to HN to indicate the end of the transaction. Since processing core 0 eventually needs to get the UC state, the dirty data in the master node needs to be written back to the memory. After the master node receives all the listening responses, it sends WriteNoSnpFull to SN to write the dirty data back to the memory.

[0129] In summary, although the entire transaction process has an additional process of writing back to memory compared to before, the above operations can maintain cache consistency between the master node and the processing core.

[0130] In one embodiment, referring to Figure 6 , Figure 6 The figure is a flow chart of the CleanInvalid transaction. Similarly, the processing core is the processing core, and the CleanInvalid transaction is a request transaction for cleaning invalid data. For details, please refer to the previous description. The cache state of the processing core 0 is I state, the cache state of the processing core 1 is SD state, and the cache state of the processing core 7 is I state.

[0131] Likewise, Figure 6 The left figure is the unoptimized CleanInvalid transaction flow chart. HN receives the CleanInvalid transaction request, queries the node directory and finds a directory hit, and sends a snoop request to processing core 1. There is dirty data in processing core 1 that needs to be written back to the memory. The master node receives the data sent by processing core 1. Since CleanInvalid is a request without data, the master node sends Comp_I to processing core 0 to indicate the end of the transaction. At the same time, the master node sends WriteNoSnpFull to SN to write the dirty data back to the memory.

[0132] Therefore, in the embodiment of the present invention, by Figure 1 The architecture can optimize the request transaction processing flow. Specifically, the arrow indicates the optimized CleanInvalid transaction flow chart. The initial state of the master node cache is I state. The master node receives the data sent by processing core 1, keeps a copy of the data in the master node cache, updates the master node cache state to SD state, and sends a completion response invalid state Comp_I to processing core 0 to indicate the end of the transaction.

[0133] In summary, through Figure 6 Corresponding embodiments show that the embodiments of the present invention delay the writing of dirty data back to the memory, reduce the time of transaction process processing, and greatly improve the efficiency of maintaining cache consistency. At the same time, when the address data is requested again, the master node can directly reply without the slave node responding again.

[0134] In one embodiment, referring to Figure 7 , Figure 7 The figure is a flowchart of the WriteUniqueFull transaction. Similarly, the processing core is the processing core, and the WriteUniqueFull transaction is a write request transaction. The details can be referred to in the previous description, that is, it is used to write data into the cache and ensure that the data is unique in the cache and the written data is complete. The cache state of the processing core 0 is I state, the cache state of the processing core 1 is SC state, and the cache state of the processing core 7 is SC state.

[0135] Likewise, Figure 7 The left figure is the unoptimized WriteUnique transaction flow chart. The master node receives the WriteUniqueFull transaction, queries the node directory, finds that the directory does not hit, and sends a snoop request to other cores. At this time, there is no dirty data in the processing core that needs to be written back to the memory. Therefore, the master node receives the data that needs to be written to the memory from the processing core 0. The master node sends WriteNoSnpFull to the slave node to write the dirty data back to the memory.

[0136] Therefore, in the embodiment of the present invention, by Figure 1 The architecture can optimize the request transaction processing flow. Specifically, the arrow indicates the optimized WriteUniqueFull transaction flow chart. The initial state of the master node cache is SD. The master node receives data that needs to be written to the memory from processing core 0. Since the initial state of the master node cache is SD, only the data in the master node cache needs to be overwritten. After the update, HN is still in SD state and does not need to be written back to the memory.

[0137] Further, in the case where it is detected that the request transaction is a read request transaction, determining whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, replying to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table includes:

[0138] When it is detected that the read request transaction is the first read request transaction and the initial state of the cache corresponding to the master node is a shared dirty state, it is determined whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction. If so, a reply response is directly made to the first read request transaction.

[0139] It should be noted that, in one embodiment, the first read request transaction is a ReadOnce transaction.

[0140] Reference Figure 8 , Figure 8 This is a flowchart of a ReadOnce transaction. ReadOnce (RO) is used for a single read operation, usually for non-cached data. A ReadOnce transaction is a read request transaction. The cache state of processing core 0 is I state, and the cache states of other cores are all I state.

[0141] Figure 8 The left figure is a flowchart of the unoptimized ReadOnce transaction. Specifically, the master node receives the ReadOnce transaction, queries the node directory, finds that the directory is hit and the states of other nodes are all I state, and there is no dirty data in the core that needs to be written back to the memory. The master node needs to read data from the memory. The master node sends ReadNoSnp to SN. The master node receives the data read from the memory and forwards it to the requesting core processing core 0. Processing core 0 replies CompAck to HN to indicate the end of the transaction.

[0142] Therefore, in the embodiment of the present invention, by Figure 1 The arrow indicates the optimized ReadOnce transaction flow chart. The initial state of the master node cache is SD. Therefore, after receiving the ReadOnce transaction request, the master node finds a cache hit and directly replies CompData_I to the processing core 0. The processing core 0 replies CompAck to HN to indicate the end of the transaction. The cache state of the master node is still SD.

[0143] It should be noted that, in the embodiment of the present invention, it can be seen that after optimization, the memory access frequency is reduced, and the efficiency of maintaining cache consistency within the cluster is greatly improved.

[0144] The multi-core on-chip network transaction processing system provided by an embodiment of the present invention includes: at least one cluster, the cluster includes a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and the cache data stored in the first cache is the dirty state data of the processing core; the master node is used to receive a request transaction sent by any processing core; in the case where it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; or, in the case where it is detected that the request transaction is a write request transaction, directly write the data carried by the write request transaction to the local cache directory corresponding to the first cache. Enter the first cache, that is, in the embodiment of the present invention, a first cache is set on the master node, and the first cache can store the latest dirty data from multiple processing cores. The dirty data is not written back to the memory immediately after passing through the master node but is stored in the HN cache. Therefore, when the processing core initiates a request transaction as a requesting core node, if it is a read transaction, it can be directly read in the first cache of the master node after hitting, and the cache consistency between multiple cores can be maintained based on the preset cache consistency maintenance table. If it is a write transaction, it can be determined based on the preset cache consistency maintenance table whether to write the carried data into the first cache. Therefore, storing the latest dirty data through the master node can reduce the frequency of memory access while improving the cache hit rate, thereby greatly improving the efficiency of maintaining cache consistency within the cache cluster.

[0145] The embodiment of the present invention also provides a communication device, such as Figure 3 As shown, it includes a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0146] Memory 703, used for storing computer programs;

[0147] The processor 701, when used to execute the program stored in the memory 703, can implement the following steps:

[0148] When a request transaction is received, the request transaction is processed by the four-stage pipeline architecture to generate a target response message or target response data, wherein the four-stage pipeline architecture is constructed based on different clock cycles to periodically process the request transaction.

[0149] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits together, which are all well known in the art, so this article will not further describe them. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor can be transmitted through a wired medium or transmitted on a wireless medium through an antenna. Further, the antenna also receives data and transmits the data to the processor. The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor when performing operations.

[0150] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0151] The communication interface is used for communication between the above terminal and other devices.

[0152] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0153] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0154] In another embodiment of the present invention, a computer-readable storage medium is provided, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the multi-core on-chip network transaction processing method described in any of the above embodiments.

[0155] In another embodiment of the present invention, a computer program product including instructions is provided, which, when executed on a computer, enables the computer to execute the multi-core on-chip network transaction processing method described in any one of the above embodiments.

[0156] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.

[0157] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0158] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0159] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A multi-core network-on-chip transaction processing system, characterized in that: The multi-core on-chip network transaction processing system includes: at least one cluster, the cluster includes a plurality of processing cores and at least one master node, wherein the master node is provided with a first cache, and the cache data stored in the first cache is dirty data written back to the memory; The master node is used to receive a request transaction sent by any processing core; when it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and the preset cache consistency maintenance table; or, when it is detected that the request transaction is a write request transaction, write the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table; wherein the preset cache consistency maintenance table includes the cache line status corresponding to the processing core, the cache line status corresponding to the monitored node, the cache line status corresponding to the master node, and whether to write back to the memory, and the cache line status includes invalid state, shared interference clean state, shared dirty state, exclusive clean state and exclusive dirty state; when it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction to the first cache based on the preset cache consistency maintenance table includes: when it is detected that the write request transaction is a first write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the shared clean state, the data corresponding to the first write request transaction is stored in the first cache, and the node state corresponding to the master node is converted from the invalid state to the shared dirty state; when it is detected that the write request transaction is a third write request transaction, the node state corresponding to the processing core is converted from the exclusive clean state to the invalid state, and the data carried by the third write request transaction is prohibited from being written to the first cache.

2. The system according to claim 1, characterized in that The cluster also includes at least one slave node, a memory control unit and the memory. The slave node is communicatively connected to the memory through the memory control unit, and the master node interacts with the memory through the slave node. Each processing core is provided with a second cache.

3. The system according to claim 1, characterized in that The cluster further includes a cross switch, wherein the cross switch is communicatively connected to a plurality of the processing cores, at least one slave node, and at least one master node; The crossbar switch is used for scheduling cluster tasks, and the scheduling cluster tasks includes sending a message to the master node in each preset clock cycle.

4. The system according to claim 1, characterized in that The cache data stored in the first cache includes at least one of data actively written by the processing core, dirty data evicted and written back to the memory by the processing core, and dirty data written back to the memory after receiving monitoring by the processing core.

5. A multi-core on-chip network transaction processing method, characterized in that: Applied to a master node in a multi-core network-on-chip transaction processing system as claimed in any one of claims 1 to 4, the method comprising: Receive request transactions sent by any processing core; In the case where it is detected that the request transaction is a read request transaction, determine whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, reply to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table; wherein the preset cache consistency maintenance table includes a cache line status corresponding to the processing core, a cache line status corresponding to the monitored node, a cache line status corresponding to the master node, and whether to write back to the memory, and the cache line status includes an invalid state, a shared clean state, a shared dirty state, an exclusive clean state, and an exclusive dirty state; or, When it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table; Wherein, when it is detected that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table includes: In the case where it is detected that the write request transaction is the first write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the shared clean state, the data corresponding to the first write request transaction is stored in the first cache, and the node state corresponding to the master node is converted from the invalid state to the shared dirty state; When it is detected that the write request transaction is the third write request transaction, the node state corresponding to the processing core is converted from the exclusive clean state to the invalid state, and the data carried by the third write request transaction is prohibited from being written into the first cache.

6. The method according to claim 5, characterized in that When detecting that the request transaction is a write request transaction, writing the data carried by the write request transaction into the first cache based on the preset cache consistency maintenance table comprises: When it is detected that the write request transaction is a second write request transaction, the node state corresponding to the processing core is converted from the exclusive dirty state to the invalid state, the data corresponding to the second write request transaction is stored in the first cache, and the node state corresponding to the master node is determined based on the node state corresponding to the processing core.

7. The method according to claim 5, characterized in that After the step of receiving the request transaction sent by any one of the processing cores, the method further includes: When it is detected that the request transaction sent by the processing core is any one of an exclusive read request transaction, a unique exclusive request transaction, and an exclusive upgrade request transaction, the node state corresponding to the processing core is converted from the invalid state to the exclusive clean state, the dirty state data stored in the first cache is used as cache data to be evicted, the cache data to be evicted is written back to the memory, and new cache data is received again.

8. The method according to claim 5, characterized in that The method further comprises: When it is detected that the first cache is in a full state, cache data to be evicted is determined in the cache data according to the usage frequency of the processing core, the cache data to be evicted is sent to the memory, and new cache data is received again.

9. The method according to claim 5, characterized in that In the case where it is detected that the request transaction is a read request transaction, determining whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction, and if so, replying to the request transaction based on the cache data stored in the first cache and a preset cache consistency maintenance table includes: When it is detected that the read request transaction is the first read request transaction and the initial state of the cache corresponding to the master node is a shared dirty state, it is determined whether the local cache directory corresponding to the first cache hits the target address carried by the read request transaction. If so, a reply response is directly made to the first read request transaction.

10. A communication device, characterized in that: include: A transceiver, a memory, a processor, and a program stored on the memory and executable on the processor; The processor is used to read the program in the memory to implement the multi-core on-chip network transaction processing method as described in any one of claims 5-9.

11. A readable storage medium for storing a program, characterized in that: When the program is executed by the processor, the multi-core on-chip network transaction processing method as described in any one of claims 5 to 9 is implemented.

12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by the processor, the multi-core on-chip network transaction processing method as described in any one of claims 5-9 is implemented.

Citation Information

Patent Citations

  • Data processing method of multi-core processor, server, product and medium

    CN118838863A