Virtual channel partitioning strategy based on flow type under CPU-GPU heterogeneous network-on-chip architecture

By adopting traffic-based virtual channel partitioning strategy and differentiated routing algorithms in CPU-GPU heterogeneous on-chip networks, traffic congestion and resource contention problems caused by high load are solved, and more efficient network performance and system efficiency are achieved.

CN119988000APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411956936.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the CPU-GPU heterogeneous on-chip network architecture, existing research is difficult to effectively solve the problems of traffic congestion and resource contention caused by high load.

Method used

The virtual channel partitioning strategy based on traffic type is adopted, and the input buffer is divided into the available partitions for requesting traffic and the available partitions for replying traffic. Different routing algorithms are used in different traffic partitions to optimize the allocation ratio of virtual channels to adapt to the network traffic mode.

Benefits of technology

It effectively alleviates traffic congestion and resource contention problems, reduces local link pressure, improves overall network performance, and improves system efficiency by optimizing the partition proportion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988000A_ABST
    Figure CN119988000A_ABST
Patent Text Reader

Abstract

The invention designs a virtual channel partitioning strategy based on a flow type under a CPU-GPU heterogeneous network-on-chip architecture. Under an NoC architecture, aiming at different characteristics of request traffic and reply traffic in a transmission path, a virtual channel is divided into a request traffic partition and a reply traffic partition according to traffic types, and the problems of traffic congestion and resource contention caused by a high load condition are relieved. Meanwhile, a differentiated routing strategy is adopted, independent transmission paths are designed for request flow and reply flow respectively, and the local hot spot phenomenon is relieved. Besides, aiming at the condition that the flow proportion is unbalanced, the optimal partition strategy is sought by changing the proportion of different partitions of the virtual channel, so that the optimal partition strategy is more adaptive to the current network flow mode, and the system performance is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer architecture and high performance computing, and specifically designs a virtual channel partitioning strategy based on traffic type in a CPU-GPU heterogeneous on-chip network architecture. Background Art

[0002] In recent years, with the continuous emergence and development of new generation information technologies such as artificial intelligence, the Internet of Things, and cloud computing, the computing power demand of chips has ushered in exponential growth. The traditional bus structure has gradually shown its limitations in multi-core processors and complex computing tasks. When the number of cores in the system increases, multiple devices simultaneously requesting access to the bus will cause competition, resulting in serious communication bottlenecks and making it difficult to meet the diverse computing power requirements. In order to solve the problem of inefficient inter-core communication in multi-core processors, the network-on-chip (NoC) has gradually developed as an efficient and scalable communication method. By adopting distributed routers and network topology structures, NoC connects nodes such as various processing cores, last-level caches (LLCs), and memory controllers (MCs). Data is divided into data packets and transmitted between nodes through links, ensuring low latency and high throughput in high-concurrency computing tasks.

[0003] In order to adapt to the needs of heterogeneous computing architecture and meet the different characteristics of CPU and GPU in communication mode, performance requirements and resource usage, CPU-GPU heterogeneous processors have begun to become the mainstream architecture in on-chip networks. By integrating the CPU's ability to perform complex tasks and the GPU's ability to process in parallel, this architecture can better adapt to the needs of complex and diversified computing tasks and support large-scale data interaction in multi-core and heterogeneous environments. However, a large number of shared resources will be competed for by multiple data flows at the same time in high-load scenarios, which will lead to traffic congestion and bring about resource contention issues that cannot be ignored.

[0004] Existing research generally optimizes the performance of heterogeneous NoC from two aspects: data flow mechanism and micro-architecture. On the one hand, it draws on the traditional homogeneous NoC optimization method to improve the overall throughput by designing fault-tolerant routing algorithms and traffic adaptive scheduling mechanisms. However, in networks for large-scale heterogeneous computing cores, due to the significant differences in memory access characteristics of different cores, traditional homogeneous NoC routing algorithms are usually difficult to adapt to heterogeneous environments, and the effect of reducing resource contention and conflict is relatively limited.

[0005] On the other hand, the performance of NoC is optimized from the perspective of the router microarchitecture, which includes input buffers, virtual channel allocators, switch arbitrators, and crossbar switches. During traffic transmission, data packets first enter the input buffer through the port and are temporarily stored for subsequent processing. Subsequently, the virtual channel allocator allocates appropriate virtual channels for each data packet. The switch allocator then selects data packets for transmission from multiple competing input flows according to the priority and arbitration rules and the predefined routing algorithm. Finally, the crossbar switch completes the path switching, transfers the data packet from the input port to the output port, and injects it into the corresponding port buffer of the next router to complete a communication process. Existing research generally optimizes NoC from two aspects: the virtual channel allocation method of static routing and the dynamic allocation of router buffers. However, in existing research, the interconnected networks are mostly symmetrically distributed, and a relatively simple average allocation method is used for buffer resources. With the development of large-scale on-chip networks, traffic distribution is no longer balanced, and router buffer resources need to be more reasonably allocated to computing nodes of different properties. Therefore, the optimization of large-scale on-chip network communication with complex traffic load conditions has become an urgent problem to be solved. Summary of the invention

[0006] The present invention proposes a virtual channel partitioning strategy based on traffic type in a CPU-GPU heterogeneous on-chip network architecture.

[0007] The present invention proposes a virtual channel partitioning strategy, which aims to solve the traffic congestion and resource contention problems caused by high load in NoC. A virtual channel is an independent queue with a fixed length in the input buffer of a router. By dividing the input buffer into multiple virtual channels, the head-of-line blocking problem can be alleviated, thereby effectively reducing traffic congestion. The present invention realizes the parallel transmission of request type traffic and reply type traffic through virtual channel partitioning technology. The request traffic flows from the CPU / GPU to the MC or LLC, while the reply traffic is returned from the MC and LLC to the computing core, carrying the corresponding calculation or storage results. In view of the different characteristics of the request traffic and the reply traffic in the transmission path, the virtual channel is divided into a request traffic available partition and a reply traffic available partition by type, and isolation is achieved at the virtual channel level. In addition, in order to alleviate the local hot spot problem, the present invention adopts differentiated routing algorithms in different traffic partitions, and reduces the path overlap and resource contention problems that may occur in the network topology by designing independent transmission paths for the request traffic and the reply traffic, dispersing the traffic load in the hot spot area, and reducing the pressure of the local link, thereby improving the overall network performance. Finally, in the case of unbalanced traffic proportion, the fixed uniform allocation strategy of virtual channels may lead to insufficient channel utilization or overload. Therefore, the present invention optimizes and adjusts the allocation ratio of virtual channels, and optimizes the allocation ratio by observing and analyzing the load of request traffic and reply traffic, so that it is more suitable for the current network traffic mode, thereby further improving system performance.

[0008] The method of the present invention and its implementation principle are described in detail below.

[0009] Step 1: Identify traffic types based on traffic characteristics.

[0010] In the NoC architecture, the traffic in the network can be divided into two types, namely request traffic and reply traffic. Request traffic carries computing requests or storage read instructions, and flows from the CPU / GPU to the memory controller and the last-level cache; reply traffic is the data packet returned by the memory controller or the last-level cache in response to the request, which contains the calculation results or data reading results. During the communication process, each node continuously injects data packets into the network. The core function of the traffic type identification module is to quickly parse the data packet header information. Through the built-in optimized retrieval mechanism, this module can efficiently extract the key identifier fields of the data packet, including the source address, destination address and operation type information. Subsequently, according to the predefined discrimination rules, the traffic type identification module classifies the data packet into request traffic or reply traffic. The predefined discrimination rules are as follows:

[0011] (1) When the operation type is identified as "read instruction" or "write request" and the destination address is the memory controller or the last-level cache, the traffic identification module identifies the data packet as request traffic.

[0012] (2) When the operation type is identified as "read response" or "write confirmation" and the source address is the memory controller or the last-level cache, the traffic identification module identifies the data packet as reply traffic.

[0013] The schematic paths of the two types of traffic flows are as follows: Figure 1 As shown in the figure. In the NoC architecture, the traffic in the network can be divided into two types, namely request traffic and reply traffic. Request traffic carries computing requests or storage read instructions, and flows from the CPU / GPU to the MC and LLC target nodes; reply traffic is the data packet returned by the MC or LLC in response to the request, which contains the computing results or data read results. On the transmission path, these two types of traffic show different characteristics, such as the distribution of traffic load and the concentration of communication direction.

[0014] In order to achieve accurate traffic type identification, it is necessary to first analyze the characteristics of the data packets in the network. The local port of each node in the NoC can initiate or receive traffic and transmit it through the external port. The present invention monitors the header information of the data packet, extracts specific identifiers (such as source node address, destination node address and operation type information), can clarify the attributes of the data packet, and divide it into request traffic or reply traffic according to predefined rules. In order to ensure real-time performance, the system has designed a set of efficient traffic identification mechanisms, which can quickly complete type discrimination when the data packet enters the router, laying the foundation for subsequent virtual channel partitioning and routing calculation.

[0015] Step 2. Implement virtual channel partition design.

[0016] In order to achieve effective management of request traffic and reply traffic, the present invention divides virtual channels into two categories, namely, request traffic available partitions and reply traffic available partitions. These virtual channels achieve independent support for different traffic types through logical isolation, which can reduce the possibility of resource contention caused by traffic congestion.

[0017] In the design of virtual channel partitions, the present invention makes targeted improvements to the virtual channel distributor. The improved distributor maps request traffic and reply traffic to independent virtual channel partitions based on the identification result of the traffic type in the previous step. This improvement can ensure the independent storage and transmission of the two types of traffic, and also lays the foundation for the implementation of differentiated routing strategies.

[0018] Step 3. Implement virtual channel partitioning and differentiated routing design.

[0019] In order to alleviate the local hot spot problem and further improve network performance, the present invention adopts different routing algorithms in the request traffic partition and the reply traffic partition. This design optimizes the traffic path selection in a targeted manner, directs different types of traffic to different network areas, and alleviates the resource contention and congestion problems caused by traffic concentration.

[0020] In the specific implementation, the present invention adopts the classic XY routing algorithm for request traffic. This algorithm is based on determinism. The data packet is preferentially transmitted along the X-axis and then along the Y-axis, ensuring that the request traffic can reach the target node along the shortest path, thereby reducing the transmission delay. For reply traffic, the YX routing algorithm is adopted, and its algorithm characteristics are similar to XY routing but the routing order is opposite. The data packet is preferentially transmitted along the Y-axis and then along the X-axis. Through this reverse routing method, direct conflicts between request traffic and reply traffic on the transmission path can be avoided, further improving the parallelism of the routing process. By adopting differentiated routing algorithms in different traffic partitions, the present invention effectively balances the global traffic distribution, can reduce the impact of local link congestion on system performance, and provide a more stable basic network environment for subsequent virtual channel partition optimization.

[0021] Step 4. Optimize virtual channel partition ratio based on traffic load.

[0022] The present invention uses the spec 2006 CPU benchmark suite and some CUDA GPGPU benchmark suites as benchmarks for experimental operation. The proportions of the two types of traffic obtained after running multiple sets of CPU-GPU hybrid benchmarks are as follows: Figure 2 As shown in the figure, the average proportion of the total data of the reply traffic type can reach 71.6%, while the average proportion of the total data of the request traffic type is only 28.4%. Inspired by the phenomenon of uneven proportion of traffic types, the design of evenly distributing the number of virtual channels to the two traffic types is not the optimal solution. In addition, considering the uneven proportion of the two traffic types is not the only factor affecting the performance of the virtual channel partitioning strategy. In CPU-GPU heterogeneous NoC, the impact of the partitioning strategy performance may also be related to the following two traffic characteristics:

[0023] (1) Traffic generation order: For the same task, request traffic often occurs before reply traffic.

[0024] (2) Differences in traffic behavior characteristics: Request traffic and reply traffic have different behavior characteristics in the network. Generally speaking, request traffic packets tend to have a small amount of data, but a large number and a fast transmission speed. Reply traffic packets have a relatively small number of data packets, but the data volume of each packet is large.

[0025] Therefore, after the virtual channel partition design is completed, the present invention proposes a statically adjusted partition ratio strategy based on experimental analysis and theoretical calculation, and seeks the optimal partition strategy by changing the ratio of different virtual channel partitions to optimize the utilization of links and nodes. This method does not require complex dynamic monitoring and real-time adjustment mechanisms, has low computational complexity and low hardware requirements.

[0026] Beneficial Effects

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1. The present invention proposes a virtual channel partitioning strategy, which avoids resource contention and congestion problems caused by traffic mixing by partitioning and isolating request traffic and reply traffic.

[0029] 2. The present invention adopts differentiated routing algorithms in different traffic partitions, disperses local hotspot traffic through path selection, thereby reducing link conflicts and improving overall network performance.

[0030] 3. The present invention aims at the traffic proportion characteristics of different scenarios, explores the differences in the overall network performance improvement of virtual channel design under different partition ratios, and determines the optimal partition ratio. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Schematic diagram of two types of traffic paths;

[0032] Figure 2 This is a schematic diagram of the proportion of the two types of traffic;

[0033] Figure 3 This is one of the schematic diagrams of the system model of the present invention;

[0034] Figure 4 This is the second schematic diagram of the system model of the present invention;

[0035] Figure 5 This is a schematic diagram of the virtual channel partitioning strategy;

[0036] Figure 6 This is one of the performance comparison charts.

[0037] Figure 7 This is the second performance comparison chart. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples.

[0039] The system model used in the example is as follows:

[0040] System model such as Figure 3and Figure 4 As shown in the figure, each node in the 5×5 two-dimensional mesh structure of the CPU-GPU heterogeneous on-chip network is connected to a router. The message generated by each node will be divided into packets of fixed size and transmitted in the network through the router. Each router is equipped with 5 input ports, one of which is an internal port for receiving traffic from local nodes, and the other four are external ports, which are connected to adjacent routers through links to transmit traffic. Each input port is equipped with an independent input buffer, which is further divided into 4 virtual channels to optimize traffic transmission and alleviate network congestion.

[0041] The present invention designs a virtual channel partitioning strategy based on traffic type in a CPU-GPU heterogeneous on-chip network architecture. In the NoC architecture, by dividing the virtual channel into request traffic partition and reply traffic partition according to traffic type, the traffic congestion and resource contention problems caused by high load conditions are alleviated. By adopting virtual channel isolation and differentiated routing strategies, independent transmission paths are designed for request traffic and reply traffic, respectively, to alleviate local hot spots and improve overall network performance. In addition, in the case of unbalanced traffic proportion, the optimal partitioning strategy is sought by changing the proportion of different partitions of the virtual channel to further improve system efficiency.

[0042] The specific steps are as follows:

[0043] Step 1: Design a traffic type identification module based on the traffic type.

[0044] In this embodiment, each node continuously injects data packets into the network, and the data packets contain information about its source node, destination node, and operation type. When the data packet arrives at the router, the traffic identification module begins to parse its header information, and extracts the source address, destination address, and operation type information by quickly retrieving the identifier field of the data packet. Subsequently, according to the predefined discrimination rules, the traffic identification module classifies the data packet as request traffic or reply traffic. The predefined discrimination rules are as follows:

[0045] (1) When the operation type is identified as "read instruction" or "write request" and the destination address is MC or LLC, the traffic identification module identifies the data packet as request traffic.

[0046] (2) When the operation type is identified as "read response" or "write confirmation" and the source address is MC or LLC, the traffic identification module identifies the data packet as reply traffic.

[0047] Step 2: Based on the traffic type, optimize the virtual channel allocator to implement the partitioning function.

[0048] In order to achieve effective management of request traffic and reply traffic, the virtual channels are divided into two categories by optimizing the virtual channel distributor, namely, request traffic available partitions and reply traffic available partitions. These virtual channels achieve independent support for different traffic types through logical isolation. In this embodiment, after receiving a data packet, the virtual channel distributor of each router will divide the data packet into the corresponding virtual channel partition according to the traffic type identifier extracted from the data packet header information. The request traffic is allocated to the request traffic partition, and the reply traffic is allocated to the reply traffic partition. The partitioned virtual channels are managed and stored through independent queues to avoid the situation where the two types of traffic compete for resources in the buffer.

[0049] Specifically, the partitioning function is realized by designing a mapping rule table. The optimized virtual channel allocator is based on the state machine logic. It first reads the key identifier from the packet header and quickly matches the traffic type (the related work of step 1); then calls the mapping rule table to locate the corresponding virtual channel partition. During the mapping process, the specific mapping rules of the virtual channel are: Step 2. Based on the traffic type, optimize the virtual channel allocator to realize the partitioning function.

[0051] (1) Initialization of virtual channel status: For the buffers of the four virtual channels, the status values ​​of the first two virtual channels are set to “0”, marking that request traffic is available; the status values ​​of the last two virtual channels are set to “1”, marking that reply traffic is available.

[0052] (2) Traffic-aware allocation logic: After the traffic identification module detects the traffic type, the allocator automatically maps the request traffic to the virtual channel with a status value of "0" according to the virtual channel status identifier; and maps the reply traffic to the virtual channel with a status value of "1".

[0053] Step 3: Based on the traffic type, the routing calculation unit determines the traffic transmission rule.

[0054] It includes the following steps:

[0055] S1: Traffic type synchronization phase.

[0056] Establish a connection between the traffic identification module and the routing calculation unit: After determining the traffic type, the traffic identification module synchronizes this information with the routing calculation unit to ensure that the routing calculation unit can make different routing decisions based on the traffic type.

[0057] S2: Routing decision selection phase.

[0058] Design a routing adaptive selection mechanism: The routing calculation unit automatically selects the corresponding routing algorithm based on the traffic type information, specifically using the XY routing algorithm for request traffic and the YX routing algorithm for reply traffic. This prevents different types of traffic from being transmitted on the same path, thereby alleviating local hot spots and traffic congestion.

[0059] S3: Execute routing decision phase.

[0060] The routing calculation unit passes the calculated path information to the router, and the crossbar switch in the router completes the path switching of the data packet from the input port to the output port, thereby injecting the data packet into the buffer of the corresponding port of the next router, thus completing the entire communication process.

[0061] Step 4: Change the virtual channel partitioning rules to find the optimal partitioning strategy.

[0062] The factors that affect the performance of partitioning strategies are as follows:

[0063] (1) Uneven traffic proportion: The total amount of data of the two traffic types is different, and the proportion of request traffic is smaller than that of reply traffic.

[0064] (2) Traffic generation order: For the same task, request type traffic is often generated before reply type traffic.

[0065] (3) Differences in traffic behavior characteristics: Request traffic and reply traffic have different behavior characteristics in the network. Request traffic type packets often have a small amount of data, but a large number and a fast transmission speed. Reply traffic type packets have a relatively small number of packets, but the data volume of each packet is large.

[0066] After the virtual channel partition design is completed, a statically adjusted partition ratio strategy is proposed based on theoretical analysis and the optimal partition ratio is determined through experiments. Two authoritative test programs, spec 2006 CPU benchmark suite and CUDA GPGPU benchmark suite, are used as benchmarks for experimental operation. In addition, the IPC of the CPU core and the network latency of the on-chip network are used as performance indicators, and their calculation methods are shown in formulas (1) and (2) respectively.

[0067]

[0068] For Formula 1, instruction is the total number of instructions during the benchmark run, cycles is the total number of cycles during the benchmark run, and IPC is the ratio of the total number of instructions to the total number of cycles. It is used to measure the execution efficiency of the processor and is an important indicator for evaluating computing performance. For Formula 2, n is the total number of processing cores on the NoC, and Latency is the ratio of the sum of the latencies of each core to the number of cores. It reflects the average delay of data transmission between tasks and is an important manifestation of network performance. By redesigning the mapping rules of the virtual channel partitions in step 2 and experimentally evaluating the impact of three different partition ratios on system performance, the optimal virtual channel static partitioning strategy is explored and determined. The redesigned virtual channel partition mapping rules are:

[0069] (1) Initialization of virtual channel status: For the buffers of the four virtual channels, the status values ​​of the first n virtual channels are set to “0”, marking the request traffic as available; the status values ​​of the last 4-n virtual channels are set to “1”, marking the reply traffic as available.

[0070] (2) Traffic-aware allocation logic: After detecting the traffic type, the traffic identification module automatically maps the request traffic to the virtual channel with a status value of "0" according to the virtual channel status identifier; and maps the reply traffic to the virtual channel with a status value of "1".

[0071] During the experimental evaluation, we set n=1, n=2, and n=3 to run the same hybrid benchmark group and calculate the IPC and latency under different partition ratios. Figure 6 and Figure 7 As shown, compared with the benchmark test, the three different partitioning strategies reduce latency by 3.06%, 11.2% and 15.5% respectively while ensuring that the IPC performance remains basically unchanged. The experimental results show that the virtual channel partitioning scheme using a 3-request traffic-1-reply traffic ratio is superior to the other two partitioning methods in performance, and significantly reduces network latency while keeping the IPC performance basically unchanged, verifying its advantages in resource utilization and traffic management. Finally, based on the experimental data, the present invention determines the optimal virtual channel partitioning strategy of 3-request traffic partitioning-1-reply traffic partitioning.

Claims

1. A virtual channel partitioning strategy based on traffic type in a CPU-GPU heterogeneous on-chip network architecture, characterized by The following steps are involved: Step 1. Design traffic type identification based on traffic characteristics; In the NoC architecture, the traffic in the network is divided into two types, namely request traffic and reply traffic; request traffic carries computing requests or storage read instructions and flows from the CPU / GPU to the memory controller and last-level cache; Reply traffic is the data packet returned by the memory controller or last-level cache in response to the request, which contains the calculation results or data reading results. During the communication process, each node continuously injects data packets into the network. The core function of the traffic type identification module is to parse the packet header information. The module uses a built-in optimized retrieval mechanism to extract the key identifier fields of the packet, including the source address, destination address, and operation type information. Then, according to the predefined discrimination rules, the traffic type identification module classifies the packet into request traffic or reply traffic. The predefined discrimination rules are as follows: (1) When the operation type is identified as "read instruction" or "write request" and the destination address is the memory controller or the last-level cache, the traffic identification module identifies the data packet as request traffic; (2) When the operation type is identified as "read response" or "write confirmation" and the source address is a memory controller or a last-level cache, the traffic identification module identifies the data packet as reply traffic; Step 2. Based on the traffic type, optimize the virtual channel allocator to implement the partitioning function; By optimizing the virtual channel allocator, virtual channels are divided into two categories: request traffic availability partitions and reply traffic availability partitions. These virtual channels can independently support different traffic types through logical isolation. Specifically, the partitioning function is realized by designing a mapping rule table. The optimized virtual channel allocator is based on the state machine logic. It first reads the key identifier from the packet header to quickly match the traffic type. Then it calls the mapping rule table to locate the corresponding virtual channel partition. During the mapping process, the specific mapping rules of the virtual channel are: (1) Initialization of virtual channel status: For the buffers of the four virtual channels, set the status values ​​of the first two virtual channels to "0", marking that request traffic is available; set the status values ​​of the last two virtual channels to "1", marking that reply traffic is available; (2) Traffic-aware allocation logic: After the traffic identification module detects the traffic type, the allocator automatically maps the requested traffic to the virtual channel with a status value of "0" based on the virtual channel status identifier; Map the reply traffic to the virtual channel with a status value of "1"; Step 3. Implement virtual channel partitioning differentiated routing design; By optimizing the routing calculation unit, different routing algorithms are used in the request traffic partition and the reply traffic partition respectively; The specific implementation includes the following steps: S1: Traffic type synchronization phase; Establishing a connection between the traffic identification module and the routing calculation unit: After determining the traffic type, the traffic identification module synchronizes this information with the routing calculation unit to ensure that the routing calculation unit can make different routing decisions based on the traffic type; S2: Routing decision selection stage; Design a routing adaptive selection mechanism: The routing calculation unit automatically selects the corresponding routing algorithm based on the traffic type information, specifically, the request traffic uses the XY routing algorithm, and the reply traffic uses the YX routing algorithm; S3: Execute routing decision phase; The routing calculation unit passes the calculated path information to the router, and the crossbar switch in the router completes the path switching of the data packet from the input port to the output port, thereby injecting the data packet into the buffer of the corresponding port of the next router, thus completing the entire communication process; Step 4. Optimize virtual channel partition ratio based on traffic load; After the virtual channel partition design is completed, the IPC of the CPU core and the network latency of the on-chip network are used as performance indicators, and their calculation methods are shown in formula (1) and formula (2) respectively; For formula 1, instruction is the total number of instructions during the benchmark run, cycles is the total number of cycles during the benchmark run, and IPC is the ratio of the total number of instructions to the total number of cycles, which is used to measure the execution efficiency of the processor and is an important indicator for evaluating computing performance. For formula 2, n is the total number of processing cores on the NoC, and Latency is the ratio of the sum of the latency of each core to the number of cores, reflecting the average latency of data transmission between tasks. By redesigning the mapping rules of the virtual channel partitions in step 2, the optimal virtual channel static partitioning strategy is explored and determined. The redesigned virtual channel partition mapping rules are: (1) Initialization of virtual channel status: For the buffers of the four virtual channels, set the status values ​​of the first n virtual channels to "0", marking that the request traffic is available; set the status values ​​of the last 4-n virtual channels to "1", marking that the reply traffic is available; (2) Traffic-aware allocation logic: After detecting the traffic type, the traffic identification module automatically maps the request traffic to the virtual channel with a status value of "0" according to the virtual channel status identifier; and maps the reply traffic to the virtual channel with a status value of "1"; The value of n is 1, 2 or 3.

2. The strategy according to claim 1, characterized in that: The value of n is 3.