A method and apparatus for determining performance, a computer device, and a storage medium

By determining the computing nodes and computing link traffic in the on-chip network, combined with the ideal throughput, the problem of low accuracy in on-chip network performance determination in the prior art is solved, and higher performance information accuracy and high performance requirements of neural network accelerators are achieved.

CN113900917BActive Publication Date: 2025-05-27SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111165419.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-05-27
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

The existing on-chip network performance determination methods rely on delay statistics such as a single communication total hop, resulting in low accuracy and cannot meet the high latency or high throughput requirements of neural network accelerators.

Method used

By determining multiple computing nodes in the target neural network, the link traffic of the physical communication link between the data processing cores is calculated based on the correlation information between the computing nodes, the topology structure and mapping location of the on-chip network, and the performance information of the on-chip network is determined based on the ideal throughput.

Benefits of technology

It improves the accuracy of on-chip network performance information and can more effectively meet the high-performance needs of neural network accelerators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113900917B_ABST
    Figure CN113900917B_ABST
Patent Text Reader

Abstract

The present disclosure provides a performance determination method, apparatus, computer device, and storage medium. Among them, the method includes: determining a plurality of computing nodes in a target neural network of an on-chip network mapping to be evaluated; determining the link traffic of the physical communication links between data processing cores in the on-chip network based on the association relationship information between the plurality of computing nodes, the topology between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determining the ideal throughput when performing the data processing task corresponding to the target neural network using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network; determining the performance information when running the target neural network using the on-chip network based on the link traffic and the ideal throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, an apparatus, a computer device, and a storage medium for performance determination. Background Art

[0002] With the increasing richness of the functional requirements of the System-on-Chip (SoC) and the continuous rise of the design complexity, on-chip interconnection has become a challenge in chip design. The Network-on-Chip (NoC) has become the primary interconnection architecture of the SoC, especially for multi-core SoCs, due to its high performance, high scalability, and other advantages.

[0003] The optimization objectives of traditional on-chip network mapping methods often focus on delay statistics such as the total number of communication hops. Such optimization objectives often cannot meet the requirements of neural network applications. On the one hand, neural network accelerators often pursue high latency or high throughput. How to reasonably design the optimization objectives of the mapping method to meet the requirements of neural network accelerators; on the other hand, the complex hardware structure and communication protocol of the on-chip network make it difficult to accurately model the on-chip network communication. Therefore, how to determine the performance of the on-chip network so that the obtained performance information is close to the simulation results. However, many existing technologies still often use single delay statistics such as the total number of communication hops to determine the performance information of the on-chip network; this way of determining performance information has the problem of low accuracy. Summary of the Invention

[0004] The embodiments of the present disclosure at least provide a method, an apparatus, a computer device, and a storage medium for performance determination.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for performance determination, including: determining a plurality of computing nodes in a target neural network of an on-chip network mapping to be evaluated; determining the link traffic of physical communication links between data processing cores in the on-chip network based on the association relationship information between the plurality of computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determining the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network; determining the performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput.

[0006] In an alternative embodiment, determining the link traffic of the physical communication links between the data processing cores in the on-chip network based on the association relationship information between multiple computing nodes, the topological structure between the data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network includes: traversing the computing nodes corresponding to other network layers except the last network layer, for each network layer, determining the computing nodes in this network layer that have data flow interaction with the computing nodes in the adjacent network layer, using the computing nodes existing in this network layer as the transmission starting points, and using the computing nodes in the adjacent network layer corresponding to this computing node as the transmission end points; based on the mapping positions of the determined transmission starting points and corresponding transmission end points on the on-chip network respectively, and the topological structure, determining the logical communication links between each transmission starting point and the corresponding transmission end point respectively; based on the data transmission volume between different network layers, determining the link traffic of the logical communication links between each transmission starting point and the corresponding transmission end point; and determining the link traffic of the physical communication links between the data processing cores in the on-chip network according to the link traffic of the logical communication links between each transmission starting point and the corresponding transmission end point.

[0007] In an alternative embodiment, the determining the logical communication links between each transmission starting point and the corresponding transmission end point respectively based on the mapping positions of the determined transmission starting points and corresponding transmission end points on the on-chip network respectively, and the topological structure includes: for each transmission starting point, based on the mapping positions of this transmission starting point and the corresponding transmission end point on the on-chip network respectively, determining the first data processing core mapped by this transmission starting point on the on-chip network and the second data processing core mapped by the corresponding transmission end point on the on-chip network; based on the topological structure, determining the physical communication link connecting the first data processing core and the second data processing core; and generating the logical communication link between this transmission starting point and the corresponding transmission end point based on the physical communication link connecting the first data processing core and the second data processing core.

[0008] In an alternative embodiment, the determining the link traffic of the physical communication links between the data processing cores in the on-chip network according to the link traffic of the logical communication links between each transmission starting point and the corresponding transmission end point includes: for each physical communication link between the data processing cores in the on-chip network, based on the logical communication links between each transmission starting point and the corresponding transmission end point, determining the target logical communication link including this physical communication link; and determining the link traffic of this physical communication link based on the link traffic corresponding to the target logical communication link.

[0009] In an alternative embodiment, determining the ideal throughput when using the network-on-chip to execute the data processing task corresponding to the target neural network based on the association relationship information and the data computation amounts respectively corresponding to each network layer in the target neural network includes: for each network layer of the target neural network, determining the data computation time respectively corresponding to each computing node in this network layer according to the computing node corresponding to this network layer and the data computation amount corresponding to this network layer; determining the ideal throughput based on the data computation times respectively corresponding to multiple computing nodes.

[0010] In an alternative embodiment, determining the ideal throughput based on the data computation times respectively corresponding to multiple computing nodes includes: determining the longest data computation time from the data computation times respectively corresponding to multiple computing nodes; determining the ideal throughput corresponding to a preset time duration based on the longest data computation time.

[0011] In an alternative embodiment, determining the performance information when using the network-on-chip to run the target neural network based on the link traffic and the ideal throughput includes: for each physical communication link in the network-on-chip, determining the required link bandwidth corresponding to this physical communication link based on the link traffic corresponding to this physical communication link and the ideal throughput; determining the loss rate respectively corresponding to each physical communication link based on the required link bandwidths respectively corresponding to each physical communication link in the network-on-chip and the actual link bandwidth; determining the performance information based on the loss rates respectively corresponding to each physical communication link.

[0012] In an alternative embodiment, determining the performance information based on the loss rates respectively corresponding to each physical communication link includes: determining the maximum loss rate among the loss rates respectively corresponding to each physical communication link as the network-on-chip loss rate; in response to the network-on-chip loss rate being greater than a preset loss rate threshold, determining the performance information of the network-on-chip as communication limited; in response to the network-on-chip loss rate being less than the preset loss rate threshold, determining the performance information of the network-on-chip as computation limited; in response to the network-on-chip loss rate being equal to the preset loss rate threshold, determining the performance information of the network-on-chip as full-load transmission.

[0013] In a second aspect, an embodiment of the present disclosure further provides a performance determination device, including: a first determination module, configured to determine a plurality of computing nodes in a target neural network of an on-chip network mapping to be evaluated; a processing module, configured to determine the link traffic of physical communication links between data processing cores in the on-chip network based on the association relationship information between the plurality of computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determine the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network; a second determination module, configured to determine performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput.

[0014] In a third aspect, an alternative implementation of the present disclosure further provides a computer device, including a processor and a memory. The memory stores machine-readable instructions executable by the processor. The processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions execute the steps in the first aspect or any possible implementation manner in the first aspect.

[0015] In a fourth aspect, an alternative implementation of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run, it executes the steps in the first aspect or any possible implementation manner in the first aspect.

[0016] For the effect descriptions of the above performance determination device, computer device, and computer-readable storage medium, refer to the description of the above performance determination method, which will not be elaborated here.

[0017] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the accompanying drawings. Description of the Drawings

[0018] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the embodiments will be briefly introduced below. The accompanying drawings are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to explain the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 The figure shows a flowchart of a performance determination method provided by an embodiment of the present disclosure;

[0020] Figure 2 The figure shows a schematic diagram of a computing node and a network-on-chip provided by an embodiment of the present disclosure;

[0021] Figure 3 The figure shows a specific flowchart for specifically determining link traffic provided by an embodiment of the present disclosure;

[0022] Figure 4 The figure shows a schematic diagram of a performance determination device provided by an embodiment of the present disclosure;

[0023] Figure 5 The figure shows a schematic diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The components of the embodiments of the present disclosure described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0025] Based on the above research on the background technology, the present disclosure provides a performance determination method. After determining a plurality of computing nodes corresponding to a target neural network mapped on a network-on-chip, on this premise, by determining the link traffic of each physical communication link in the network-on-chip and the ideal throughput when using the network-on-chip to execute the data processing task corresponding to the target neural network, it is possible to determine whether the link traffic that the network-on-chip can carry can support the transmission of the ideal throughput when actually executing the data processing task, so as to judge the performance information of the network-on-chip when running the target neural network. In this way, compared with the method of solely using delay statistical information such as the total number of communication hops to determine the performance of the network-on-chip, the obtained performance information has higher accuracy.

[0026] Regarding the defects existing in the above solutions, they are all the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should both be the contributions made by the inventors to the present disclosure during the process of the present disclosure.

[0027] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0028] For ease of understanding of this embodiment, first, a performance determination method disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the performance determination method provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device or a server or other processing devices. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the performance determination method can be implemented by a processor invoking computer-readable instructions stored in a memory.

[0029] The performance determination method provided in the embodiments of the present disclosure will be described below.

[0030] See Figure 1 As shown, it is a flowchart of a performance determination method provided in the embodiments of the present disclosure. The method includes steps S101 to S104, where:

[0031] S101: Determine multiple computing nodes in the target neural network of the on-chip network mapping to be evaluated;

[0032] S102: Based on the association relationship information between the multiple computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network, determine the link traffic of the physical communication links between the data processing cores in the on-chip network;

[0033] S103: Based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network, determine the ideal throughput when using the on-chip network to execute the data processing tasks corresponding to the target neural network;

[0034] S104: Based on the link traffic and the ideal throughput, determine the performance information when using the on-chip network to run the target neural network.

[0035] In the case of determining the network-on-chip to be evaluated in the embodiments of the present disclosure, the link traffic of the communication links in the network-on-chip can be determined based on the association relationship information between multiple computing nodes mapped on the network-on-chip, the topological structure between data processing cores in the network-on-chip, and the mapping positions of the computing nodes on the network-on-chip. Additionally, based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network, the ideal throughput when performing the data processing task corresponding to the target neural network using the network-on-chip can be determined. Since the traffic of the communication links in the network-on-chip restricts the data throughput between the data processing cores corresponding to the computing nodes in the network-on-chip, the ideal throughput and the link traffic can be used to determine whether the link traffic that the network-on-chip can carry can support the transmission of the ideal throughput during the actual execution of the data processing task, thereby determining the performance information of the network-on-chip to be evaluated when running the target neural network. In this way, compared with the method of determining the performance of the network-on-chip solely using delay statistical information such as the total number of communication hops, the obtained performance information has higher accuracy.

[0036] The above S101 to S104 will be described in detail below.

[0037] Regarding the above S101, in one possible case when determining the network-on-chip to be evaluated, the target neural network can be first determined according to the actual data processing task requirements. Based on the computing tasks corresponding to each network layer of the target neural network, the target neural network is divided into multiple computing nodes, and then each computing node in the target neural network is mapped to different data processing cores in the system-on-chip, so that communication data streams are generated between different data processing cores due to the data flow relationship between different network layers in the target neural network, obtaining the network-on-chip to be evaluated (here, the system-on-chip with on-chip interconnection generated after mapping the target neural network is called the network-on-chip), and the performance determination method provided by the embodiments of the present disclosure is used to determine the performance information corresponding to the network-on-chip to be evaluated.

[0038] In a specific implementation, when determining the on-chip network to be evaluated in this case, first determine the target neural network according to the actual data processing task requirements. Among them, for the target neural network, in different application scenarios, the corresponding target neural network is also different. For example, in the image processing scenario, such as when applied to image recognition, the target neural network can include, for example, Convolutional Neural Networks (CNN), or Deep Neural Networks (DNN); in the natural semantic processing scenario, such as when applied to speech recognition, the target neural network can include, for example, the self-attention mechanism model (Transformer), or the Recurrent Neural Network (RNN). It can be specifically determined according to the actual situation and will not be elaborated here.

[0039] For the determined target neural network, it usually has multiple network layers; among them, corresponding to each network layer, there are multiple neurons. The data processing tasks processed by the multiple neurons in any one network layer are similar. Therefore, at least one computing node can be divided for any one network layer, and each computing node correspondingly undertakes the data processing tasks corresponding to some neurons in the corresponding network layer. For the multiple network layers of the target neural network, determine the corresponding computing nodes for the multiple network layers respectively, then the multiple computing nodes corresponding to the target neural network can be determined.

[0040] Among them, when determining the multiple computing nodes corresponding to the target neural network, for example, a random division method can be used to determine, or other division methods can be used to determine the multiple computing nodes for the target neural network. It can be specifically determined according to the actual situation and no limitation is made here.

[0041] The on-chip network includes a computing subsystem and a communication subsystem. Among them, the computing subsystem is composed of, for example, processing elements (PE) to complete computing tasks; in addition, the processing element can also be a central processing unit (CPU), or can be a System on Chip (SoC), or can also be an Intellectual Property core (IP core) with a dedicated function or a memory array, reconfigurable hardware, etc. The communication subsystem is responsible for connecting the processing elements to achieve high-speed communication between computing resources.

[0042] In the embodiments of the present disclosure, the processing engine in the computing subsystem is referred to as a data processing core, and the communication subsystem includes a communication router. That is, in the embodiments of the present disclosure, the system-on-chip includes multiple data processing cores, and each data processing core corresponds to a communication router. When mapping a target neural network onto the system-on-chip, for the obtained on-chip network, the data processing cores on the on-chip network correspondingly undertake the specific tasks of data processing, and the communication routers corresponding to the data processing cores complete the transfer of data between different data processing cores.

[0043] Specifically, when determining the mapping positions of the corresponding multiple computing nodes of the target neural network on the on-chip network obtained by mapping the target neural network onto the system-on-chip, a random determination method can also be correspondingly adopted, or other methods for determining the mapping positions can be correspondingly adopted, which are not limited herein.

[0044] In this way, when processing an actual data processing task, the performance determination method provided by the embodiments of the present disclosure can, when determining the node partitioning method of the computing nodes corresponding to the target neural network and the mapping positions of the computing nodes on the system-on-chip, randomly determine the performance information of the on-chip network obtained under the current node partitioning method and mapping positions, so as to determine whether the current node partitioning method and mapping positions are optimal, thereby providing targeted guidance for further adjusting the node partitioning method and mapping positions.

[0045] In another possible case, when determining the on-chip network to be evaluated, the already determined on-chip network can be used as the on-chip network to be evaluated. Specifically, the performance determination method provided by the embodiments of the present disclosure can directly determine the performance information of the on-chip network when running the target neural network, which is determined when the computing node partitioning method and the mapping positions of the computing nodes on the on-chip network have been determined. That is, for the already determined on-chip network, the performance determination method provided by the embodiments of the present disclosure can also be directly used to determine the corresponding performance information.

[0046] In this way, the performance determination method provided by the embodiments of the present disclosure can be applied to determine the corresponding performance information for any determined on-chip network, so the applicable range is wider.

[0047] Regarding the above S102, when determining the on-chip network to be evaluated, the mapping positions of the multiple computing nodes of the target neural network on the on-chip network, the correlation relationship information between the multiple computing nodes, and the topological structure between the data processing cores in the on-chip network can be correspondingly determined.

[0048] Among them, for the multiple computing nodes of the target neural network, by way of example, refer to Figure 2 As shown, it is a schematic diagram of a computing node and an on-chip network provided by the embodiments of the present disclosure. Among them, inFigure 2 In (a), the corresponding relationship between multiple computing nodes and the target neural network is shown. Among them, V i represents a computing node. In Figure 2 (a), V 0 to V 5 a total of 6 computing nodes are shown; N i represents a network layer. Exemplarily, N 1 represents the first network layer, and N 2 represents the second network layer. For Figure 2 the computing node V 1 shown in (a), its corresponding network layer is N 1 , that is, the first network layer. Accordingly, it can be determined that the first network layer N 1 corresponds to the computing node V 1 , the second network layer N 2 corresponds to the computing nodes V 0 and V 3 , the third network layer N 3 corresponds to the computing nodes V 2 and V 4 , and the fourth network layer N 4 corresponds to the computing node V 5 .

[0049] Regarding the mapping positions of multiple computing nodes on the on-chip network respectively, exemplarily, see Figure 2 shown in (b). Regarding Figure 2 the on-chip network shown in (b), it includes data processing cores u 0 to u 5 , and the communication routes R i corresponding to each data processing core u i respectively. In Figure 2 (b), the specific positions of the computing nodes V Figure 2 determined in (a) mapped to the data processing cores in the on-chip network are also marked accordingly. For example, the mapping position of the computing node V 0 to the data processing core u 5 in the on-chip network, the mapping position of the computing node V 0 to the data processing core u 0 in the on-chip network, the mapping position of the computing node V 3 to the data processing core u 1 in the on-chip network, etc.

[0050] Since multiple computing nodes have been determined, the association relationship information between multiple computing nodes can be determined accordingly. Among them, the association relationship information between multiple computing nodes can be determined, for example, according to the data flow direction between different computing nodes. Exemplarily, according to Figure 2The correspondence between the multiple computing nodes shown in (a) and the target neural network can be determined. The first-layer neural network N 1 corresponding computing node V 1 , and the processed result data corresponding thereto flows to the second-layer neural network N 2 corresponding computing node V 0 and V 3 , that is, there is an association relationship between computing node V 1 and computing node V 0 , and there is an association relationship between computing node V 1 and computing node V 3 . Since there is a determined data flow direction among the data computing nodes, the association relationship between the computing nodes included in the association relationship information among the multiple computing nodes can also be directed.

[0051] For the topological structure among the data processing cores in the network-on-chip, after determining the mapping positions of the computing nodes in the network-on-chip, since the data flow direction is determined when the computing nodes actually perform data processing tasks, the data processing cores will send data to other data processing cores or receive data sent by other data processing cores according to the task flow direction of the computing nodes during data processing. For example, as shown in Figure 2 (b), since the processed result data corresponding to computing node V 1 will be sent to computing node V 0 and computing node V 3 , the data processing core u 1 corresponding to computing node V 5 will, when specifically executing data processing tasks, have a logical connection relationship with the data processing core u 0 corresponding to computing node V 0 and the data processing core u 3 corresponding to computing node V 1 . In the embodiments of the present disclosure, this logical connection relationship among the data processing cores in the network-on-chip is referred to as the topological structure among the data processing cores.

[0052] Correspondingly, according to the association relationship information, the topological structure among the data processing cores in the network-on-chip, and the mapping positions of the computing nodes on the network-on-chip, the link traffic of the physical communication links among the data processing cores in the network-on-chip can be determined.

[0053] In a specific implementation, as shown in Figure 3 , it is a specific flowchart for specifically determining the link traffic provided by the embodiments of the present disclosure; wherein,

[0054] S301: Traverse the computing nodes corresponding to the network layers except the last network layer. For each network layer, determine the computing nodes in this network layer that have data flow interaction with the computing nodes of the adjacent network layer. Take the computing nodes existing in this network layer as the transmission starting points, and take the computing nodes of the adjacent network layer corresponding to this computing node as the transmission ending points.

[0055] Among them, since the computing nodes of the last network layer have completed all the computing steps of the target neural network after processing the received processed data, the last computing node does not need to send the corresponding processed result data to other computing nodes. That is, corresponding to the computing nodes of the last network layer, there are no computing nodes of the next adjacent network layer.

[0056] Here, "adjacent" refers to an adjacent network layer, that is, adjacent in the same traversal direction. It can be determined whether the adjacent network layer is the next adjacent network layer or the previous adjacent network layer according to the traversal order. If traversing starts from the input end, the determined adjacent network layer is the next adjacent network layer in the data flow direction. If traversing starts from the output end, the determined adjacent network layer is the adjacent network layer from which the data flow comes. The present disclosure does not limit the traversal direction.

[0057] In addition, according to the above specific description of the association relationship information, it can be known that two computing nodes with an association relationship correspond to adjacent network layers of the target neural network, and the association relationship between the two computing nodes is directed according to the data flow between the corresponding network layers. Therefore, for the computing nodes that can be traversed, the computing nodes of the adjacent network layer corresponding to the traversed computing node can be determined according to the association relationship information. Since there may be multiple computing nodes of the adjacent network layer corresponding to the network layer corresponding to the traversed computing node in the target neural network, for the traversed computing node, its corresponding adjacent network layer computing nodes may include multiple.

[0058] Exemplarily, take Figure 2 the multiple computing nodes shown in (a) and the mapping positions of the multiple computing nodes on the on-chip network as an example for description. When traversing the multiple computing nodes, for example, the computing nodes corresponding to the first network layer of the target neural network can be used as the first traversed computing nodes for traversal. That is, take the computing node V 1 corresponding to the first network layer as the first traversed computing node, and correspondingly take this computing node V 1 as the transmission starting point.

[0059] In this example, this computing node V 1The computing node belonging to the first - layer network layer, according to the association relationship information, the result data of the first - layer network layer will be used as the input data of the second - layer network layer and transmitted from the first - layer network layer to the second - layer network layer. Thus, related to computing node V 1 The corresponding adjacent network - layer computing nodes include the computing nodes corresponding to the second - layer network layer: V 0 and V 3 .

[0060] Here, the adjacent network - layer computing nodes corresponding to the traversed computing nodes can also be used as the transmission end - points. That is, when taking computing node V 1 as the transmission starting - point, the corresponding transmission end - points include computing node V 0 and V 3 .

[0061] Taking computing node V 0 as the next traversed computing node, and taking computing node V 0 as the transmission starting - point, the target computing nodes corresponding to computing node V 0 include: V 2 and V 4 ; that is, the transmission end - points include: V 2 and V 4 ;

[0062] ……

[0063] Taking computing node V 5 as the last traversed computing node, according to the association relationship information among multiple computing nodes, there is no computing - node pair with computing node V 5 as the transmission starting - point among the multiple computing nodes, and the traversal of the multiple computing nodes ends.

[0064] In this way, by determining to take computing node V 1 as the transmission starting - point and correspondingly determining to take computing node V 0 as the transmission end - point, the communication between the computing nodes corresponding to the adjacent two - layer network layers can be determined.

[0065] In addition, for other traversed computing nodes, the corresponding adjacent network - layer computing nodes can also be determined in a similar way. For example, for computing node V 0 in the second - layer network layer, taking it as the transmission starting - point, its corresponding adjacent network - layer computing nodes can include, for example, computing node V 2 or V 4 . In the case of taking computing node V 2 as the adjacent network - layer computing node, computing node V 2 is computing node V 0The corresponding transmission end point; when the computing node V 4 is used as an adjacent network layer computing node, the computing node V 4 is the transmission end point corresponding to the computing node V 0 That is, for the same traversed computing node, taking it as the transmission starting point, there can be multiple corresponding transmission end points.

[0066] It can be seen that in one embodiment, the transmission starting point and the transmission end point represent two computing nodes belonging to adjacent network layers and having a data flow interaction relationship.

[0067] S302: Based on the mapped positions of the determined transmission starting points and corresponding transmission end points on the on-chip network respectively, and the topological structure, determine the physical communication links between each transmission starting point and the corresponding transmission end point respectively.

[0068] It should be noted that the communication link between the transmission starting point and the corresponding transmission end point can be a logical communication link. Since the data processing cores where the transmission starting point and the corresponding transmission end point are located may not be directly connected physically and need to be forwarded through at least one other data processing core to complete the communication, the logical communication link between the transmission starting point and the corresponding transmission end point can include at least one physical communication link, and this physical communication link represents the link between two directly physically connected data processing cores.

[0069] In a specific implementation, for example, the following method can be used to determine the logical communication links between each transmission starting point and the transmission end point: for each transmission starting point, based on the mapped positions of the transmission starting point and the corresponding transmission end point on the on-chip network respectively, determine the first data processing core mapped by the transmission starting point on the on-chip network and the second data processing core mapped by the corresponding transmission end point on the on-chip network; based on the topological structure, determine the physical communication link connecting the first data processing core and the second data processing core; based on the physical communication link connecting the first data processing core and the second data processing core, generate the logical communication link between the transmission starting point and the corresponding transmission end point.

[0070] Here, for the convenience of explanation, Figure 3 in (b) is simplified to Figure 3 the form of (c) in, and in Figure 3 only the topological structure between the communication routes is retained in (c), and the mapping relationship between the data processing cores corresponding to each communication route and the computing nodes is marked accordingly.

[0071] Exemplarily, taking the computing node V 1 as the transmission starting point and the computing node V 0 as the transmission end point, using Figure 3The Network-on-Chip shown in (c) can determine that when the processed result data flows from the data processing core mapped by computing node V 1 to the computing node V 0 mapped data processing core, the processed result data passes through the computing node V 1 corresponding communication route R 5 and flows to the computing node V 0 corresponding communication route R 0 .

[0072] The communication route R 0 and the communication route R 1 corresponding data processing cores are physically connected in the Network-on-Chip. Therefore, a physical communication link can be determined between the communication route R 0 and the communication route R 1 . For the communication route R 0 and the communication route R 2 , the data processing cores corresponding to the two communication routes are not directly physically connected in the Network-on-Chip. Therefore, there is no physical communication link between the communication route R 0 and the communication route R 2 . However, a logical communication link between the communication route R 0 and the communication route R 1 can be established through the physical communication links between the corresponding data processing cores of the communication route R 1 and the communication route R 2 to transfer data. 0 2 2 In addition, for the case where the processed result data is transmitted from the communication route R

[0073] to the communication route R 5 in the above example, according to the description of the communication link above, it can be known that since the data processing cores corresponding to the communication route R 0 and the communication route R 5 are not directly physically connected, multiple physical communication links will be passed through when transmitting the processed result data. Exemplarily, the logical communication link between the communication route R 0 and the communication route R 5 to the communication route R 0 can be represented as S1. For the convenience of representation in Figure 2 (c), this logical communication link S1 is represented by a dotted arrow.

[0074] In this logical communication link S1, since the computing node V 1 and the computing node V 0The separately mapped data processing cores are not directly physically connected, so the corresponding logical communication link passes through the computing node V successively 5 The communication route R corresponding to the mapped data processing core 2 , and the computing node V 3 The communication route R corresponding to the mapped data processing core 1 . Then, corresponding to the logical communication link between the computing node V 1 and the computing node V 0 , it includes the physical communication link s1 between the computing node V 1 and the computing node V 5 , the physical communication link s2 between the computing node V 5 and the computing node V 3 , and the physical communication link s3 between the computing node V 3 and the computing node V 0 .

[0075] Here, only one logical communication link S1 between the computing node V 1 and the computing node V 0 is given. When actually determining the logical communication link, the logical communication link can also pass through the communication routes corresponding to other computing nodes. Specific limitations are not made here

[0076] In addition, for Figure 3 the communication links between other computing nodes and the corresponding adjacent network layer computing nodes in (a) can also be determined based on a similar method, which will not be elaborated here

[0077] S303: Determine the link communication volume of the logical communication link between each transmission starting point and the corresponding transmission end point based on the data transmission volume between different network layers

[0078] In the case of determining the physical communication links corresponding to multiple computing nodes respectively, by calculating the data transmission volume when the computing nodes perform data processing, the link communication volume of the physical communication links between the topologically adjacent data processing cores in the on-chip network can be determined accordingly

[0079] In Figure 3 (a) shows the data transmission volume when the computing nodes perform data transmission between different network layers. The data transmission volume between two directly data-transmitting computing nodes V i and V j is e(i, j). Exemplarily, the data transmission volume between the computing node V 1 and the computing node V 0 is e(1, 0)

[0080] Correspondingly, corresponding to Figure 3The logical communication link S1 in (c), due to the computing node V 1 The corresponding data processing core that maps transmits the communication data to the computing node V through the logical communication link S1 0 For the data processing core that maps, since the logical communication link S1 includes multiple physical communication links, correspondingly, the traffic volumes respectively corresponding to the multiple physical communication links s1, s2, and s3 also include the computing node V 1 And the computing node V 0 The data transmission volume e(1,0) between them.

[0081] S304: Determine the traffic volume of the physical communication link between the data processing cores in the on-chip network according to the traffic volume of the logical communication link between each transmission start point and the corresponding transmission end point.

[0082] As can be seen from the above, the data processing cores where the transmission start point and the corresponding transmission end point are located may not be directly physically connected, and at least one data processing core may be required for forwarding to complete the communication. In this step, it is necessary to determine the traffic volume of the physical communication link between the directly physically connected data processing cores.

[0083] In a specific implementation, when determining the traffic volume of the physical communication link between the data processing cores in the on-chip network, for example, the following method can be adopted: for each physical communication link between the data processing cores in the on-chip network, based on the logical communication link between each transmission start point and the corresponding transmission end point, determine the target logical communication link that includes this physical communication link; based on the traffic volume corresponding to the target logical communication link, determine the traffic volume of this physical communication link.

[0084] Continue with Figure 2 As an example, Figure 2 (c) The logical communication link S1 includes physical communication links s1, s2, and s3. Taking one of the physical communication links s1 as an example, if there are other logical communication links that include this communication link s1, then the corresponding logical communication links can also determine the data transmission volume between the corresponding transmission start point and the transmission end point in the above manner. In the case of determining all the logical communication links that include the physical communication link s1 and the data transmission volumes between the two computing nodes corresponding to all the determined logical communication links, the multiple data transmission volumes corresponding to the physical communication link s1 can be correspondingly determined. By adding the multiple data transmission volumes, the traffic volume corresponding to this physical communication link s1 can be correspondingly determined.

[0085] Regarding the above S103, the above S103 and S102 can be executed synchronously or asynchronously.

[0086] In a specific implementation, based on the association relationship information determined in S101 above and the data computation amounts corresponding to each network layer in the target neural network, the ideal throughput of the on-chip network can be correspondingly determined when using the on-chip network to execute the data processing task corresponding to the target neural network.

[0087] In a specific implementation, when determining the ideal throughput when using the on-chip network to execute the data processing task corresponding to the target neural network, for example, the following method can be adopted: for each network layer of the target neural network, based on the computing nodes corresponding to this network layer and the data computation amount corresponding to this network layer, determine the data computation time corresponding to each computing node corresponding to this network layer; based on the data computation times corresponding to multiple said computing nodes, determine the ideal throughput.

[0088] Among them, the data computation amounts corresponding to each network layer in the target neural network can be determined by using the number of neurons corresponding to each network layer. Exemplarily, for the pooling layer in the target neural network, if the output neurons of the pooling layer total o_h×o_w×o_ch, and the computation amount of each output neuron is k_size×k_size. Here, o_h is used to represent the output height of the output feature map, o_w is used to represent the output width of the output feature map, and o_ch is used to represent the output channel number of the output feature map. Then correspondingly, the data computation amount corresponding to this pooling layer is o_h×o_w×o_ch×k_size×k_size. Here, k_size is used to represent the size of the pooling parameter (kernel size) corresponding to the pooling layer.

[0089] Furthermore, after determining the computing nodes corresponding to each network layer respectively, since the data computation amounts corresponding to the network layers can be determined by the above method, the data computation amounts that the computing nodes corresponding to each network layer need to bear can be correspondingly determined. Since the data computation amounts corresponding to different network layers are different, and the number of computing nodes allocated to each network layer is also different, the data computation amounts that each computing node needs to bear are different, and the corresponding data computation times are also different.

[0090] Exemplarily, for the first network layer, if the first network layer is a convolutional layer, the total number of corresponding output neurons is o1_h × o1_w × o1_ch, and 4 computing nodes are allocated for this first network layer. Then, the number of output neurons corresponding to each computing node is o1_h × o1_w × o1_ch / 4. Here, o1_h is used to represent the output height of the output feature map in the convolutional layer, o1_w is used to represent the output width of the output feature map in the convolutional layer, and o1_ch is used to represent the output number of channels of the output feature map in the convolutional layer. If the computational amount of each output neuron is k1_size × k1_size × i1_ch, where k1_size is used to represent the size of the convolutional kernel corresponding to the convolutional layer, then for any one of the computing nodes in the first network layer, if the average distribution method is adopted, the corresponding data processing amount is o1_h × o1_w × o1_ch / 4 × k1_size × k1_size × i1_ch. Here, i1_ch is used to represent the number of channels of the input feature map. Specifically, without considering the difference in time between processing the image boundary and the interior, since the data processing amount corresponding to each computing node can be determined, the time required for each computing node to process its respective data processing amount can be determined accordingly by combining the processing capabilities of the data processing cores.

[0091] Exemplarily, for the second network layer, if the second network layer is a fully connected layer, the total number of corresponding output neurons is o2_h × o2_w × o2_ch, and 4 computing nodes are also allocated for this second network layer in the same way. Then, the number of output neurons corresponding to each computing node is o2_h × o2_w × o2_ch / 4. Here, o2_h is used to represent the output height of the output feature map in the fully connected layer, o2_w is used to represent the output width of the output feature map in the fully connected layer, and o2_ch is used to represent the output number of channels of the output feature map in the fully connected layer. If the computational amount of each output neuron is k2_size × k2_size × i2_ch, where k2_size is used to represent the size of the weight matrix corresponding to the fully connected layer, then for any one of the computing nodes in the second network layer, if the average distribution method is adopted, the corresponding data processing amount is o2_h × o2_w × o2_ch / 4 × k2_size × k2_size × i2_ch. Here, i2_ch is used to represent the number of channels of the input feature map of the fully connected layer. Specifically, since the time required for convolutional processing is longer, the time for each computing node corresponding to the fully connected layer to process its respective data processing amount is longer than the time for the computing nodes corresponding to the above convolutional layer to process their respective data processing amounts.

[0092] In this way, when determining the data processing volume corresponding to each network layer and the computing nodes corresponding to each network layer, the data calculation time corresponding to each computing node in each network layer can be determined.

[0093] After determining the data calculation time corresponding to each computing node, the ideal throughput when the network-on-chip executes the data processing task corresponding to the target neural network can be determined accordingly. Specifically, for example, the following method can be adopted: determine the longest data calculation time from the data calculation times corresponding to multiple computing nodes; based on the longest data calculation time, determine the corresponding ideal throughput under a preset time period.

[0094] In a specific implementation, since the data calculation time corresponding to all computing nodes in the target neural network can be determined, the computing node corresponding to the longest data calculation time max(cmpCycle) can be determined accordingly. For an image, when processed by multiple computing nodes, it can be determined that the time spent on data calculation at the computing node corresponding to the longest data calculation time is the longest. Then, when multiple computing nodes continuously process multiple images in a pipeline form, the computing node with the longest data calculation time will affect the data processing volume that multiple computing nodes in the target neural network can complete under a preset time period.

[0095] Exemplarily, when the longest data calculation time is determined to be 0.2 seconds, it can be determined that for multiple computing nodes, when continuously processing multiple images in a pipeline form, for any one of the images, the longest processing time of the computing node is 0.2 seconds. If the unit time is used as the preset time period, then correspondingly, if multiple computing nodes are all in a working state, affected by the computing node corresponding to the longest processing time, the maximum amount of data that can be processed (i.e., the ideal throughput thpt ideal ) satisfies the following formula (1):

[0096]

[0097] According to the above, when the longest data calculation time max(cmpCycle) is 0.2 seconds, the ideal throughput thpt within the unit time period can be determined accordingly. ideal is the data volume corresponding to 5 images.

[0098] For the above S104, when determining the link traffic and the ideal throughput, the performance information when running the target neural network using the network-on-chip can be determined accordingly. Specifically, for example, the following method can be adopted: for each physical communication link in the network-on-chip, based on the link traffic corresponding to the physical communication link and the ideal throughput, determine the required link bandwidth corresponding to the physical communication link; based on the required link bandwidths corresponding to the respective physical communication links in the network-on-chip and the actual link bandwidths, determine the loss rates corresponding to the respective physical communication links; based on the loss rates corresponding to the respective physical communication links, determine the performance information.

[0099] In a specific implementation, when determining the link traffic corresponding to each physical communication link and the ideal throughput, the required link bandwidth corresponding to the physical communication link can be determined. Among them, in the following example, the physical communication link is represented by s i and the link traffic of each physical communication link s i is represented by Exemplarily, in the case where the physical communication link corresponding to the longest data calculation time is s 1 , the link traffic corresponding to s 1 is, for example, 8 Mbit.

[0100] The required link bandwidths corresponding to the respective physical communication links satisfy the following formula (2):

[0101]

[0102] Exemplarily, for the physical communication link s 1 , the corresponding required link bandwidth is, for example, 40 Mbit / s.

[0103] In addition, the actual link bandwidths corresponding to the respective physical communication links s i can also be determined accordingly, denoted as Since the network-on-chip has been determined, its corresponding performance is also determined accordingly. Therefore, the respective physical communication links in the network-on-chip and the actual link bandwidths corresponding to the respective physical communication links can be directly determined. Exemplarily, when determining the actual link bandwidths corresponding to the respective physical communication links, they can be directly obtained according to the relevant parameters of the network-on-chip, which will not be elaborated here. In a possible case, the link traffic corresponding to the respective physical communication links of the network-on-chip can be the same or different.

[0104] When the required link bandwidth and the actual link bandwidth corresponding to each physical communication link in the network-on-chip are determined, the loss rate corresponding to each physical communication link can also be determined accordingly. Specifically, for each physical communication link s i The corresponding loss rate For example, it can satisfy the following formula (3):

[0105]

[0106] In a possible case, for any physical communication link, if the required link bandwidth corresponding to it is less than the actual link bandwidth, the data transmission requirements of this physical communication link can be satisfied. Accordingly, according to the above formula (3), for the physical communication link s i , in the case where the required link bandwidth is less than the actual link bandwidth , the determined range of the loss rate is (0, 1]. In this case, it can be determined that for this physical communication link s i , the physical communication link can correspondingly undertake the data transmission requirements.

[0107] In this case, this physical communication link can support the transmission of the required data volume. However, since the required link bandwidth is less than the actual link bandwidth, the required data volume does not reach the maximum data transmission volume that can be supported, and this physical communication link does not achieve a high data transmission utilization rate.

[0108] Exemplarily, for the physical communication link s 1 , the corresponding required link bandwidth is, for example, 40 Mbit / s, and the corresponding actual link bandwidth If it is 80 Mbit / s, then the corresponding loss rate is 0.5.

[0109] In another possible case, for any physical communication link, if the required link bandwidth corresponding to it is greater than the actual link bandwidth, the data transmission requirements of this physical communication link cannot be satisfied. Accordingly, according to the above formula (3), for the physical communication link s i , in the case where the required link bandwidth is greater than the actual link bandwidth , the determined range of the loss rate is greater than 1. In this case, it can be determined that for this physical communication link s i , the physical communication link cannot undertake the corresponding data transmission requirements.

[0110] In this case, this physical communication link can ensure that when transmitting the required amount of data to be transmitted, it occupies the actual link bandwidth to the maximum extent for data transmission. However, due to the limited actual link bandwidth, it is impossible to transmit all the data to be transmitted at once through the actual link bandwidth. In a possible case, it is necessary to adopt a method of transmitting in batches, and complete the transmission of all data through multiple transmissions, that is, make up for the shortage of the actual link bandwidth by spending more time. Therefore, for this link, there is a problem that the actual link bandwidth cannot meet the actual communication requirements.

[0111] Exemplarily, for the physical communication link s 1 , the corresponding required link bandwidth is, for example, 40 Mbit / s, and the corresponding actual link bandwidth If it is 20 Mbit / s, then the corresponding loss rate is 2.

[0112] In addition, in the case of determining the loss rate corresponding to each physical communication link, the performance information of the on-chip network can be determined accordingly. Specifically, for example, the following method can be adopted: determine the maximum loss rate among the loss rates corresponding to each physical communication link as the on-chip network loss rate; in response to the on-chip network loss rate being greater than the preset loss rate threshold, determine that the performance information of the on-chip network is communication-limited; in response to the on-chip network loss rate being less than the preset loss rate threshold, determine that the performance information of the on-chip network is computation-limited.

[0113] Specifically, according to the above description of the loss rate corresponding to the physical communication link, for example, the preset loss rate threshold can be determined as "1"; when the loss rate of the physical communication link is less than 1, there is a problem that this physical communication link fails to achieve a high data transmission utilization rate; when the loss rate of the physical communication link is greater than 1, there is a problem that the actual link bandwidth of this link cannot meet the actual communication requirements; and when the loss rate of the physical communication link is equal to 1, it can be considered that this physical communication link is in a full-load data transmission state. For multiple physical communication links of the on-chip network, the physical communication link that has the greatest impact on communication will affect the overall data transmission of the on-chip network. For example, if there is a physical communication link with a too large loss rate, for example, the corresponding loss rate is 2, then for this physical communication link, it can be determined that 2 times the transmission time is required to make up for the inability to complete the transmission of all data at once due to the insufficient time link bandwidth.

[0114] Here, for the case where the loss rate corresponding to the physical communication link is small, although there is a problem that this physical communication link fails to achieve a high data transmission utilization rate, this physical communication link can correspondingly support the transmission of the required amount of data to be transmitted, so it will not affect the data transmission of the on-chip network.

[0115] Therefore, when the loss rates corresponding to the respective physical communication links are determined, the maximum loss rate among the loss rates corresponding to the respective physical communication links can be determined as the on-chip network loss rate lRate. By comparing the on-chip network loss rate with a preset loss rate threshold, the performance information of the on-chip network can be determined accordingly.

[0116] In a possible case, if the on-chip network loss rate is greater than the preset loss rate threshold, since there are physical communication links that take more time to complete data transmission. Taking data processing of an image as an example, when multiple data processing cores in the on-chip network perform data processing in a pipeline form, due to the time limit of this physical communication link, the number of images that can be processed will be reduced. Here, the performance information corresponding to the on-chip network in this case is determined as communication-limited.

[0117] Exemplarily, if the on-chip network loss rate is determined to be 0.5, the performance information of the on-chip network can be determined as compute-limited. In this case, the actual throughput thpt of the on-chip network real can be determined by the ideal throughput thpt ideal Specifically, for example, it can satisfy the following formula (4):

[0118] thpt real = thpt ideal (4)

[0119] In this case, since the on-chip network is compute-limited, it does not affect communication, and the determined actual throughput can be the same as or close to the above-determined ideal throughput.

[0120] If the on-chip network loss rate is determined to be 2, the performance information of the on-chip network can be determined as communication-limited. In this case, the actual throughput thpt of the on-chip network rea l can be determined by the ideal throughput thpt ideal and the on-chip network loss rate lRate. Specifically, for example, it can satisfy the following formula (5):

[0121]

[0122] In this case, since the on-chip network is communication-limited, it is impossible to complete the communication transmission of the ideal throughput within the longest data calculation time. Here, according to the ideal throughput and the on-chip network loss rate, the throughput that can be communicated within the longest calculation time can be determined accordingly as the actual throughput.

[0123] In another embodiment of the present disclosure, in the case of determining the performance information of the on-chip network when running the target neural network through the above performance determination method, the partitioning method for the computing nodes for the target neural network and / or the mapping positions of the multiple computing nodes on the on-chip network can be adjusted accordingly based on the performance information, so that after the adjusted multiple target computing nodes determined for the target neural network and the mapping positions of the multiple target computing nodes in the on-chip network are mapped into the on-chip network, the on-chip network can achieve a balance between computing and communication when running the target neural network, that is, fully utilize the link traffic corresponding to the physical communication links on the on-chip network and complete the data transmission relatively quickly.

[0124] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process, and the specific execution order of each step should be determined according to its function and possible internal logic.

[0125] Based on the same inventive concept, a performance determination device corresponding to the performance determination method is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above performance determination method in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0126] Refer to Figure 4 As shown in the figure, it is a schematic diagram of a performance determination device provided by an embodiment of the present disclosure. The device includes: a first determination module 41, a processing module 42, and a second determination module 43; wherein,

[0127] The first determination module 41 is configured to determine multiple computing nodes in the target neural network mapped by the to-be-evaluated on-chip network;

[0128] The processing module 42 is configured to determine the link traffic of the physical communication links between the data processing cores in the on-chip network based on the association relationship information between the multiple computing nodes, the topological structure between the data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determine the ideal throughput when using the on-chip network to execute the data processing tasks corresponding to the target neural network based on the association relationship information and the data calculation amounts corresponding to each network layer in the target neural network;

[0129] The second determination module 43 is configured to determine the performance information of using the on-chip network to run the target neural network based on the link traffic and the ideal throughput.

[0130] In an alternative embodiment, when determining the link traffic of the physical communication links between the data processing cores in the on-chip network based on the association relationship information between multiple computing nodes, the topological structure between the data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network, the processing module 42 is configured to: traverse the computing nodes corresponding to other network layers except the last network layer, and for each network layer, determine the computing nodes in this network layer that have data flow interaction with the computing nodes in the adjacent network layer, use the computing nodes existing in this network layer as the transmission starting points, and use the computing nodes in the adjacent network layer corresponding to this computing node as the transmission ending points; based on the mapping positions of the determined transmission starting points and corresponding transmission ending points on the on-chip network respectively, and the topological structure, determine the logical communication links between each transmission starting point and the corresponding transmission ending point respectively; based on the data transmission volume between different network layers, determine the link traffic of the logical communication links between each transmission starting point and the corresponding transmission ending point; and determine the link traffic of the physical communication links between the data processing cores in the on-chip network according to the link traffic of the logical communication links between each transmission starting point and the corresponding transmission ending point.

[0131] In an alternative embodiment, when determining the logical communication links between each transmission starting point and the corresponding transmission ending point respectively based on the mapping positions of the determined transmission starting points and corresponding transmission ending points on the on-chip network respectively, and the topological structure, the processing module 42 is configured to: for each transmission starting point, based on the mapping positions of this transmission starting point and the corresponding transmission ending point on the on-chip network respectively, determine the first data processing core mapped by this transmission starting point on the on-chip network, and the second data processing core mapped by the corresponding transmission ending point on the on-chip network; based on the topological structure, determine the physical communication link connecting the first data processing core and the second data processing core; and generate the logical communication link between this transmission starting point and the corresponding transmission ending point based on the physical communication link connecting the first data processing core and the second data processing core.

[0132] In an alternative embodiment, when determining the link traffic of the physical communication links between the data processing cores in the on-chip network according to the link traffic of the logical communication links between each transmission starting point and the corresponding transmission ending point, the processing module 42 is configured to: for each physical communication link between the data processing cores in the on-chip network, based on the logical communication links between each transmission starting point and the corresponding transmission ending point, determine the target logical communication link including this physical communication link; and determine the link traffic of this physical communication link based on the link traffic corresponding to the target logical communication link.

[0133] In an alternative embodiment, when the processing module 42 determines the ideal throughput for executing the data processing task corresponding to the target neural network by using the network-on-chip based on the association relationship information and the data computation amount corresponding to each network layer in the target neural network, it is configured to: for each network layer of the target neural network, determine the data computation time corresponding to each computing node in this network layer according to the computing node corresponding to this network layer and the data computation amount corresponding to this network layer; and determine the ideal throughput based on the data computation times corresponding to the multiple computing nodes.

[0134] In an alternative embodiment, when the processing module 42 determines the ideal throughput based on the data computation times corresponding to the multiple computing nodes, it is configured to: determine the longest data computation time from the data computation times corresponding to the multiple computing nodes; and determine the ideal throughput corresponding to a preset time duration based on the longest data computation time.

[0135] In an alternative embodiment, when the second determination module 43 determines the performance information for running the target neural network by using the network-on-chip based on the link traffic and the ideal throughput, it is configured to: for each physical communication link in the network-on-chip, determine the required link bandwidth corresponding to this physical communication link based on the link traffic corresponding to this physical communication link and the ideal throughput; determine the loss rate corresponding to each physical communication link based on the required link bandwidths corresponding to the respective physical communication links in the network-on-chip and the actual link bandwidth; and determine the performance information based on the loss rates corresponding to the respective physical communication links.

[0136] In an alternative embodiment, when the second determination module 43 determines the performance information based on the loss rates corresponding to the respective physical communication links, it is configured to: determine the maximum loss rate among the loss rates corresponding to the respective physical communication links as the network-on-chip loss rate; in response to the network-on-chip loss rate being greater than a preset loss rate threshold, determine that the performance information of the network-on-chip is communication-limited; in response to the network-on-chip loss rate being less than the preset loss rate threshold, determine that the performance information of the network-on-chip is computation-limited; and in response to the network-on-chip loss rate being equal to the preset loss rate threshold, determine that the performance information of the network-on-chip is full-load transmission.

[0137] For the description of the processing procedures of the respective modules in the device and the interaction procedures between the respective modules, reference may be made to the relevant descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0138] The embodiments of the present disclosure further provide a computer device, such as Figure 5As shown in the figure, it is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure, including:

[0139] A processor 10 and a memory 20; the memory 20 stores machine-readable instructions executable by the processor 10, and the processor 10 is used to execute the machine-readable instructions stored in the memory 20. When the machine-readable instructions are executed by the processor 10, the processor 10 performs the following steps:

[0140] Determine multiple computing nodes in the target neural network of the on-chip network mapping to be evaluated; based on the association relationship information between the multiple computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network, determine the link traffic of the physical communication links between data processing cores in the on-chip network; and based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network, determine the ideal throughput when using the on-chip network to execute the data processing task corresponding to the target neural network; based on the link traffic and the ideal throughput, determine the performance information when using the on-chip network to run the target neural network.

[0141] The above-mentioned memory 20 includes an internal memory 210 and an external memory 220; the internal memory 210 here is also called the main memory, which is used to temporarily store the operation data in the processor 10 and the data exchanged with the external memory 220 such as a hard disk. The processor 10 exchanges data with the external memory 220 through the internal memory 210.

[0142] The specific execution process of the above instructions can refer to the steps of the performance determination method described in the embodiment of the present disclosure, which will not be elaborated here.

[0143] The embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the performance determination method described in the above method embodiment. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.

[0144] The embodiment of the present disclosure also provides a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the performance determination method described in the above method embodiment. For details, please refer to the above method embodiment, which will not be elaborated here.

[0145] Among them, the above computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), and so on.

[0146] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system and device can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed system, device, and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0147] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, in each embodiment of the present disclosure, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0149] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0150] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for determining performance, characterized in that, comprising: determining a plurality of computing nodes in a target neural network of an on-chip network mapping to be evaluated; determining the link traffic of physical communication links between data processing cores in the on-chip network based on the association relationship information between the plurality of computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determining the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network; determining the performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput; The determining the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network includes: for each network layer of the target neural network, determining the data computation time corresponding to each computing node in this network layer according to the computing nodes corresponding to this network layer and the data computation amount corresponding to this network layer; determining the ideal throughput based on the data computation times corresponding to the plurality of computing nodes; The determining the performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput includes: for each physical communication link in the on-chip network, determining the required link bandwidth corresponding to this physical communication link based on the link traffic corresponding to this physical communication link and the ideal throughput; determining the loss rate corresponding to each physical communication link based on the required link bandwidths corresponding to each physical communication link in the on-chip network and the actual link bandwidth; determining the performance information based on the loss rates corresponding to each physical communication link.

2. The performance determination method according to claim 1, characterized in that, the determining the link traffic of physical communication links between data processing cores in the on-chip network based on the association relationship information between the plurality of computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network includes: traversing the computing nodes corresponding to other network layers except the last network layer, for each network layer, determining the computing nodes in this network layer that have data stream interaction with the computing nodes of the adjacent network layer, taking the computing nodes existing in this network layer as the transmission starting points, and taking the computing nodes of the adjacent network layer corresponding to this computing node as the transmission ending points; respectively determining the logical communication links between each transmission starting point and the corresponding transmission ending point based on the mapping positions of the determined transmission starting points and corresponding transmission ending points on the on-chip network and the topological structure; Determine the link traffic of the logical communication link between each transmission start point and the corresponding transmission end point based on the data traffic between different network layers; Determine the link traffic of the physical communication link between data processing cores in the on-chip network according to the link traffic of the logical communication link between each transmission start point and the corresponding transmission end point.

3. The performance determination method according to claim 2, characterized in that, The determining of the logical communication link between each transmission start point and the corresponding transmission end point respectively based on the mapped positions of the determined transmission start points and corresponding transmission end points on the on-chip network and the topological structure includes: For each transmission start point, based on the mapped positions of the transmission start point and the corresponding transmission end point on the on-chip network, determine the first data processing core mapped by the transmission start point on the on-chip network and the second data processing core mapped by the corresponding transmission end point on the on-chip network; Based on the topological structure, determine the physical communication link connecting the first data processing core and the second data processing core; Based on the physical communication link connecting the first data processing core and the second data processing core, generate the logical communication link between the transmission start point and the corresponding transmission end point.

4. The performance determination method according to claim 2 or 3, characterized in that, The determining of the link traffic of the physical communication link between data processing cores in the on-chip network according to the link traffic of the logical communication link between each transmission start point and the corresponding transmission end point includes: For each physical communication link between data processing cores in the on-chip network, based on the logical communication links between each transmission start point and the corresponding transmission end point, determine the target logical communication link including the physical communication link; Based on the link traffic corresponding to the target logical communication link, determine the link traffic of the physical communication link.

5. The performance determination method according to claim 1, characterized in that, The determining of the ideal throughput based on the data calculation times respectively corresponding to multiple computing nodes includes: Determine the longest data calculation time from the data calculation times respectively corresponding to multiple computing nodes; Based on the longest data calculation time, determine the corresponding ideal throughput under a preset time duration.

6. The performance determination method according to claim 1, characterized in that, The determining of the performance information based on the loss rates respectively corresponding to each physical communication link includes: Determine the maximum loss rate among the loss rates respectively corresponding to each physical communication link as the on-chip network loss rate; In response to the on-chip network loss rate being greater than a preset loss rate threshold, determine that the performance information of the on-chip network is communication limited; In response to the on-chip network loss rate being less than the preset loss rate threshold, determine that the performance information of the on-chip network is calculation limited; In response to the on-chip network loss rate being equal to the preset loss rate threshold, determine that the performance information of the on-chip network is full-load transmission.

7. A performance determination device, characterized in that, comprises: A first determination module, configured to determine multiple computing nodes in a target neural network mapped by the on-chip network to be evaluated; A processing module, configured to determine the link traffic of the physical communication links between data processing cores in the on-chip network based on the association relationship information between multiple computing nodes, the topological structure between data processing cores in the on-chip network, and the mapping positions of the computing nodes on the on-chip network; and determine the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network. A second determination module, configured to determine the performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput. When determining the ideal throughput when executing the data processing task corresponding to the target neural network by using the on-chip network based on the association relationship information and the data computation amounts corresponding to each network layer in the target neural network, the processing module is configured to: For each network layer of the target neural network, determine the data computation time corresponding to each computing node in this network layer according to the computing node corresponding to this network layer and the data computation amount corresponding to this network layer. Determine the ideal throughput based on the data computation times corresponding to multiple computing nodes. When determining the performance information when running the target neural network by using the on-chip network based on the link traffic and the ideal throughput, the processing module is configured to: For each physical communication link in the on-chip network, determine the required link bandwidth corresponding to this physical communication link based on the link traffic corresponding to this physical communication link and the ideal throughput. Determine the loss rate corresponding to each physical communication link based on the required link bandwidths corresponding to each physical communication link in the on-chip network and the actual link bandwidths. Determine the performance information based on the loss rates corresponding to each physical communication link.

8. A computer device Characterized in that It includes: A processor and a memory. The memory stores machine-readable instructions executable by the processor. The processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor executes the steps of the performance determination method according to any one of claims 1 to 6.

9. A computer-readable storage medium Characterized in that A computer program is stored on the computer-readable storage medium. When the computer program is run by a computer device, the computer device executes the steps of the performance determination method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • IP (Internet Protocol) core fast mapping method for network on chip based on region division

    CN102065019A

  • Schedule optimization method for communication energy consumption in on-chip network

    CN103631659A