Mapping Scheme Optimization Method and Apparatus, Electronic Device, Readable Storage Medium
By reconstructing and optimizing the inter-core connections in the initial mapping scheme, the routing pressure problem of neural networks when executed in the multi-core system is solved, and more efficient mapping schemes and system efficiency are achieved.
Patent Information
- Application Number
- CN202210539850.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The prior art is difficult to effectively solve the routing pressure problem caused by the mapping scheme of neural networks when executed in multi-core systems, affecting system efficiency.
By reconstructing the inter-core connections in the initial mapping scheme, it is determined whether the preset inter-core optimization conditions are met. If so, the optimized mapping scheme will be determined to reduce the pressure of inter-core data transmission and routing.
The cohesion of mapping schemes is realized, reducing data transmission between cores, reducing routing pressure on the multi-core system, and thus improving the efficiency of executing neural networks.
Smart Images

Figure CN114881221B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a method and device for optimizing a mapping scheme, an electronic device, and a readable storage medium. Background Art
[0002] The popularization of intelligent applications has made neural network computing more and more important. However, general-purpose processors are difficult to meet the requirements of the large computing volume and large storage volume required by neural networks. For this reason, relevant technical personnel have developed a many-core system with a neural network acceleration architecture. When the many-core system executes a neural network, many neurons of the neural network are spatially distributed on different processing cores, that is, a mapping scheme is obtained. Since the data transfer between processing cores is realized through the routing system (network on chip) of the many-core system, and there is a large amount of data transfer in the neural network itself, the data transfer volume between processing cores is very large, and the unprocessed mapping scheme affects the efficiency of the many-core system in executing the neural network. Summary of the Invention
[0003] The present disclosure provides a method and device for optimizing a mapping scheme based on a many-core system, an electronic device, and a readable storage medium.
[0004] In a first aspect, the present disclosure provides a method for optimizing a mapping scheme based on a many-core system. The method for optimizing a mapping scheme based on a many-core system includes:
[0005] Obtain an initial first mapping scheme, where the first mapping scheme is used to map a first neural network to be executed to multiple processing cores of the many-core system, each processing core is used to execute at least one neuron of the first neural network, and the first mapping scheme includes inter-core connections between the respective processing cores;
[0006] Reconstruct the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction, where n ≥ 1 and n is an integer, and the mapping scheme of the 0-th inter-core reconstruction is the first mapping scheme;
[0007] Determine whether the mapping scheme of the n-th inter-core reconstruction satisfies the inter-core optimization condition;
[0008] When the mapping scheme of the n-th inter-core reconstruction satisfies the preset inter-core optimization condition, determine an optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction.
[0009] In a second aspect, the present disclosure provides a device for optimizing a mapping scheme based on a many-core system. The device for optimizing a mapping scheme based on a many-core system includes:
[0010] An acquisition module, configured to acquire an initial first mapping scheme, where the first mapping scheme is used to map a first neural network to be executed to a plurality of processing cores of the many-core system, each of the processing cores is used to execute at least one neuron of the first neural network, and the first mapping scheme includes inter-core connections between the processing cores;
[0011] A reconstruction module, configured to reconstruct the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction, where n ≥ 1 and n is an integer, and the mapping scheme of the 0-th inter-core reconstruction is the first mapping scheme;
[0012] A judgment module, configured to judge whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition;
[0013] A determination module, configured to, when the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, determine an optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction.
[0014] In a third aspect, the present disclosure provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the above-mentioned mapping scheme optimization method based on a many-core system.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which a computer program is stored, where the computer program, when executed by a processor / processing core, implements the above-mentioned mapping scheme optimization method based on a many-core system.
[0016] In the embodiments provided by the present disclosure, the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction are reconstructed to obtain the mapping scheme of the n-th inter-core reconstruction, and then it is judged whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition. When the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, an optimized second mapping scheme is determined according to the mapping scheme of the n-th inter-core reconstruction, which optimizes the mapping scheme of the neural network, realizes the cohesion of the mapping scheme, reduces the inter-core data transmission, alleviates the routing pressure of the many-core system, and improves the efficiency of the many-core system in executing the neural network.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are provided to further understand the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. By describing the detailed exemplary embodiments with reference to the drawings, the above and other features and advantages will become more obvious to those skilled in the art. In the drawings:
[0019] Figure 1 is a flowchart of an optimization method for a mapping scheme based on a many-core system provided by an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a mapping scheme obtained by mapping a neural network to different processing cores of a many-core system in an embodiment of the present disclosure;
[0021] Figure 3 is a flowchart of an inter-core reconstruction optimization method provided by an embodiment of the present disclosure;
[0022] Figure 4 is a schematic diagram of the change of the mapping scheme during the inter-core pruning optimization provided by an embodiment of the present disclosure;
[0023] Figure 5 is a flowchart of another inter-core reconstruction optimization method provided by an embodiment of the present disclosure;
[0024] Figure 6 is a schematic diagram of the change of the mapping scheme during the inter-core reconstruction in an embodiment of the present disclosure;
[0025] Figure 7 is a flowchart of an intra-core reconnection optimization method provided by an embodiment of the present disclosure;
[0026] Figure 8 is a flowchart of an intra-core reconnection optimization method provided by an embodiment of the present disclosure;
[0027] Figure 9 is a schematic diagram of the change process of the mapping scheme during the intra-core reconnection optimization provided by an embodiment of the present disclosure;
[0028] Figure 10 is a flowchart of an optimization method for a mapping scheme provided by an embodiment of the present disclosure;
[0029] Figure 11 is a schematic diagram of the change process of the mapping scheme during two rounds of inter-core reconstruction optimization and intra-core reconnection optimization provided by an embodiment of the present disclosure;
[0030] Figure 12 is a block diagram of an optimization device for a mapping scheme based on a many-core system provided by an embodiment of the present disclosure;
[0031] Figure 13 A block diagram of an electronic device provided by an embodiment of the present disclosure. Specific embodiments
[0032] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the following provides a description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0033] Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0034] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0035] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms "comprise" and / or "consist of" are used in this specification, it specifies the presence of the stated features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups. "Connection" or "connected" and other similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0036] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.
[0037] Current neural network mapping schemes consider factors such as the memory and load balancing of many-core systems more, and rarely consider routing factors. Therefore, the routing pressure of the mapping scheme is relatively large.
[0038] According to the mapping scheme optimization method of the embodiments of the present disclosure, the mapping scheme can be optimized, the routing pressure can be reduced, and thus the efficiency of the many-core system in executing neural networks can be improved.
[0039] The mapping scheme optimization method according to an embodiment of the present disclosure may be executed by an electronic device, which may be an electronic device such as a terminal device or a server. The terminal device may be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method may be implemented by a processor invoking computer-readable program instructions stored in a memory. Alternatively, the method may be executed by a server.
[0040] In some embodiments, a many-core system may include multiple processing units, which may be chips or processing cores. That is, the many-core system includes multiple chips, and each chip is used to execute a sub-task to be processed; or, the many-core system includes multiple processing cores, and each processing core is used to execute a sub-task to be processed, where the sub-task to be processed may be at least one neuron of a neural network. For ease of description, the following embodiments will be described by taking processing cores as an example.
[0041] In some embodiments, for a neural network to be executed (including multiple neurons), a corresponding mapping scheme may be generated by an external device (such as a compiler) so as to map the neural network to all or part of the processing cores or chips of the many-core system, so that the many-core system can execute the processing task corresponding to the neural network.
[0042] Figure 1 It is a flowchart of a mapping scheme optimization method based on a many-core system provided by an embodiment of the present disclosure. Referring to Figure 1 , the mapping scheme optimization method based on a many-core system provided by an embodiment of the present disclosure includes:
[0043] In step S101, an initial first mapping scheme is obtained.
[0044] Among them, the first mapping scheme is used to map a first neural network to be executed to multiple processing cores of the many-core system, and each processing core is used to execute at least one neuron of the first neural network. The first mapping scheme includes inter-core connections between the processing cores.
[0045] In some embodiments, a many-core system includes multiple processing cores, and physical parameters such as the storage capacity and computing capacity of each processing core may be the same or different. The first mapping scheme is the mapping scheme obtained after mapping the first neural network to be executed to each processing core of the many-core system. Exemplarily, the first mapping scheme is a scheme determined by mapping the first neural network to the many-core system based on the physical parameters of the many-core system. In the first mapping scheme, one processing core may map one or more neurons in the first neural network. Among them, the information transmission between neurons located in different processing cores requires occupying the routing resources of the many-core system, and the information transmission between neurons located in the same processing core does not occupy the routing resources of the many-core system. Therefore, the intra-core connections in the first neural network are the connections between neurons located in the same processing core, and the inter-core connections in the first neural network are the connections between neurons located in different processing cores.
[0046] In the embodiments of the present disclosure, in order to reduce the occupation of the routing resources of the many-core system, the first mapping scheme can be optimized to obtain the optimized second mapping scheme.
[0047] In some embodiments, the first mapping scheme not only includes the processing cores involved in the first neural network and the inter-core connections between the processing cores, but also includes the processing nodes corresponding to the neurons in the first neural network and the intra-core connections between the processing nodes. Each processing node corresponds to at least one neuron.
[0048] Figure 2 It is a schematic diagram of the mapping scheme obtained by mapping the neural network to different processing cores of the many-core system in the embodiments of the present disclosure. In Figure 2 it, the squares represent processing cores, the black dots represent processing nodes, the processing nodes correspond to one or a group (a layer) of neurons, and the lines between the processing nodes represent data transfer. Figure 2 The shown mapping scheme includes four processing cores, namely the first processing core 21, the second processing core 22, the third processing core 23, and the fourth processing core 24. Among them, each of the first processing core 21, the second processing core 22, and the third processing core 23 includes three processing nodes, and the fourth processing core 24 is provided with four processing nodes. A total of nine inter-core connections are provided between the four processing cores. Among them, there is one inter-core connection 201 between the first processing core 21 and the second processing core 22, there are two inter-core connections 202A and 202B between the first processing core 21 and the third processing core 23, there is one inter-core connection 203 between the first processing core 21 and the fourth processing core 24, there are two inter-core connections 204A and 204B between the second processing core 22 and the third processing core 23, there are two inter-core connections 205A and 205B between the second processing core 22 and the fourth processing core 24, and there is one inter-core connection 206 between the third processing core 23 and the fourth processing core 24. In Figure 2In the mapping scheme shown, there are many and complex inter-core connections between processing cores. Therefore, it is necessary to optimize this mapping scheme to relieve the routing pressure of the many-core system.
[0049] In step S102, the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction are reconstructed to obtain the mapping scheme of the n-th inter-core reconstruction.
[0050] Where n is an integer greater than or equal to 1, and the mapping scheme of the 0-th inter-core reconstruction is the first mapping scheme in step S101, that is, when optimizing the mapping scheme, the first mapping scheme adopted is the initial mapping scheme corresponding to the first neural network to be executed.
[0051] In some embodiments, the way to reconstruct the inter-core connections may include: deleting some inter-core connections with smaller weights from the mapping scheme of the (n - 1)-th inter-core reconstruction; deleting some connections with higher costs; deleting some longer inter-core connections and correspondingly adding some shorter inter-core connections, etc. The present disclosure does not limit this. In this way, after reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction, the mapping scheme of the n-th inter-core reconstruction can be obtained.
[0052] In step S103, it is judged whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition.
[0053] When the mapping scheme of the n-th inter-core reconstruction does not meet the preset inter-core optimization condition, return to step S102 to reconstruct the inter-core connections again. When the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, execute step S104.
[0054] In step S104, when the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, determine the optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction.
[0055] In some embodiments, the inter-core optimization condition is that the number of cycles of inter-core optimization reaches a preset number of cycles. For example, when the preset number of cycles is N times, in step S103, when it is judged that n is less than N, return to step S102; when n is equal to N, determine the mapping scheme of the n-th inter-core reconstruction as the second mapping scheme. Where N is a positive integer.
[0056] In some embodiments, the inter-core optimization condition is that the mapping scheme of inter-core reconstruction meets the routing load requirement. For example, a routing load threshold is preset in advance. When the mapping scheme of the n-th inter-core reconstruction can meet the preset routing load threshold, determine the mapping scheme of the n-th inter-core reconstruction as the second mapping scheme.
[0057] According to an embodiment of the present disclosure, the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction are reconstructed to obtain the mapping scheme of the n-th inter-core reconstruction. Then, it is determined whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition. When the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, an optimized second mapping scheme is determined according to the mapping scheme of the n-th inter-core reconstruction. This method optimizes the mapping scheme of the neural network, makes the mapping scheme more cohesive, reduces the data transfer between cores, alleviates the routing pressure of the many-core system, and improves the efficiency of the many-core system in executing the neural network.
[0058] In an embodiment of the present disclosure, in the optimized second mapping scheme, the inter-core connections of the first neural network are re-determined, and the inter-core connections are reduced. Thus, under the condition of meeting the accuracy requirements of the first neural network, the routing pressure of the many-core system is reduced. When the optimized second mapping scheme is mapped to the many-core system and the many-core system is used to execute the first neural network, both the accuracy requirements of the first neural network can be met and the routing pressure of the many-core system can be reduced.
[0059] Figure 3 The figure is a flowchart of an inter-core reconstruction optimization method provided by an embodiment of the present disclosure. This inter-core reconstruction optimization method optimizes the mapping scheme through an inter-core pruning method. As Figure 3 shown, reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction in step S102 to obtain the mapping scheme of the n-th inter-core reconstruction includes:
[0060] Step S301, determining the connection cost and connection weight of each inter-core connection in the mapping scheme of the (n - 1)-th inter-core reconstruction.
[0061] Among them, the connection cost indicates the cost of the connection between the processing cores corresponding to the inter-core connection, that is, the routing resources that this inter-core connection needs to occupy when the neural network is executed using the mapping scheme. The connection weight indicates the weight of the connection between the neurons corresponding to the inter-core connection, that is, the contribution size of this inter-core connection among all the inter-core connections involved in the mapping scheme when the neural network is executed.
[0062] In some embodiments, the connection cost is determined based on the per-unit-time data communication volume A of the inter-core connection and the distance D between the starting processing core and the ending processing core corresponding to the inter-core connection. That is, the connection cost C is a function of the per-unit-time data communication volume A and the distance D between the starting processing core and the ending processing core corresponding to the inter-core connection, as shown in Equation (1):
[0063] cost(c) = f(A, D) (1)
[0064] Among them, cost(c) represents the connection cost, f() represents the cost function, A represents the data communication volume per unit time, and D represents the distance between the source processing core and the destination processing core. The present disclosure places no restrictions on the specific representation form of the cost function.
[0065] Step S302: Set the m-th cost threshold and the m-th first weight threshold for the inter-core connection, where m ≥ 1 and m is an integer.
[0066] In some embodiments, each loop can set the cost threshold Cost of the inter-core connection according to the situation th and the first weight threshold W of the inter-core connection th .
[0067] Step S303: Delete the inter-core connections whose connection cost is greater than the m-th cost threshold and whose absolute value of the connection weight is less than the m-th first weight threshold, to obtain the m-th inter-core pruning mapping scheme.
[0068] In some embodiments, delete the inter-core connections whose connection cost is greater than the m-th cost threshold and whose absolute value of the connection weight is less than the m-th first weight threshold, to obtain the m-th inter-core pruning mapping scheme.
[0069] Step S304: Retrain the second neural network corresponding to the m-th inter-core pruning mapping scheme using the training data set, to obtain the m-th retrained second neural network.
[0070] In some embodiments, the loss function used for training the second neural network using the training data set includes a connection cost regularization term and a connection weight regularization term.
[0071] Exemplarily, the loss function Loss for training is Equation (2):
[0072]
[0073] Among them, Loss′ represents the original loss function on the training data set, c represents the inter-core connection, C represents all the inter-core connections in the mapping scheme, Cost(c) represents the connection cost, that is, the connection cost regularization term, and W(c) represents the connection weight, that is, the connection weight regularization term.
[0074] Step S305: Determine whether the accuracy of the m-th retrained second neural network reaches a preset accuracy threshold. If the accuracy of the m-th retrained second neural network does not reach the preset accuracy threshold, return to execute Step S302, that is, set the cost threshold and the first weight threshold of the inter-core connection again; if the accuracy of the m-th retrained second neural network reaches the preset accuracy threshold, execute Step S306.
[0075] Step S306: When the accuracy of the m-th retrained second neural network reaches the preset accuracy threshold, determine the inter-core pruning mapping scheme corresponding to the m-th retrained second neural network as the mapping scheme for the n-th inter-core reconstruction.
[0076] Figure 4 FIG. is a schematic diagram showing the change of the mapping scheme during the inter-core pruning optimization provided by the embodiments of the present disclosure. In Figure 4 , the boxes represent processing cores and the black dots represent processing nodes. Figure 4 The mapping scheme shown includes four processing cores, namely the first processing core 41, the second processing core 42, the third processing core 43, and the fourth processing core 44. In Figure 4 , (a) is the mapping scheme before inter-core pruning. There is an inter-core connection 401 between the first processing core 41 and the second processing core 42, two inter-core connections 402A and 402B between the first processing core 41 and the third processing core 43, one inter-core connection 403 between the first processing core 41 and the fourth processing core 44, two inter-core connections 404A and 404B between the second processing core 42 and the third processing core 43, two inter-core connections 405A and 405B between the second processing core 42 and the fourth processing core 44, and one inter-core connection 406 between the third processing core 43 and the fourth processing core 44.
[0077] Assume that in step S303, it is determined that the connection costs of the inter-core connections 402A, 403, and 405B are greater than the m-th cost threshold, and the absolute value of the connection weight is less than the m-th first weight threshold. Therefore, the inter-core connections 402A, 403, and 405B are deleted (i.e., inter-core pruning), and the second neural network obtained after pruning is retrained. When the accuracy of the m-th retrained second neural network reaches the preset accuracy threshold, the inter-core connections of the obtained m-th inter-core pruning mapping scheme include the inter-core connection 401, the inter-core connection 402B, the inter-core connection 404A, the inter-core connection 404B, the inter-core connection 405A, and the inter-core connection 406, as shown in Figure 4 FIG. (b).
[0078] By optimizing the mapping scheme through the above inter-core pruning method, the mapping scheme is made more cohesive, reducing the inter-core connections and the data transfer between cores, thereby reducing the routing pressure of the mapping scheme.
[0079] In some mapping schemes, there are cases where the inter-core connections are long, that is, the starting processing core and the ending processing core corresponding to the inter-core connection are far apart. The longer the inter-core connection, the greater the routing burden, and the shorter the inter-core connection, the smaller the routing burden. Therefore, replacing the longer inter-core connections with shorter inter-core connections can also reduce the data transfer between cores.
[0080] In some embodiments, reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction includes: reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction using Hebb's rule to obtain the mapping scheme of the n-th inter-core reconstruction.
[0081] Figure 5 The flowchart of another method for optimizing inter-core reconstruction provided by the embodiments of the present disclosure. This method for optimizing inter-core reconstruction optimizes the mapping scheme through inter-core reconnection. As Figure 5 shown, reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction using Hebb's rule to obtain the mapping scheme of the n-th inter-core reconstruction includes:
[0082] Step S501: Determine the processing cores in the mapping scheme of the (n - 1)-th inter-core reconstruction and the processing nodes in each processing core.
[0083] Wherein, each processing node includes at least one neuron.
[0084] Step S502: Reconstruct the inter-core connections between any two processing nodes according to a pre-agreed connection probability to obtain a third mapping scheme.
[0085] Wherein, the two processing nodes involved in the reconstructed inter-core connections are located in different processing cores.
[0086] In some embodiments, the connection probability is negatively correlated with the distance between the processing cores corresponding to the two processing nodes, that is, the closer the distance between the processing nodes in different processing cores, the higher the connection probability, and vice versa, the lower the connection probability. The inter-core connections determined using this connection probability can minimize the length of the inter-core connections, thereby reducing the routing pressure.
[0087] Step S503: Retrain the third neural network corresponding to the third mapping scheme using the training data set to obtain the retrained third neural network.
[0088] In some embodiments, the loss function used for training the third neural network using the training data set includes a connection cost regularization term and a connection weight regularization term.
[0089] Wherein, the loss function Loss for training can adopt the function shown in Equation (2), which will not be elaborated here.
[0090] Step S504: Delete the inter-core connections in the retrained third neural network whose connection weights are less than a preset second weight threshold to obtain the mapping scheme of the n-th inter-core reconstruction.
[0091] Figure 6A schematic diagram of the change in the mapping scheme during the inter-core reconstruction provided by the embodiments of the present disclosure. In Figure 6 , the boxes represent processing cores, and the black dots represent processing nodes. Figure 6 The mapping scheme shown in Figure 6 includes six processing cores, namely, a first processing core 61, a second processing core 62, a third processing core 63, a fourth processing core 64, a fifth processing core 65, and a sixth processing core 66. As shown in (a) of Figure 6 , in the original mapping scheme before inter-core reconstruction, an inter-core connection 601 is provided between the first processing core 61 and the second processing core 62, an inter-core connection 602 is provided between the first processing core 61 and the fourth processing core 64, an inter-core connection 603 is provided between the third processing core 63 and the fourth processing core 64, an inter-core connection 604 is provided between the second processing core 62 and the fifth processing core 65, and an inter-core connection 605 is provided between the fifth processing core 65 and the sixth processing core 66. Among them, the first processing core 61 and the fourth processing core 64 are relatively far apart, and the corresponding inter-core connection 602 is relatively long. After inter-core reconstruction, an inter-core connection 607 is reconstructed between the first processing core 61 and the sixth processing core 66, and an inter-core connection 608 is reconstructed between the fourth processing core 64 and the fifth processing core 65. The first processing core 61 and the sixth processing core 66, as well as the fourth processing core 64 and the fifth processing core 65, are relatively close, and the lengths of the inter-core connections 607 and 608 are relatively short. Using the inter-core connections 607 and 608 to replace the inter-core connection 602 shortens the length of the inter-core connection and reduces the routing burden.
[0092] In step S104, the mapping scheme of the nth inter-core reconstruction can be directly determined as the second mapping scheme. However, since the second mapping scheme is only a mapping scheme obtained by reconstructing the inter-core connection, and reconstructing the inter-core connection may cause a decrease in the accuracy of the neural network, especially the inter-core pruning mapping scheme is likely to cause a decrease in the accuracy of the neural network. Therefore, some means can be adopted to improve the accuracy of the neural network.
[0093] It should be noted that the first mapping scheme also includes intra-core connections, that is, intra-core connections between processing nodes within each processing core, and each processing node includes at least one neuron. In some embodiments, the mapping scheme of the nth inter-core reconstruction can be optimized by using the intra-core reconnection method to determine the optimized second mapping scheme, improving the capacity of the neural network, and thus improving the accuracy of the neural network.
[0094] Figure 7 A flowchart of an intra-core reconnection optimization method provided by the embodiments of the present disclosure. The intra-core reconnection optimization method further optimizes the mapping scheme of the inter-core reconstruction by using the intra-core reconnection method to compensate for the impact of the inter-core reconstruction on the neural network. As shown in Figure 7As shown, in step S104, according to the mapping scheme of the nth inter-core reconstruction, determine the optimized second mapping scheme, including:
[0095] Step S701, determine the intra-core connections between the processing nodes within each processing core in the mapping scheme of the (r - 1)th intra-core reconnection.
[0096] Where r ≥ 1 and r is an integer, and the mapping scheme of the 0th intra-core reconnection is the mapping scheme of the nth inter-core reconstruction.
[0097] In some embodiments, first determine all the processing cores in the mapping scheme of the (r - 1)th intra-core reconnection, and then determine the processing nodes within each processing core and the intra-core connections between the processing nodes.
[0098] Step S702, connect the unconnected processing nodes in at least one processing core to obtain the intermediate mapping scheme of the rth intra-core reconnection.
[0099] In some embodiments, all the unconnected processing nodes in one processing core or each processing core can be connected, and then some of the intra-core connections are randomly deleted, or the unconnected processing nodes in one processing core or each processing core are selectively connected to obtain the intermediate mapping scheme of the rth intra-core reconnection.
[0100] Step S703, retrain the fourth neural network corresponding to the intermediate mapping scheme of the rth intra-core reconnection using the training dataset to obtain the fourth neural network after the rth retraining.
[0101] In some embodiments, the loss function used during retraining can be the original loss function, such as Loss′ mentioned above.
[0102] Step S704, use the intra-core reconnection mapping scheme corresponding to the fourth neural network after the rth retraining as the mapping scheme of the rth intra-core reconnection.
[0103] Step S705, determine whether the fourth neural network after the rth retraining meets the intra-core optimization condition. If not, return to step S702; if so, execute step S706.
[0104] Step S706, when the mapping scheme of the rth intra-core reconnection meets the intra-core optimization condition, determine the mapping scheme of the rth intra-core reconnection as the optimized second mapping scheme.
[0105] In the embodiments of the present disclosure, a mapping scheme is reconstructed by means of in-core reconnecting to improve the capacity of the neural network, so as to compensate for the decrease in the accuracy of the neural network caused by inter-core reconstruction, such that the mapping scheme not only reduces the routing load but also does not affect the accuracy of the neural network. In the optimized second mapping scheme, the inter-core connections and in-core connections of the first neural network are re-determined, and the inter-core connections are reduced by increasing the in-core connections, thereby reducing the routing pressure of the many-core system while meeting the accuracy requirements of the first neural network. When the optimized second mapping scheme is mapped to the many-core system and the many-core system executes the first neural network, both the accuracy requirements of the first neural network can be met and the routing pressure of the many-core system can be reduced.
[0106] Figure 8 It is a flowchart of an in-core reconnecting optimization method provided by the embodiments of the present disclosure. This in-core reconnecting optimization method also further optimizes the mapping scheme of inter-core reconstruction by means of in-core reconnecting to compensate for the impact of inter-core reconstruction on the neural network. The difference lies in the construction method of in-core reconnecting. As Figure 8 shown, after obtaining the fourth neural network after the r-th retraining in step S704, when determining the optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction in step S104, it further includes:
[0107] Step S801: Delete the in-core connections with connection weights less than a preset fifth weight threshold in the (p - 1)-th in-core mapping scheme to obtain the p-th in-core mapping scheme.
[0108] Where p ≥ 1 and p is an integer, and the 0-th in-core mapping scheme is the in-core reconnecting mapping scheme corresponding to the fourth neural network after the r-th retraining. Except for the first in-core reconnecting mapping scheme, other in-core reconnecting mapping schemes are obtained by deleting the in-core connections with connection weights less than a preset third weight threshold in the (p - 1)-th in-core reconnecting mapping scheme.
[0109] Step S802: Retrain the fifth neural network corresponding to the p-th in-core mapping scheme using the training data set to obtain the fifth neural network after the p-th retraining.
[0110] In some embodiments, the loss function used during retraining can be the original loss function, such as Loss′ mentioned above.
[0111] Step S803: Determine whether the fifth neural network after the p-th retraining meets the preset accuracy condition. If not, return to step S801; if so, execute step S804.
[0112] Step S804: When the p-th retrained fifth neural network meets the preset accuracy condition, determine the in-core mapping scheme corresponding to the p-th retrained fifth neural network as the mapping scheme for the r-th in-core reconnection. The preset accuracy condition can be, for example, that the processing accuracy of the neural network for test samples reaches an accuracy threshold. The present disclosure does not limit the specific setting method of the accuracy condition.
[0113] In the embodiments of the present disclosure, the p-th in-core mapping scheme is determined based on the connection weights on the basis of the (p - 1)-th in-core reconnection mapping scheme. Compared with Figure 7 the in-core reconnection optimization method shown, it can enable the fifth neural network to meet the in-core optimization condition faster and shorten the training time.
[0114] Figure 9 FIG. is a schematic diagram showing the change process of the mapping scheme during the in-core reconnection optimization provided by the embodiments of the present disclosure. In Figure 9 , the square boxes represent processing cores, and the black dots represent processing nodes. Figure 9 The mapping scheme shown in FIG. includes four processing cores, namely the first processing core 91, the second processing core 92, the third processing core 93, and the fourth processing core 94. In Figure 9 FIG. (a) is the mapping scheme before in-core reconnection. In the first processing core 91, there is an in-core connection 9101 between the first processing node and the third processing node. In the second processing core 92, there is an in-core connection 9201 between the second processing node and the third processing node. In the third processing core 93, there is an in-core connection 9301 between the first processing node and the third processing node. In the fourth processing core 94, there is an in-core connection 9401 between the first processing node and the fourth processing node, and there is an in-core connection 9402 between the third processing node and the fourth processing node.
[0115] After in-core reconnection optimization, as shown in Figure 9 FIG. (b), in the first processing core 91, an in-core connection 9102 between the first processing node and the second processing node is added. In the second processing core 92, an in-core connection 9202 between the first processing node and the third processing node is added. In the fourth processing core 94, an in-core connection 9403 between the first processing node and the second processing node is added, and an in-core connection 9404 between the second processing node and the fourth processing node is added. The in-core connections 9102, 9202, 9403, and 9404 are used to compensate for the impact of inter-core pruning on the accuracy of the neural network.
[0116] In some embodiments, after determining the second mapping scheme in one round of inter-core reconstruction and intra-core reconnection, especially the second mapping scheme determined when the number of cycles of inter-core optimization reaches a preset number of cycles, it can be further optimized to further reduce the routing pressure of the many-core system and improve the accuracy of the neural network simultaneously.
[0117] Figure 10 The flowchart of a mapping scheme optimization method provided by an embodiment of the present disclosure. This mapping scheme optimization method is to perform at least one round of inter-core reconstruction and intra-core reconnection after obtaining the second mapping scheme to further reduce the routing pressure of the many-core system and improve the accuracy of the neural network simultaneously. As Figure 10 shown, after determining the optimized second mapping scheme, it further includes:
[0118] Step S1001, perform inter-core reconstruction optimization on the mapping scheme optimized in the (k - 1)-th round to obtain an inter-core optimized mapping scheme optimized in the k-th round that meets the inter-core optimization conditions.
[0119] Where k ≥ 2 and k is an integer, and the mapping scheme optimized in the first round is the second mapping scheme.
[0120] In step S1001, for the specific method of performing inter-core reconstruction optimization on the mapping scheme optimized in the (k - 1)-th round, refer to steps S102 to S104, which will not be elaborated here.
[0121] Step S1002, perform intra-core reconnection optimization on the inter-core optimized mapping scheme optimized in the k-th round to obtain an intra-core optimized mapping scheme optimized in the k-th round that meets the intra-core optimization conditions.
[0122] In step S1002, for the specific method of performing intra-core reconnection optimization on the inter-core optimized mapping scheme optimized in the k-th round, refer to steps S501 to S504, which will not be elaborated here.
[0123] Step S1003, determine whether the intra-core optimized mapping scheme optimized in the k-th round meets the preset target conditions. If so, execute step S1004; if not, return to step S1001.
[0124] Step S1004, when the intra-core optimized mapping scheme optimized in the k-th round meets the preset target conditions, determine the intra-core optimized mapping scheme optimized in the k-th round as the target mapping scheme.
[0125] Figure 11 The schematic diagram of the change process of the mapping scheme during two rounds of inter-core reconstruction optimization and intra-core reconnection optimization provided by an embodiment of the present disclosure. In Figure 11 it, the boxes represent processing cores, and the black dots represent processing nodes. Figure 11The mapping scheme shown in the figure includes four processing cores, namely the first processing core 1101, the second processing core 1102, the third processing core 1103, and the fourth processing core 1104.
[0126] In the original mapping scheme before inter-core reconstruction, as Figure 11 shown in (a) of the figure, there is an inter-core connection 111 between the first processing core 1101 and the second processing core 1102, there are two inter-core connections 112A and 112B between the first processing core 1101 and the third processing core 1103, there is an inter-core connection 113 between the first processing core 1101 and the fourth processing core 1104, there are two inter-core connections 114A and 114B between the second processing core 1102 and the third processing core 1103, there are two inter-core connections 115A and 115B between the second processing core 1102 and the fourth processing core 1104, and there is an inter-core connection 116 between the third processing core 1103 and the fourth processing core 1104.
[0127] Inside the first processing core 1101, there is an intra-core connection 1301 between the first processing node and the third processing node. Inside the second processing core 1102, there is an intra-core connection 2301 between the second processing node and the third processing node. Inside the third processing core 1103, there is an intra-core connection 3301 between the first processing node and the third processing node. Inside the fourth processing core 1104, there is an intra-core connection 4301 between the first processing node and the fourth processing node, and there is an intra-core connection 4302 between the third processing node and the fourth processing node.
[0128] After the first round of inter-core reconstruction optimization, the inter-core connections 112A, 113, and 115B are deleted. The mapping scheme includes the inter-core connections 111, 112B, 114A, 114B, 115A, and 116, as Figure 11 shown in (b) of the figure.
[0129] After the first round of intra-core reconnection optimization, inside the first processing core 1101, an intra-core connection 1302 is added between the first processing node 1201 and the third processing node 1203. Inside the fourth processing core 1104, an intra-core connection 4303 is added between the first processing node 4201 and the second processing node 4202, and an intra-core connection 4304 is added between the second processing node 422 and the fourth processing node 4204, as Figure 11 shown in (c) of the figure.
[0130] After the second round of inter-core reconstruction optimization, the inter-core connection 114A between the second processing core 1102 and the third processing core 1103 is deleted. The mapping scheme includes inter-core connections 111, 112B, 114B, 115A, and 116, as Figure 11 shown in (d) of
[0131] After the first round of intra-core reconnection optimization, within the first processing core 1101, the intra-core connection 1303 between the second processing node and the third processing node is added. Within the second processing core 1102, the intra-core connection 2302 between the first processing node and the second processing node is added. Within the third processing core 1103, the intra-core connection 3302 between the second processing node and the third processing node is added, as Figure 11 shown in (e) of
[0132] The embodiments of the present disclosure further optimize the mapping scheme of the neural network through multiple rounds of inter-core reconstruction and intra-core reconnection, further reducing the routing pressure of the many-core system while improving the accuracy of the neural network.
[0133] Each neural network (the first to the fifth neural networks) provided by the embodiments of the present disclosure is used to execute any one of image processing tasks, speech processing tasks, text processing tasks, and video processing tasks. The present disclosure does not limit the specific task types executed by the neural network.
[0134] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0135] Figure 12 It is a block diagram of an optimization device for a mapping scheme based on a many-core system provided by an embodiment of the present disclosure.
[0136] Referring to Figure 12 , the embodiments of the present disclosure provide an optimization device for a mapping scheme based on a many-core system. The optimization device for the mapping scheme based on the many-core system includes:
[0137] An acquisition module 121, configured to acquire an initial first mapping scheme. The first mapping scheme is used to map a to-be-executed first neural network to multiple processing cores of the many-core system. Each processing core is used to execute at least one neuron of the first neural network. The first mapping scheme includes inter-core connections between the respective processing cores.
[0138] A reconstruction module 122, configured to reconstruct the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction, where n ≥ 1 and n is an integer, and the mapping scheme of the 0-th inter-core reconstruction is the first mapping scheme.
[0139] A determination module 123, configured to determine whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition;
[0140] A determination module 124, configured to, when the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, determine an optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction.
[0141] The mapping scheme optimization device based on a many-core system provided by an embodiment of the present disclosure can be used to implement any one of the mapping scheme optimization methods based on a many-core system provided by the present disclosure. For the corresponding technical solutions and descriptions, refer to the corresponding records in the method section, which will not be elaborated here.
[0142] In the embodiment provided by the present disclosure, an acquisition module obtains an initial first mapping scheme, a reconstruction module reconstructs the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction, a determination module determines whether the mapping scheme of the n-th inter-core reconstruction meets the inter-core optimization condition, and a determination module, when the mapping scheme of the n-th inter-core reconstruction meets the preset inter-core optimization condition, determines an optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction. This device optimizes the mapping scheme of a neural network, realizes the cohesion of the mapping scheme, reduces the data transfer between cores, alleviates the routing pressure of the many-core system, and improves the efficiency of the many-core system in executing a neural network.
[0143] Figure 13 It is a block diagram of an electronic device provided by an embodiment of the present disclosure.
[0144] Referring to Figure 13 , an embodiment of the present disclosure provides an electronic device, which includes: at least one processor 13001; and a memory 13002 and an I / O interface 13003 communicatively connected to the at least one processor 13001; wherein, the memory 13002 stores one or more computer programs executable by the at least one processor 13001, and the one or more computer programs are executed by the at least one processor 13001 so that the at least one processor 13001 can execute the above-mentioned mapping scheme optimization method based on a many-core system.
[0145] In some embodiments, the electronic device may be a brain-inspired chip. Since the brain-inspired chip can adopt a vectorized computing method and needs to load parameters such as weight information of a neural network model from an external memory, such as a Double Data Rate (DDR) synchronous dynamic random access memory. Therefore, the batch processing adopted in the embodiments of the present disclosure has relatively high operation efficiency.
[0146] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. Wherein, the computer program, when executed by a processor / processing core, implements the above-mentioned mapping scheme optimization method based on a many-core system. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0147] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned mapping scheme optimization method based on a many-core system.
[0148] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be executed by several physical components in cooperation. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium).
[0149] As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. Additionally, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable program instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery media.
[0150] The computer-readable program instructions described herein can be downloaded to each computing / processing device from a computer-readable storage medium or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0151] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0152] The computer program product described herein may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0153] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0154] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0155] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0156] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0157] Example embodiments have been disclosed herein, and although specific terms are employed, they are used in a general and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, the features, characteristics, and / or elements described in connection with a particular embodiment may be used singly or in combination with those described in connection with other embodiments. Accordingly, it will be understood by those skilled in the art that various forms and details may be changed without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. An optimization method for mapping schemes based on many-core systems, characterized in that, Including: Obtain an initial first mapping scheme, where the first mapping scheme is used to map a to-be-executed first neural network to multiple processing cores of the many-core system, each processing core is used to execute at least one neuron of the first neural network, and the first mapping scheme includes inter-core connections between the respective processing cores; Reconstruct the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction, where n≥1 and n is an integer, and the mapping scheme of the 0-th inter-core reconstruction is the first mapping scheme; Determine whether the mapping scheme of the n-th inter-core reconstruction satisfies the inter-core optimization condition; When the mapping scheme of the n-th inter-core reconstruction satisfies the preset inter-core optimization condition, determine an optimized second mapping scheme according to the mapping scheme of the n-th inter-core reconstruction; The reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction includes: Determine the connection cost and connection weight of each inter-core connection in the mapping scheme of the (n - 1)-th inter-core reconstruction, where the connection cost indicates the cost of the connection between the processing cores corresponding to the inter-core connection, and the connection weight indicates the weight of the connection between the neurons corresponding to the inter-core connection; Set the m-th cost threshold and the m-th first weight threshold of the inter-core connection, where m≥1 and m is an integer; Delete the inter-core connections whose connection cost is greater than the m-th cost threshold and the absolute value of the connection weight is less than the m-th first weight threshold to obtain the m-th inter-core pruning mapping scheme; Retrain a second neural network corresponding to the m-th inter-core pruning mapping scheme using a training data set to obtain the m-th retrained second neural network; When the accuracy of the m-th retrained second neural network reaches the preset accuracy threshold, determine the inter-core pruning mapping scheme corresponding to the m-th retrained second neural network as the mapping scheme of the n-th inter-core reconstruction.
2. The optimization method for mapping schemes based on many-core systems according to claim 1, characterized in that, The connection cost of the inter-core connection is determined based on the per-unit-time data communication volume of the inter-core connection and the distance between the starting processing core and the ending processing core corresponding to the inter-core connection.
3. The optimization method for mapping schemes based on many-core systems according to claim 1, characterized in that, The reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction to obtain the mapping scheme of the n-th inter-core reconstruction includes: Reconstruct the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction using Hebb's rule to obtain the mapping scheme of the n-th inter-core reconstruction.
4. The optimization method for mapping schemes based on many-core systems according to claim 3, characterized in that, The reconstructing the inter-core connections in the mapping scheme of the (n - 1)-th inter-core reconstruction using Hebb's rule to obtain the mapping scheme of the n-th inter-core reconstruction includes: Determine the processing cores in the mapping scheme of the (n - 1)-th inter-core reconstruction and the processing nodes in each processing core, where each processing node includes at least one neuron; Reconstruct the inter-core connections between any two processing nodes according to a pre-agreed connection probability to obtain a third mapping scheme; where the two processing nodes are located in different processing cores; Retrain the third neural network corresponding to the third mapping scheme using the training data set to obtain the retrained third neural network; Delete the inter-core connections with connection weights less than a preset second weight threshold in the retrained third neural network to obtain the mapping scheme of the nth inter-core reconstruction.
5. The optimization method for mapping schemes based on many-core systems according to claim 4, characterized in that, The connection probability is negatively correlated with the distance between the processing cores corresponding to the two processing nodes.
6. The optimization method for mapping schemes based on many-core systems according to claim 1, characterized in that, The loss function used to train the second neural network using the training data set includes a connection cost regularization term and a connection weight regularization term.
7. The optimization method for mapping schemes based on many-core systems according to claim 1, characterized in that, The inter-core optimization condition is that the number of cycles of the inter-core optimization reaches a preset number of cycles or the mapping scheme of the inter-core reconstruction meets the routing load requirement.
8. The optimization method for mapping schemes based on many-core systems according to claim 1, characterized in that, The first mapping scheme further includes intra-core connections between processing nodes within each processing core, and each processing node includes at least one neuron; Among them, determining the optimized second mapping scheme according to the mapping scheme of the nth inter-core reconstruction includes: Determine the intra-core connections between the processing nodes within each processing core in the mapping scheme of the (r - 1)th intra-core reconnection, where r ≥ 1 and r is an integer. Among them, the mapping scheme of the 0th intra-core reconnection is the mapping scheme of the nth inter-core reconstruction; Connect the unconnected processing nodes in at least one of the processing cores to obtain the intermediate mapping scheme of the rth intra-core reconnection; Retrain the fourth neural network corresponding to the intermediate mapping scheme of the rth intra-core reconnection using the training data set to obtain the fourth neural network after the rth retraining, and use the intra-core reconnection mapping scheme corresponding to the fourth neural network after the rth retraining as the mapping scheme of the rth intra-core reconnection; When the mapping scheme of the rth intra-core reconnection meets the intra-core optimization condition, determine the mapping scheme of the rth intra-core reconnection as the optimized second mapping scheme.
9. The optimization method for mapping scheme based on a many-core system according to claim 8, wherein, After obtaining the fourth neural network after the rth retraining, determining the optimized second mapping scheme according to the mapping scheme of the nth inter-core reconstruction further includes: Delete the intra-core connections with connection weights less than a preset fifth weight threshold in the (p - 1)th intra-core mapping scheme to obtain the pth intra-core mapping scheme, where p ≥ 1 and p is an integer. Among them, the 0th intra-core mapping scheme is the intra-core reconnection mapping scheme corresponding to the fourth neural network after the rth retraining; Retrain the fifth neural network corresponding to the pth intra-core mapping scheme using the training data set to obtain the fifth neural network after the pth retraining; When the fifth neural network after the pth retraining meets the preset accuracy condition, determine the intra-core mapping scheme corresponding to the fifth neural network after the pth retraining as the mapping scheme of the rth intra-core reconnection.
10. The optimization method for mapping scheme based on a many-core system according to claim 8 or 9, wherein, After determining the optimized second mapping scheme, the method further includes: Perform inter-core reconstruction optimization on the mapping scheme optimized in the (k - 1)th round to obtain the inter-core optimization mapping scheme optimized in the kth round that meets the inter-core optimization condition, where k ≥ 2 and k is an integer. Among them, the mapping scheme optimized in the first round is the second mapping scheme; Perform in - core reconnection optimization on the inter - core optimized mapping scheme after the k - th round of optimization to obtain an in - core optimized mapping scheme after the k - th round that meets the in - core optimization conditions. When the in - core optimized mapping scheme after the k - th round meets the preset target conditions, determine the in - core optimized mapping scheme after the k - th round as the target mapping scheme.
11. An optimization device for mapping scheme based on a many-core system, wherein, It includes: An acquisition module, configured to acquire an initial first mapping scheme, where the first mapping scheme is used to map a first neural network to be executed to multiple processing cores of the many - core system, and each processing core is used to execute at least one neuron of the first neural network. The first mapping scheme includes the inter - core connections between the respective processing cores. A reconstruction module, configured to reconstruct the inter - core connections in the mapping scheme of the (n - 1) - th inter - core reconstruction to obtain the mapping scheme of the n - th inter - core reconstruction, where n≥1 and n is an integer. Among them, the mapping scheme of the 0 - th inter - core reconstruction is the first mapping scheme. A judgment module, configured to judge whether the mapping scheme of the n - th inter - core reconstruction meets the inter - core optimization conditions. A determination module, configured to, when the mapping scheme of the n - th inter - core reconstruction meets the preset inter - core optimization conditions, determine an optimized second mapping scheme according to the mapping scheme of the n - th inter - core reconstruction. The step of reconstructing the inter - core connections in the mapping scheme of the (n - 1) - th inter - core reconstruction to obtain the mapping scheme of the n - th inter - core reconstruction includes: Determine the connection cost and connection weight of each inter - core connection in the mapping scheme of the (n - 1) - th inter - core reconstruction. The connection cost indicates the cost of the connection between the processing cores corresponding to the inter - core connection, and the connection weight indicates the weight of the connection between the neurons corresponding to the inter - core connection. Set the m - th cost threshold and the m - th first weight threshold of the inter - core connection, where m≥1 and m is an integer. Delete the inter - core connections whose connection cost is greater than the m - th cost threshold and whose absolute value of the connection weight is less than the m - th first weight threshold to obtain the m - th inter - core pruning mapping scheme. Retrain the second neural network corresponding to the m - th inter - core pruning mapping scheme using the training data set to obtain the m - th retrained second neural network. When the accuracy of the m - th retrained second neural network reaches the preset accuracy threshold, determine the inter - core pruning mapping scheme corresponding to the m - th retrained second neural network as the mapping scheme of the n - th inter - core reconstruction.
12. An electronic device, wherein, It includes: At least one processor; And A memory communicatively connected to the at least one processor; where The memory stores one or more computer programs executable by the at least one processor. The one or more computer programs are executed by the at least one processor so that the at least one processor can execute the mapping scheme optimization method based on a many - core system according to any one of claims 1 - 10.
13. A computer-readable storage medium, on which a computer program is stored, wherein, The computer program, when executed by the processor, implements the mapping scheme optimization method based on a many - core system according to any one of claims 1 - 10.
Citation Information
Patent Citations
NoC mapping method based on improved genetic algorithm
CN108153592A