Route generation method and device, many kernel system, computer readable medium
Patent Information
- Application Number
- CN202110647124.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-06-10
AI Technical Summary
[0048]本公开实施例提供一种路由生成方法,能够根据软件层面定义的核心对象簇的任务对接关系,确定各个核心对象簇中核心对象的连接关系,并根据该连接关系生成将多个核心对象簇映射到众核系统后的硬件路由,提升了众核系统路由生成的效率,便于将由多个核心对象簇映射到众核系统。
Smart Images

Figure CN115470174B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a route generation method, a route generation device, a many-core system, and a computer-readable medium. Background Technology
[0002] A many-core system can consist of at least one chip, each chip having multiple computing units. The smallest computing unit in each chip that can be independently scheduled and has complete computing capabilities is called a processing core.
[0003] In many-core systems, multiple processors can work together, and each processor core can run program instructions independently, leveraging parallel computing capabilities to accelerate program execution and provide multitasking capabilities.
[0004] Currently, there is an urgent need for an efficient method to generate routes for many-core systems. Summary of the Invention
[0005] This disclosure provides a route generation method, a route generation apparatus, a many-core system, and a computer-readable medium.
[0006] In a first aspect, embodiments of this disclosure provide a route generation method, including:
[0007] Based on the task docking relationship between the first core object cluster and the second core object cluster, the connection relationship between at least one first core object in the first core object cluster and at least one second core object in the second core object cluster is determined. Each first core object describes the configuration information of a processing core of the many-core system, and each second core object describes the configuration information of a processing core of the many-core system.
[0008] The hardware routing of the many-core system is determined based on the connection relationship.
[0009] In some embodiments, the step of determining the connection relationship between at least one first core object in the first core object cluster and at least one second core object in the second core object cluster, based on the task docking relationship between the first core object cluster and the second core object cluster, includes:
[0010] Obtain the output vector of the at least one first core object, where each output vector corresponds to one first core object;
[0011] The output vector is scheduled to obtain at least one input vector, and each input vector corresponds to a second core object;
[0012] By obtaining the corresponding input vector of each of the at least one second core objects, the input vector of each second core object is obtained.
[0013] The connection relationship is determined based on the correspondence between at least one output vector and at least one input vector.
[0014] In some embodiments, each first core object corresponds to at least one output vector, each output vector carries a source address, and each input vector carries the source address of its corresponding output vector; each second core object corresponds to at least one input vector, and each input vector carries a destination address; the step of determining the connection relationship based on the correspondence between at least one output vector and at least one input vector includes:
[0015] The source address carried by the corresponding input vector is obtained through the at least one second core object;
[0016] A correspondence between at least one source address and at least one destination address is determined to obtain a receive table, which represents the connection relationship.
[0017] In some embodiments, the step of determining the connection relationship based on the correspondence between at least one output vector and at least one input vector further includes:
[0018] A sending table is generated based on the receiving table, the sending table representing the connection relationship, and one source address in the sending table corresponds to at least one destination address.
[0019] In some embodiments, the step of scheduling the output vectors to obtain at least one input vector, which includes multiple output vectors, comprises:
[0020] Generate an output matrix based on the multiple output vectors;
[0021] Multiple input vectors are generated based on the output matrix.
[0022] In some embodiments, the step of generating an output matrix based on a plurality of output vectors includes:
[0023] The output vectors are arranged to construct the output matrix.
[0024] In some embodiments, the step of generating a plurality of input vectors based on the output matrix includes:
[0025] Arrange the output matrix to obtain the input matrix;
[0026] The input matrix is split into multiple input vectors.
[0027] In some embodiments, the step of obtaining the output vector of the at least one first core object includes:
[0028] Call the output function of each of the first core objects to obtain the output vector of each of the first core objects.
[0029] In some embodiments, the step of obtaining the corresponding input vectors of the at least one second core object includes:
[0030] Call the input function of each of the second core objects to obtain the input vector of each of the second core objects.
[0031] In some embodiments, each output vector carries a time identifier, the time identifier representing the time when one of the plurality of second core objects performs the calculation corresponding to the output vector; each input vector carries the time identifier of its corresponding output vector; the route generation method further includes:
[0032] The time identifier carried by the corresponding input vector is obtained by using multiple second core objects;
[0033] Determine the correspondence between the multiple time markers and the multiple input vectors.
[0034] In some embodiments, the second core object corresponds to multiple storage subspaces, and the computation corresponding to at least one of the input vectors stored in the storage subspaces is performed at the same time; the step of determining the correspondence between the multiple time markers and the multiple input vectors includes:
[0035] Each storage subspace is determined to store the input vector corresponding to each of the time markers, and a calculation time table representing the correspondence between the time markers and the storage subspaces is generated.
[0036] In some embodiments, the step of determining the hardware route of the many-core system based on the connection relationship includes:
[0037] The first core object cluster and the second core object cluster are respectively mapped to the many-core system, with each first core object corresponding to one processing core and each second core object corresponding to one processing core;
[0038] A hardware routing table is generated based on the correspondence between at least one first core object and at least one processing core, the correspondence between at least one second core object and at least one processing core, and the topology of the many-core system, to represent the hardware routing.
[0039] Secondly, embodiments of this disclosure provide a route generation apparatus, comprising:
[0040] One or more processors;
[0041] A memory having stored one or more programs that, when executed by one or more processors, enable the one or more processors to implement any of the route generation methods described in the first aspect of the present disclosure.
[0042] One or more I / O interfaces are connected between the processor and the memory and configured to enable information interaction between the processor and the memory.
[0043] Thirdly, embodiments of this disclosure provide a many-core system, including:
[0044] Multiple processing cores; and
[0045] The on-chip network is configured to interact with data between the multiple processing cores and external data;
[0046] One or more processing cores store one or more instructions, which are executed by one or more processing cores to enable the one or more processing cores to perform any of the route generation methods described in the first aspect of the present disclosure.
[0047] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the route generation methods described in the first aspect of this disclosure.
[0048] This disclosure provides a route generation method that can determine the connection relationship of core objects in each core object cluster based on the task docking relationship of core object clusters defined at the software level, and generate hardware routes that map multiple core object clusters to the many-core system based on the connection relationship. This improves the efficiency of route generation in the many-core system and facilitates the mapping of multiple core object clusters to the many-core system.
[0049] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0050] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0051] Figure 1 This is a flowchart of a route generation method according to an embodiment of this disclosure;
[0052] Figure 2 This is a schematic diagram of a task docking in an embodiment of this disclosure;
[0053] Figure 3 This is a flowchart of some steps in another route generation method in this disclosure embodiment;
[0054] Figure 4 This is a flowchart of some steps in another route generation method according to an embodiment of this disclosure;
[0055] Figure 5 This is a schematic diagram illustrating one embodiment of generating a display lookup table in this disclosure.
[0056] Figure 6 This is a flowchart of some steps in another route generation method in this disclosure embodiment;
[0057] Figure 7 This is a block diagram of a route generation device according to an embodiment of the present disclosure;
[0058] Figure 8 This is a block diagram of a many-core system according to an embodiment of the present disclosure. Detailed Implementation
[0059] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0060] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0061] As used herein, the term "and / or" includes any and all combinations of one or more related enumerated items.
[0062] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0063] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0064] Firstly, referring to Figure 1 This disclosure provides a route generation method, including:
[0065] In step S100, based on the task docking relationship between the first core object cluster and the second core object cluster, the connection relationship between at least one first core object in the first core object cluster and at least one second core object in the second core object cluster is determined. Each first core object describes the configuration information of a processing core of the many-core system, and each second core object describes the configuration information of a processing core of the many-core system.
[0066] In step S200, the hardware route of the many-core system is determined based on the connection relationship.
[0067] In this embodiment of the disclosure, the processing core of the many-core system can be regarded as a reconfigurable functional unit. That is, each processing core has a basic function and can be configured as a functional core with a specific function. Multiple processing cores configured with a specific function can form a processing core cluster with a specific function.
[0068] The core object is instantiated from the processing core class. The core object is a software-level description of the processing core of the many-core system. The core object cluster, consisting of at least one core object, is a software-level description of the processing core cluster consisting of at least one processing core.
[0069] This disclosure does not impose any special limitations on the functional core class. For example, there can be multiple functional core classes, each defining functional cores with different functionalities. In some embodiments, multiple first core objects can be instantiated from different functional core classes; multiple second core objects can also be instantiated from different functional core classes.
[0070] In this embodiment, the first core object cluster and the second core object cluster refer to two core object clusters with a task docking relationship. The task docking relationship means that the first core object cluster performs a computational task and transmits the output data to the second core object cluster, which then receives the output data and performs its own computational task. It should also be noted that the first core object cluster and the second core object cluster can be any pair of core object clusters with a task docking relationship. For example, when building a neural network using multiple core object clusters, and each core object cluster executes a computational node of the neural network, the first core object cluster and the second core object cluster correspond to a computational node in an adjacent layer of the neural network, respectively. In the program built from multiple core object clusters, step S100 determines the connection relationship between any two core object clusters, and step S200 generates hardware routes that map all core object clusters to the many-core system.
[0071] In some embodiments, the inputs, outputs, and routing of the first and second core object clusters may undergo transformations such as data transposition, sorting interchange, and neural network connection relationships. (Refer to...) Figure 2Taking Task A and Task B as examples, the output of Task A is D, and D is also the input of Task B. Task A is mapped to the core object cluster Cluster A, and Task A is executed as an input routing process (Cluster A In Routing), a computation process (Cluster A Computing), and an output routing process (Cluster A Out Routing). Task B is mapped to the core object cluster Cluster B, and Task B is executed as Cluster B In Routing, Cluster B Computing, and Cluster B Out Routing. Through route recording, Cluster A In Routing and Cluster B In Routing can be parallelized (InParallel), Cluster A Computing and Cluster B Computing can be parallelized, and Cluster A Out Routing and Cluster B Out Routing can be parallelized. Through route extraction, the docking process between Task A and Task B can be represented as an explicit lookup table (LUT), a data transfer process, and parallel Cluster A Computing and Cluster B Computing. Figure 2 This illustrates various permuting opportunities. For example, the output of Task A and the input of Task B may undergo transformations such as data transposition, sorting interchange, and neural network connections. Parallel routing inputs, parallel routing outputs, and parallel computation processes may also undergo transformations such as data transposition, sorting interchange, and neural network connections. In this embodiment, the connection relationship determined in step S100 can characterize at least one of the above transformations.
[0072] This disclosure provides a route generation method that can determine the connection relationship of core objects in each core object cluster based on the task docking relationship of core object clusters defined at the software level, and generate hardware routes that map multiple core object clusters to the many-core system based on the connection relationship. This improves the efficiency of route generation in the many-core system and facilitates the mapping of multiple core object clusters to the many-core system.
[0073] In some embodiments, an abstract routing model is introduced. This abstract routing model can obtain the output vector of a first core object and the input vector of a second core object, and supports various permutation opportunities such as input permutation, output permutation, and system-level permutation. For example, the abstract routing model supporting permutation opportunities means that it can schedule the output object of at least one first core object based on the input and output of the first and second core object clusters, as well as transformations that may occur during routing transmission, such as data transposition, sorting interchange, and neural network connection relationships, thereby establishing a logical connection between the first and second core objects in the abstract routing model. Based on this logical connection, the connection relationship between at least one first core object and at least one second core object can be determined.
[0074] Accordingly, refer to Figure 3 In some embodiments, step S100 includes:
[0075] In step S110, the output vector of the at least one first core object is obtained, and each output vector corresponds to one first core object;
[0076] In step S120, the output data is scheduled to obtain at least one input vector, and each input vector corresponds to a second core object;
[0077] In step S130, the corresponding input vector of each second core object is obtained by acquiring the input vector of each second core object through at least one second core object;
[0078] In step S140, the connection relationship is determined based on the correspondence between at least one output vector and at least one input vector.
[0079] In some embodiments, the abstract routing model can call the output functions of each first core object to obtain the output vector of each first core object.
[0080] In some embodiments, the abstract routing model can call the input functions of each second core object to obtain the input vector of each second core object.
[0081] This disclosure does not impose any special limitations on how the connection relationship is expressed. In some embodiments, the connection relationship is expressed as a look-up table (LUT). The LUT includes a receive table and a send table, and records the correspondence between source addresses and destination addresses.
[0082] It should be noted that, in this embodiment, the first core object cluster is considered as the sending end, and the second core object cluster is considered as the receiving end. Accordingly, relative to the second core object, the address corresponding to the output vector of the first core object is the source address; relative to the first core object, the address corresponding to the input vector of the second core object is the destination address.
[0083] In some embodiments, by simulating the data transmission process from the first core object to the second core object through an abstract routing model, a receive table can be determined, indexed by the destination address and with the source address as the table value. In the receive table, each destination address corresponds to one source address.
[0084] Accordingly, in some embodiments, each of the first core objects corresponds to at least one output vector, each output vector carries a source address, and each input vector carries the source address of its corresponding output vector; each of the second core objects corresponds to at least one input vector, and each input vector carries a destination address; refer to Figure 4 Step S140 includes:
[0085] In step S141, the source address carried by the corresponding input vector is obtained through the at least one second core object;
[0086] In step S142, the correspondence between at least one source address and at least one destination address is determined to obtain a receiving table, which represents the connection relationship.
[0087] In some embodiments, given a receive table, a send table indexed by the source address and with the destination address as the table value can be obtained by reverse engineering. Since the first core object and the second core object may have a multicast relationship—that is, the output vector of one first core object is transmitted to multiple second core objects—each source address in the send table corresponds to at least one destination address.
[0088] Accordingly, in some embodiments, reference is made to Figure 4 Step S140 further includes:
[0089] In step S143, a sending table is generated based on the receiving table. The sending table represents the connection relationship, and one source address in the sending table corresponds to at least one destination address.
[0090] This disclosure does not impose any special limitations on the form of the source and destination addresses. In some embodiments, each core object has a core identifier (e.g., a core number), each output vector of the core object has an output address, and each input vector has an input address; the destination address is calculated based on the core identifier and the input address, and the source address is calculated based on the core identifier and the output address. In some embodiments, a virtual sequence number is configured for each output vector of the core object as the output address of the output vector, and the source address is represented by a tuple of the core number and the virtual sequence number; simultaneously, a virtual sequence number is configured for each input vector as the input address of the input vector, and the destination address is represented by a tuple of the core number and the virtual sequence number.
[0091] In some embodiments, the output vector and its source address are obtained by calling the output function of the first core object; the input vector and its source address are obtained by investigating the input function of the second core object, and the source address and destination address are written into the LUT's receive table. Then, the LUT's send table can be obtained by reverse calculation using the receive table.
[0092] In some embodiments, such as Figure 5 As shown, Core 1 and Core 2 are the first core objects, and Core 3, Core 4, and Core 5 are the second core objects. Core 1 includes output vectors with source addresses (Src.) of (1,10) and (1,20); Core 2 includes output vectors with source addresses (2,15) and (2,30); Core 3 includes input vectors with destination addresses (Dest.) of (3,5), (3,10), and (3,15); Core 4 includes an input vector with a destination address of (4,5); and Core 5 includes input vectors with destination addresses of (5,10) and (5,15). The output vector of Core 2 with source addresses (2,15) and (2,30) is multicast to Core 3 and Core 5. Figure 5 The receiving table (Rtable) and the sending table (Ftable) derived from the receiving table are also shown.
[0093] This disclosure does not impose any special limitations on how step S120 is performed to schedule the output vectors and obtain the input vectors. For example, an abstract routing model can shape multiple output vectors into an output matrix, and then generate multiple input vectors based on the output matrix.
[0094] Accordingly, in some embodiments, the step of scheduling the output data to obtain multiple input vectors includes:
[0095] Generate an output matrix based on the multiple output vectors;
[0096] Multiple input vectors are generated based on the output matrix.
[0097] In some embodiments, during the output routing process, when generating an output matrix based on multiple output vectors, the multiple output vectors are first sorted, and the output matrix is generated based on the sorted output vectors. This achieves output permutation.
[0098] Accordingly, in some embodiments, the step of generating an output matrix based on the plurality of said output vectors includes:
[0099] The output vectors are arranged to construct the output matrix.
[0100] In some embodiments, during the input routing process, the output matrix can be arranged to obtain the input matrix, and then the input matrix can be split into multiple input vectors. This achieves input permutation.
[0101] Accordingly, in some embodiments, a plurality of output vectors are included, and the step of generating a plurality of input vectors based on the output matrix includes:
[0102] Arrange the output matrix to obtain the input matrix;
[0103] The input matrix is split into multiple input vectors.
[0104] In some embodiments, the first core object in the first core object cluster of the sending end carries a time identifier when sending data, and the second core object in the second core object cluster of the receiving end receives the time identifier while receiving data. Based on the time identifier, the corresponding calculation is performed when the time indicated by the time identifier arrives. Compared to some related technologies where the sending end sends data at a specific time so that the receiving end performs the corresponding calculation at that specific time, in this embodiment of the disclosure, the sending end does not need to send data according to a specific time.
[0105] In some embodiments, data corresponding to different times is stored in different storage subspaces at the receiving end, and the data received by the receiving end is written to the corresponding storage subspace according to the time identifier. When the corresponding time arrives, the core object of the receiving end reads the data required to perform the calculation at that time from the corresponding storage subspace.
[0106] In this embodiment of the disclosure, while generating the LUT, the correspondence between the time identifier and the input vector is also determined at the receiving end to generate the first core object cluster and the timing relationship of the first core object cluster.
[0107] In some embodiments, each output vector carries a time identifier, the time identifier representing the time when one of the plurality of second core objects performs the calculation corresponding to the output vector; each input vector carries the time identifier of its corresponding output vector; the route generation method further includes:
[0108] The time identifier carried by the corresponding input vector is obtained by using multiple second core objects;
[0109] Determine the correspondence between the multiple time markers and the multiple input vectors.
[0110] In some embodiments, a timetable is configured to represent the correspondence between time markers and storage subspaces.
[0111] Accordingly, in some embodiments, the second core object corresponds to multiple storage subspaces, and the computation corresponding to at least one of the input vectors stored in the storage subspaces is performed at the same time; the step of determining the correspondence between the multiple time markers and the multiple input vectors includes:
[0112] Each storage subspace is determined to store the input vector corresponding to each of the time markers, and a calculation time table representing the correspondence between the time markers and the storage subspaces is generated.
[0113] For example, data 1 and data 2 correspond to time t1, and data 3 corresponds to time t2. When the receiving end receives data 1 and data 2, it stores data 1 and data 2 in the storage subspace Memory 1; when the receiving end receives data 3, it stores data 3 in the storage subspace Memory 2. Simultaneously, a calculation timetable is generated. The calculation timetable stores entries representing the correspondence between time t1 and storage subspace Memory 1, and also stores entries representing the correspondence between time t2 and storage subspace Memory 2. When time t1 arrives, the second core object can read data 1 and data 2 from Memory 1 corresponding to time t1 according to the calculation timetable and perform the corresponding calculations; when time t2 arrives, the second core object can read data 3 from Memory 2 corresponding to time t2 according to the calculation timetable and perform the corresponding calculations.
[0114] In this embodiment of the disclosure, when the receiving end receives multiple data points in a neural network, the multiple data points can be placed on one axon, or each data point can be placed on a separate axon. This embodiment of the disclosure does not impose any special limitations on this.
[0115] In some embodiments, refer to Figure 6 Step S200 includes:
[0116] In step S210, the first core object cluster and the second core object cluster are mapped to the many-core system, with each first core object corresponding to one processing core and each second core object corresponding to one processing core.
[0117] In step S220, a hardware routing table is generated based on the correspondence between at least one first core object and at least one processing core, the correspondence between at least one second core object and at least one processing core, and the topology of the many-core system, to characterize the hardware routing.
[0118] In some embodiments, in step S100, multiple first core objects and multiple second core objects are regarded as a one-dimensional (1D) sequence. In step S210, the 1D core objects are mapped to a two-dimensional (2D) many-core system. When the correspondence between the first core objects and the second core objects is expressed by LUT, the routing path in the many-core system is determined according to the LUT, and a hardware routing table is generated.
[0119] In some embodiments, after generating the LUT in step S100, the correctness of the LUT table can also be verified. The abstract routing model calls the output functions of each core object cluster and sends data to each destination address according to the LUT's receive table. If data sent to a certain destination address cannot be delivered, an error is reported. In some embodiments, a dummy core can be used to collect invalid data that cannot be delivered.
[0120] In some embodiments, verifying the correctness of the LUT table can be performed through multicast detection. Multicast detection is performed by checking whether the input buffer addresses of core object clusters with multicast relationships are consistent. In some embodiments, when inconsistencies exist, the data is reordered by a rescheduling core with a new routing table.
[0121] Secondly, referring to Figure 7 This disclosure provides a route generation apparatus, including:
[0122] One or more processors 101;
[0123] The memory 102 stores one or more programs that, when executed by one or more processors, enable the one or more processors to implement any of the route generation methods described in the first aspect of the present disclosure.
[0124] One or more I / O interfaces 103 are connected between the processor and the memory and configured to enable information exchange between the processor and the memory.
[0125] The processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102, enabling information exchange between the processor 101 and the memory 102, including but not limited to a data bus (Bus).
[0126] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via bus 104, and thus connected to other components of the computing device.
[0127] Thirdly, referring to Figure 8 This disclosure provides a many-core system, including:
[0128] It includes multiple processing cores 201 and an on-chip network 202, wherein the multiple processing cores 201 are all connected to the on-chip network 202, and the on-chip network 202 is used to exchange data between the multiple processing cores and external data.
[0129] One or more processing cores 201 store one or more instructions, and the one or more instructions are executed by one or more processing cores 201 to enable one or more processing cores 201 to execute any of the route generation methods described in the first aspect of the present disclosure.
[0130] Fourthly, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the route generation methods described in the first aspect of this disclosure.
[0131] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0132] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for general illustrative purposes only and should not be construed as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A route generation method, comprising: Based on the task docking relationship between the first core object cluster and the second core object cluster, the connection relationship between at least one first core object in the first core object cluster and at least one second core object in the second core object cluster is determined. Each first core object describes the configuration information of a processing core of the many-core system, and each second core object describes the configuration information of a processing core of the many-core system. The hardware routing of the many-core system is determined based on the connection relationship; the connection relationship represents at least one of the following in the routing transmission: the input and output of the first core object cluster and the second core object cluster, as well as data transposition, sorting interchange, and neural network connection relationship transformation. The first core object and the second core object are instantiated based on the processing core class. The first core object cluster and the second core object cluster are used to characterize the software-level description of a processing core cluster consisting of at least one processing core. The step of determining the connection relationship between at least one first core object in the first core object cluster and at least one second core object in the second core object cluster, based on the task docking relationship between the first core object cluster and the second core object cluster, includes: Obtain the output vector of the at least one first core object, where each output vector corresponds to one first core object; The output vector is scheduled to obtain at least one input vector, and each input vector corresponds to a second core object; By obtaining the corresponding input vector of each of the at least one second core objects, the input vector of each second core object is obtained. The connection relationship is determined based on the correspondence between at least one output vector and at least one input vector.
2. The route generation method according to claim 1, wherein, Each of the first core objects corresponds to at least one of the output vectors, each of the output vectors carries a source address, and each of the input vectors carries the source address of its corresponding output vector. Each of the second core objects corresponds to at least one of the input vectors, and each of the input vectors carries a destination address; The step of determining the connection relationship based on the correspondence between at least one output vector and at least one input vector includes: The source address carried by the corresponding input vector is obtained through the at least one second core object; A correspondence between at least one source address and at least one destination address is determined to obtain a receive table, which represents the connection relationship.
3. The route generation method according to claim 2, wherein, The step of determining the connection relationship based on the correspondence between at least one output vector and at least one input vector further includes: A sending table is generated based on the receiving table, the sending table representing the connection relationship, and one source address in the sending table corresponds to at least one destination address.
4. The route generation method according to claim 1, wherein, The step of scheduling multiple output vectors to obtain at least one input vector includes: Generate an output matrix based on the multiple output vectors; Multiple input vectors are generated based on the output matrix.
5. The route generation method according to claim 4, wherein, The step of generating an output matrix based on the multiple output vectors includes: The output vectors are arranged to construct the output matrix.
6. The route generation method according to claim 4 or 5, wherein, The step of generating multiple input vectors based on the output matrix includes: Arrange the output matrix to obtain the input matrix; The input matrix is split into multiple input vectors.
7. The route generation method according to any one of claims 1 to 5, wherein, The steps for obtaining the output vector of at least one first core object include: Call the output function of each of the first core objects to obtain the output vector of each of the first core objects.
8. The route generation method according to any one of claims 1 to 5, wherein, The steps of obtaining the corresponding input vectors of each of the at least one second core objects include: Call the input function of each of the second core objects to obtain the input vector of each of the second core objects.
9. The route generation method according to claim 1, wherein, Each of the output vectors carries a time identifier, which represents the time when one of the plurality of second core objects performs the calculation corresponding to the output vector; Each input vector carries a time identifier for its corresponding output vector; The route generation method further includes: The time identifier carried by the corresponding input vector is obtained by using multiple second core objects; Determine the correspondence between the multiple time markers and the multiple input vectors.
10. The route generation method according to claim 9, wherein, The second core object corresponds to multiple storage subspaces, and the computation corresponding to at least one of the input vectors stored in the storage subspaces is executed at the same time. The step of determining the correspondence between the multiple time markers and the multiple input vectors includes: Each storage subspace is determined to store the input vector corresponding to each of the time markers, and a calculation time table representing the correspondence between the time markers and the storage subspaces is generated.
11. The route generation method according to any one of claims 1 to 5, 9, and 10, wherein, The steps for determining the hardware routing of the many-core system based on the connection relationship include: The first core object cluster and the second core object cluster are respectively mapped to the many-core system, with each first core object corresponding to one processing core and each second core object corresponding to one processing core; A hardware routing table is generated based on the correspondence between at least one first core object and at least one processing core, the correspondence between at least one second core object and at least one processing core, and the topology of the many-core system, to represent the hardware routing.
12. A route generation apparatus, comprising: One or more processors; A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to implement the route generation method according to any one of claims 1 to 11; One or more I / O interfaces are connected between the processor and the memory and configured to enable information interaction between the processor and the memory.
13. A many-core system, comprising: Multiple processing cores; as well as The on-chip network is configured to interact with data between multiple processing cores and with external data. One or more processing cores store one or more instructions, which are executed by one or more processing cores to enable the one or more processing cores to perform the route generation method according to any one of claims 1 to 11.
14. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the route generation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Configurable multi-core / many core system based on single instruction set microprocessor computing unit
CN101751373A
Hardware description language simulation acceleration method based on net list segmentation and multithreading paralleling
CN105589736A
Apparatus and method for routing data among multiple cores
US20110252179A1