Processing unit, synchronization method for processing unit and corresponding product
By introducing a synchronization instruction mechanism in a multi-core processor, data synchronization between cores in a multi-core processor architecture is achieved, which solves the problem of inter-core collaboration in a multi-core processor system and improves processing efficiency and task scheduling flexibility.
Patent Information
- Application Number
- CN202011036272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2040-09-27
AI Technical Summary
In a multi-core processor system, how to achieve collaborative work between multiple cores to meet the needs of high storage and computing capabilities.
By providing a synchronization method for a processing unit and a multi-core processor, data synchronization between multiple cores in a multi-core processor architecture is achieved, including a synchronization instruction interaction mechanism between the first core and the second core, ensuring the synchronization and consistency of data transmission.
It improves the flexible scheduling and processing efficiency of multi-core processor tasks and ensures the collaborative work between cores.
Smart Images

Figure CN114281560B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of processors, and in particular to a processing unit, a multi-core processor, a synchronization method for a processing unit / multi-core processor, a chip, and a board. Background Art
[0002] With the advancement of computer technology, various applications (such as video structuring, ad recommendations, and intelligent translation) are placing increasing demands on machine storage and computing power. Because single-core processors are no longer sufficient for these applications, various multi-core processor systems have emerged. A key issue in multi-core processor systems is the coordinated operation of multiple cores. Therefore, how to achieve inter-core coordination within a multi-core architecture is an urgent problem to be solved. Summary of the Invention
[0003] In order to solve one or more technical problems mentioned above, the present disclosure provides a processing unit, a multi-core processor and a synchronization method for a processing unit / multi-core processor in multiple aspects, wherein the provided synchronization method can achieve data synchronization between multiple cores in a multi-core processor architecture.
[0004] In a first aspect, the present disclosure provides a processing unit comprising at least one first core, wherein: the first core is configured to: in response to a first synchronization instruction associated with a second core, query whether there is an unprocessed synchronization event associated with the second core; and based on the result of the query, perform corresponding synchronization operations.
[0005] In a second aspect, the present disclosure provides a processing unit comprising at least one second core, wherein: the second core is configured to: in response to a second synchronization instruction associated with the first core, send a synchronization request signal to the first core to indicate a synchronization event to be processed; and in response to receiving a data transmission end signal from the first core, obtain the transmitted data to be synchronized.
[0006] In a third aspect, the present disclosure provides a multi-core processor comprising the processing unit of any embodiment of the first aspect and the processing unit of any embodiment of the second aspect.
[0007] In a fourth aspect, the present disclosure provides a chip, which encapsulates a processing unit as described in any embodiment of the first aspect, or a processing unit as described in any embodiment of the second aspect, or a multi-core processor as described in any embodiment of the third aspect.
[0008] In a fifth aspect, the present disclosure provides a board comprising the chip of any one of the embodiments of the fourth aspect.
[0009] In a sixth aspect, the present disclosure provides a synchronization method for a processing unit, wherein the processing unit includes at least one first core, the method comprising: the first core queries whether there is an unprocessed synchronization event associated with the second core in response to a first synchronization instruction associated with the second core; and the first core performs corresponding synchronization operations based on the result of the query.
[0010] In the seventh aspect, the present disclosure provides a synchronization method for a processing unit, wherein the processing unit includes at least one second core, and the method includes: the second core sends a synchronization request signal to the first core in response to a second synchronization instruction associated with the first core to indicate a synchronization event to be processed; and the second core obtains the transmitted data to be synchronized in response to receiving a data transmission end signal from the first core.
[0011] In the eighth aspect, the present disclosure provides a synchronization method for a multi-core processor, wherein the multi-core processor includes a first processing unit and a second processing unit, and the method includes: the first processing unit executes the synchronization method according to any embodiment of the aforementioned sixth aspect; and the second processing unit executes the synchronization method according to any embodiment of the aforementioned seventh aspect.
[0012] Through the processing unit, multi-core processor, synchronization method for processing unit / multi-core processor, chip and board provided above, the disclosed embodiment provides a solution for data synchronization between two cores in a multi-core processor architecture, thereby ensuring the collaborative work between the cores in the multi-core processor architecture, facilitating the flexible scheduling of multi-core processor tasks and improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts.
[0014] Figure 1 An exemplary structural diagram showing a multi-core processor architecture to which the embodiments of the present disclosure may be applied;
[0015] Figure 2 An exemplary internal architectural diagram of a processor core is shown;
[0016] Figures 3A-3D A schematic flow chart showing a synchronization method according to an embodiment of the present disclosure;
[0017] Figure 4A-4B A schematic diagram showing a data storage space according to an embodiment of the present disclosure;
[0018] Figure 5 A structural diagram showing a combined processing device according to an embodiment of the present disclosure; and
[0019] Figure 6 A schematic structural diagram of a board card according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.
[0021] It should be understood that the terms "first," "second," "third," and "fourth," etc., which may be used in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0022] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0023] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0024] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0025] Figure 1An exemplary structural diagram of a multi-core processor architecture to which embodiments of the present disclosure may be applied is shown. The multi-core processor 100 may be used to process input data such as computer vision, speech, natural language, and data mining. Figure 1 The multi-core processor 100 in the embodiment of the present invention adopts a multi-core hierarchical structure design. The multi-core processor 100 can be used as a system on a chip, which can include multiple clusters (Clusters), and each cluster includes multiple cores (Cores). In other words, the multi-core processor 100 is composed of a hierarchy of system on a chip, clusters, and cores.
[0026] At the system-on-chip level, Figure 1 As shown, the multi-core processor 100 includes an external memory controller 111 , a peripheral communication module 112 , an on-chip interconnect module 113 and a plurality of clusters 115 .
[0027] There can be multiple external storage controllers 111, and two are shown as examples in the figure. They are used to respond to access requests issued by the processor core and access external storage devices (such as DRAM) to read data from outside the chip or write data. The peripheral communication module 112 is used to receive control signals from the processing device (not shown) through the interface device (not shown) to start the multi-core processor 100 to perform tasks. The on-chip interconnect module 113 connects the external storage controller 111, the peripheral communication module 112 and the multiple clusters 115 to transmit data and control signals between the modules. The multiple clusters 115 are the computing cores of the multi-core processor 100. Four are shown as examples in the figure. With the development of hardware, the multi-core processor 100 disclosed herein may also include 8, 16, 64, or even more clusters 115. The clusters 115 are used to efficiently execute deep learning algorithms.
[0028] At the cluster level, Figure 1 As shown, each cluster 115 includes multiple processor cores (IPU cores) 121 and a memory core (MEM core) 122 .
[0029] The figure shows four processor cores 121 as an example, but the present disclosure does not limit the number of processor cores 121 .
[0030] The storage core 122 is primarily used for storage and communication, namely, to store shared data or intermediate results between the processor cores 121, and to perform communication between the execution cluster 115 and the DRAM 127, between the clusters 115, and between the processor cores 121. In other embodiments, the storage core 122 has scalar operation capabilities and is used to perform scalar operations.
[0031] The memory core 122 includes a shared memory unit (SMEM) 124, a broadcast bus 123, a cluster direct memory access (CDMA) module 126, and a global direct memory access (GDMA) module 125. SMEM 124 acts as a high-performance data transfer station. Data reused between different processor cores 121 within the same cluster 115 does not need to be obtained from each processor core 121 individually through DRAM 127. Instead, it is transferred between the processor cores 121 via SMEM 124. The memory core 122 only needs to quickly distribute the reused data from SMEM 124 to multiple processor cores 121, improving inter-core communication efficiency and significantly reducing on-chip and off-chip input / output accesses. The broadcast bus 123, CDMA 126, and GDMA 125 are used for communication between processor cores 121, communication between clusters 115, and data transfer between clusters 115 and DRAM 127, respectively. These functions are described below.
[0032] The broadcast bus 123 facilitates high-speed communication between the processor cores 121 within the cluster 115. In this embodiment, the broadcast bus 123 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (i.e., single-core to single-core) data transmission. Multicast transmits a copy of data from the SMEM 124 to a specific number of processor cores 121. Broadcast, a special case of multicast, transmits a copy of data from the SMEM 124 to all processor cores 121.
[0033] The CDMA 126 is used to control memory access to the SMEM 124 between different clusters 115 within the same multi-core processor 100 .
[0034] The GDMA 125 cooperates with the external memory controller 111 to control memory access from the SMEM 124 of the cluster 115 to the DRAM 127 , or to read data from the DRAM 127 to the SMEM 124 .
[0035] Although Figure 1 The multi-core processor architecture is described by taking a multi-core processor 100 of a single system on chip as an example. Those skilled in the art will appreciate that a multi-core processor can also be constructed using multiple single-core or multi-core processing units, and the present disclosure is not limited in this regard.
[0036] Figure 2 FIG. 1 shows an exemplary internal architecture diagram of the processor core 121. Figure 2 As shown, the processor core 121 may include three modules: a control module 21 , a calculation module 22 and a storage module 23 .
[0037] The control module 21 coordinates and controls the operations of the computing module 22 and the storage module 23 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 211 and an instruction decode unit (IDU) 212. The instruction fetch unit 211 retrieves instructions from a processing device (not shown), while the instruction decode unit 212 decodes the retrieved instructions and sends the decoded results as control information to the computing module 22 and the storage module 23.
[0038] The operation module 22 includes a vector operation unit 221 and a matrix operation unit 222. The vector operation unit 221 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformation. The matrix operation unit 222 is responsible for the core calculations of deep learning algorithms, such as matrix multiplication and convolution.
[0039] The storage module 23 is used to store or transfer relevant data, including a neuron RAM (NRAM) 231, a weight RAM (WRAM) 232, an input / output direct memory access module (IODMA) 233, and a move direct memory access module (MVDMA) 234. NRAM 231 is used to store input and output data and intermediate results for calculation by the processor core 121; WRAM 232 is used to store the weights of the deep learning network; IODMA 233 is used to transmit data to the processor core 121 through the broadcast bus 123 (see Figure 1 ) controls memory access between NRAM 231 / WRAM 232 and DRAM 127; MVDMA 234 controls memory access between NRAM 231 / WRAM 232 and SMEM 124.
[0040] In some embodiments, GDMA 125 (see Figure 1 ) functions and IODMA 233 (see Figure 2) can be integrated into the same component. For ease of description, this disclosure treats GDMA 125 and IODMA 233 as different components. Those skilled in the art will recognize that as long as the functions implemented and the technical effects achieved are similar to those disclosed herein, they fall within the scope of protection of this disclosure. Furthermore, the functions of GDMA 125, IODMA 233, CDMA 126, and MVDMA 234 can also be implemented by the same component. Similarly, as long as the functions implemented and the technical effects achieved are similar to those disclosed herein, they fall within the scope of protection of this disclosure.
[0041] In a multi-core processor architecture, multiple processes can run simultaneously on multiple cores, and some processes may have certain connections. Multiple processes may need to synchronize data in order to complete the same task. Based on the aforementioned multi-core processor architecture, data synchronization can include synchronization between storage cores in different clusters, as well as synchronization between storage cores and processor cores within the same cluster. The disclosed embodiments provide a synchronization mechanism for a multi-core processor architecture, which can be applied to any of the aforementioned synchronization scenarios.
[0042] Figure 3A An exemplary flow chart showing a method for synchronous interaction between two cores in a multi-core processor architecture according to an embodiment of the present disclosure. The multi-core processor architecture may be, for example, the one previously described. Figure 1-Figure 2 An exemplary multi-core processor architecture is described. Figure 3A As shown, the synchronous interaction method involves at least two cores: a first core 310 and a second core 320, with data synchronization performed between the first core 310 and the second core 320. The first core 310 and the second core 320 can be provided by different processing units or the same processing unit, which is not a limitation of this disclosure. For ease of description, it is assumed that the first core 310 is the data generator, which generates the data to be synchronized; the second core 320 is the data acquirer, which acquires the data generated by the second core to achieve data synchronization.
[0043] Those skilled in the art will appreciate that for different synchronization scenarios, the first core and the second core may represent different cores. For example, for inter-cluster synchronization of a multi-core processor, the first core and the second core may be storage cores of different clusters in the multi-core processor (see Figure 1 For another example, for intra-cluster synchronization of a multi-core processor, the first core and the second core may be processor cores or storage cores, respectively, within the same cluster of the multi-core processor. Specifically, the first core may be a processor core and the second core may be a storage core; or vice versa, the first core may be a storage core and the second core may be a processor core.
[0044] In multi-core processing mode, each core maintains its own resource data, which is not visible to other cores, nor can it see the data of other cores. Therefore, before the data to be synchronized is transmitted, synchronization is required to ensure that both ends of the synchronization are ready before data transmission can proceed. The data generator needs to prepare the relevant data; the data acquisition end needs to prepare the corresponding storage space to store the relevant data. Therefore, in the embodiments of the present disclosure, two synchronization instructions are provided: a first synchronization instruction and a second synchronization instruction.
[0045] The first synchronization instruction is used for a core in a multi-core processor that is a data generator, such as the first core 310 in this embodiment. The first synchronization instruction can be a Send instruction, which is used to indicate that the sender of the data to be synchronized for the relevant synchronization event is ready. In other words, the data to be synchronized has been generated at this time.
[0046] The first synchronization instruction may include information related to the data to be synchronized involved in the synchronization event, information related to the synchronization object, and so on. The synchronization object may be, for example, a core in a multi-core processor that serves as a data acquisition end, such as the second core 320 in this embodiment. The information related to the synchronization object may include, for example, the identification of the core. In some embodiments, the information related to the synchronization object may also include the data receiving address of the synchronization object, that is, the predetermined address in the synchronization object for initially storing the data transmitted by the data generation end. The address may be a fixed address or a variable address, and the present disclosure is not limited in this respect. The address may also be pre-set, so that it does not need to be included in the first synchronization instruction.
[0047] The second synchronization instruction is used by a core in a multi-core processor that serves as a data acquisition end, such as the second core 320 in this embodiment. The second synchronization instruction can be a receive instruction, which indicates that the receiving end of the data to be synchronized for the relevant synchronization event is ready. In other words, the storage space for storing the data to be synchronized has been prepared. This storage space corresponds to the data receiving address or predetermined address of the synchronization object.
[0048] The second synchronization instruction may include information related to the data to be synchronized involved in the synchronization event, information related to the synchronization object, and so on. The synchronization object may be, for example, a core in a multi-core processor that serves as a data generation end, such as the first core 310 in this embodiment. The information related to the synchronization object may include, for example, the identification of the core. The information related to the data to be synchronized can be used to determine the destination address for ultimately storing the data to be synchronized at the data acquisition end. In some embodiments, the data to be synchronized is tensor data, and the second synchronization instruction includes a descriptor of the tensor data. The specific content of the descriptor of the tensor data will be described later.
[0049] See also Figure 3A, which shows the interaction process for data synchronization between the first core 310 and the second core 320. The first core 310 and the second core 320 can each execute corresponding instructions, which may include synchronization instructions, such as a first synchronization instruction or a second synchronization instruction, for indicating a data synchronization event between the cores.
[0050] In step S321, the second core may encounter a second synchronization instruction during instruction execution. As previously described, the second synchronization instruction indicates the presence of a data synchronization event and that the second core, acting as the data acquirer or receiver, is ready for the synchronization event. In other words, the second core has prepared storage space to receive the data. The second synchronization instruction may include information related to the data to be synchronized related to the synchronization event, information related to the synchronization object, and so on. In this example, the synchronization object is the first core 310.
[0051] Next, in step S322, the second core responds to the second synchronization instruction and sends a synchronization request signal to the first core as the synchronization target to indicate the pending synchronization event. The synchronization request signal does not need to carry the data information to be synchronized, but only needs to indicate the state of being ready to receive to the first core.
[0052] Then, the second core may enter a waiting state, waiting for a data transfer completion signal from the first core.
[0053] The first core and the second core may execute their respective instructions in parallel, and the instruction execution progress may be different. Figure 3A 3 shows that the first core 310 first receives the synchronization request signal sent by the second core 320, and then executes the first synchronization instruction associated with the second core.
[0054] like Figure 3A As shown, in step S311, in response to receiving a synchronization request signal from the second core, a related synchronization event is recorded, which indicates that there is a data synchronization event to be processed associated with the second core.
[0055] In step S312, the first core may encounter a first synchronization instruction during instruction execution. As previously described, a first synchronization instruction indicates the presence of a data synchronization event and that the first core, as the data generator or sender, is ready for the synchronization event. In other words, the first core has generated or prepared the data to be synchronized and can send it to the second core. The first synchronization instruction may include information related to the data to be synchronized related to the synchronization event, information related to the synchronization object, and so on. In this example, the synchronization object is the second core 320.
[0056] Next, in step S313, the first core queries whether there are any unprocessed synchronization events associated with the second core in response to the first synchronization instruction, and then performs corresponding synchronization operations based on the query result.
[0057] exist Figure 3A In the illustrated embodiment, a synchronization request signal from the second core has been received before the first synchronization instruction is executed. Therefore, in this embodiment, the query result at step S313 indicates that there is an unprocessed synchronization event. The process then proceeds to step S314.
[0058] At step S314, the data to be synchronized is transmitted to the second core 320. The data transmission address can be fixed or variable. For example, the address can be indicated in the first synchronization instruction as described above, or it can be pre-set and known to the second core, thus eliminating the need for additional indication in the first synchronization instruction.
[0059] When the data transfer is completed, at step S315 , a data transfer completion signal may be sent to the second core to indicate the completion of the data transfer to the second core.
[0060] Then, at step S322 , the second core obtains the transferred data in response to receiving the data transfer end signal from the first core.
[0061] In some embodiments, the second core obtaining the transferred data may include: determining a destination address of the data to be synchronized based on the second synchronization instruction; and reading the data transferred by the first core from the aforementioned transfer address of the data (e.g., a pre-known fixed address) and writing the data to the determined destination address. The second core may cache the data when executing the second synchronization instruction for use when determining the destination address at this time.
[0062] Furthermore, the synchronization method may further include step S323, where the second core, in response to receiving the data transfer completion signal from the first core, sends an acknowledgement signal to the first core to indicate the completion of synchronization. Subsequently, upon receiving the acknowledgement signal, the second core may terminate the synchronization operation. Receipt of the acknowledgement signal may be considered as completion of data synchronization between the two cores, eliminating the need to wait for the second core to write all data to the destination address, thereby accelerating instruction execution.
[0063] Alternatively or additionally, in some embodiments, when dependencies between processed data are maintained via a dependency data structure, the dependency data structure may be updated accordingly after synchronization is complete. In some implementations, the data structure may be represented as a data table, a data read / write counter, etc., and the present disclosure is not limited in this respect.
[0064] Alternatively or additionally, in some embodiments, after acquiring the data, the second core may compare the total amount of data received with the total amount of data expected to be received, and report an exception if there is a mismatch, or discard the excess data if there is excess data.
[0065] The physical path carrying the data transmission between the first core and the second core is sequence-preserving transmission, thereby ensuring data correctness. Various technologies in the field of communication transmission can be used to ensure sequence-preserving transmission of the physical path, and this disclosure is not limited in this regard.
[0066] The various signal transmissions between the first core and the second core described above can be implemented in various ways. These signal transmissions include, for example, synchronization request signals, data transfer end signals, response signals, etc. In some embodiments, the transmission of these signals can be implemented in the form of semaphores.
[0067] Figure 3B An exemplary flow chart of a synchronization method for a multi-core processor according to another embodiment of the present disclosure is shown. As mentioned above, the first core and the second core may execute their respective instructions in parallel, and the instruction execution progress may be different from each other. Figure 3B and Figure 3A Basically the same, except: Figure 3B 3 shows that the first core 310 first executes the first synchronization instruction associated with the second core, and then receives the synchronization request signal sent by the second core 320, that is, the first core 310 executes the synchronization instruction before the second core 320.
[0068] like Figure 3B As shown, in step S312, the first core executes a first synchronization instruction during the instruction execution process. The first synchronization instruction indicates that the first core has generated or prepared the data to be synchronized and can send it to the second core. In this example, the synchronization target is the second core 320.
[0069] Next, in step S313, the first core queries whether there are any unprocessed synchronization events associated with the second core in response to the first synchronization instruction, and then performs corresponding synchronization operations based on the query result.
[0070] exist Figure 3B The illustrated embodiment shows a situation where the synchronization request signal from the second core has not yet been received before the first synchronization instruction is executed. Therefore, in this embodiment, the query result at step S313 indicates that there are no unprocessed synchronization events. Subsequently, the process can wait for the synchronization request signal from the second core.
[0071] While the first core is waiting, the second core may execute the second synchronization instruction. Figure 3BAs shown, in step S321, the second core may encounter a second synchronization instruction during instruction execution. The second synchronization instruction indicates that the second core has prepared storage space to receive data.
[0072] Next, in step S322, the second core responds to the second synchronization instruction and sends a synchronization request signal to the first core as the synchronization target to indicate the pending synchronization event. The synchronization request signal does not need to carry the data information to be synchronized, but only needs to indicate the state of being ready to receive to the first core.
[0073] Then, the second core may enter a waiting state, waiting for a data transfer completion signal from the first core.
[0074] At this time, the first core records a related synchronization event in response to receiving a synchronization request signal from the second core during the waiting process, as shown in step S311. The record indicates that there is a data synchronization event to be processed associated with the second core.
[0075] Then, the first core can determine in the query of step S313 that there is an unprocessed synchronization event based on the record in step S311. Therefore, the process proceeds to step S314 to transmit the synchronization data. The subsequent data transmission process is the same as Figure 3A The same as in , no further description is given here.
[0076] For ease of understanding, Figure 3A-3B The synchronization mechanism provided by the embodiment of the present disclosure is illustrated in an interactive manner. However, those skilled in the art will appreciate that the embodiment of the present disclosure also provides a synchronization method for each core in the synchronization process.
[0077] Figure 3C A synchronization method for a processing unit according to an embodiment of the present disclosure is shown, where the processing unit includes at least one first core.
[0078] like Figure 3C As shown, method 300C includes step S331, in which the first core, in response to a first synchronization instruction associated with the second core, queries whether there is an unprocessed synchronization event associated with the second core. In some implementations, the first synchronization instruction can be a Send instruction, which indicates that the sender of the to-be-synchronized data of the related synchronization event is ready.
[0079] Next, in step S332 , the first core performs a corresponding synchronization operation based on the query result.
[0080] Specifically, in one embodiment, based on the determination at step S3321, method 300C may further include, when the query detects the presence of an unprocessed synchronization event ("Yes" branch at S3321), transmitting the data to be synchronized to the second core at step S3322. In some implementations, the first core may transmit the data to be synchronized to a predetermined address of the second core. The predetermined address may be, for example, specified in the first synchronization instruction or pre-set.
[0081] Next, it can be determined whether the data transmission has been completed (step S3323); and when the data transmission is completed, a data transmission completion signal is sent to the second core (step 3324).
[0082] In some implementations, the data transfer end signal may be sent in the form of a semaphore.
[0083] In another embodiment, based on the judgment at step S3321, method 300C may further include, when no unprocessed synchronization event is queried (branch "No" at S3321), in step S3325, waiting for a synchronization event, that is, waiting for a synchronization request signal from the second core.
[0084] During the waiting period, when a synchronization request signal is received from the second core (“Yes” in S3326 ), the synchronization event is recorded (step S3327 ). The method flow then returns to step S3321 to determine whether there is any unprocessed synchronization event.
[0085] Alternatively or additionally, in some embodiments, the method 300C may further include step S333 , in which the first core ends the synchronization operation in response to receiving a response signal from the second core indicating that the synchronization is complete.
[0086] Figure 3D A synchronization method for a processing unit according to an embodiment of the present disclosure is shown, where the processing unit includes at least one second core.
[0087] like Figure 3D As shown, method 300D includes step S341, in response to a second synchronization instruction associated with the first core, sending a synchronization request signal to the first core to indicate a pending synchronization event; and step S342, in response to receiving a data transfer completion signal from the first core, obtaining the transmitted data to be synchronized. In some implementations, the second synchronization instruction is a receive instruction, which indicates that the receiving end of the data to be synchronized for the associated synchronization event is ready. In other implementations, the synchronization request signal is sent as a semaphore.
[0088] In one embodiment, obtaining the transmitted data to be synchronized can be specifically implemented as multiple sub-steps. For example, in step S3421, a destination address of the data to be synchronized is determined based on the second synchronization instruction. In step S3422, the transmitted data to be synchronized is read from a predetermined address for receiving data to be synchronized transmitted from the first core. Finally, in step S3423, the read data to be synchronized is written to the destination address determined in step S3421.
[0089] In some implementations, the data to be synchronized is tensor data, and the second synchronization instruction includes a descriptor for the tensor data. Therefore, in step S3421, the destination address of the tensor data can be determined based on the descriptor in the second synchronization instruction. The descriptor can include at least shape information of the tensor data.
[0090] Alternatively or additionally, in some embodiments, the method 300D may further include step S343 , wherein the second core sends a response signal to the first core to indicate synchronization completion in response to receiving the data transfer end signal from the first core.
[0091] Figure 3C and Figure 3D The specific implementation of each step can refer to the previous combination Figure 3A and Figure 3B The description is not repeated here.
[0092] from Figures 3A-3D As can be seen from the method flow chart, the disclosed embodiment provides a synchronization mechanism of send (Send)-receive (Receive) mode, in which data is actively sent from one core (source end) to another core (destination end), and the two cores do not need to transmit data information involved in synchronization within their respective cores. Moreover, during the data transmission process, instructions exist in both the first core (source end) and the second core (destination end), and each maintains sequential consistency within the core, so that data synchronization can be achieved between the two cores without affecting the execution order of their respective instructions. In some embodiments, when the data receiving channel of the destination end receives the data, the source end can consider that the transmission between the two is completed without waiting until the destination end actually writes the data to the destination address, thereby speeding up instruction processing and improving the efficiency of the processor.
[0093] The synchronization mechanism applied to a multi-core processor according to an embodiment of the present disclosure has been described above with reference to the flowchart. It should be noted that, for the sake of simplicity of description, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited to the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present disclosure.
[0094] It should be further explained that although Figure 3A and Figure 3B The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 3A and Figure 3B At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0095] With the development of artificial intelligence technology, tasks such as image processing and pattern recognition often require operands of multidimensional vector data types (i.e., tensor data). These tensor data are typically processed on the multi-core processors of the disclosed embodiments. Therefore, how to synchronize tensor data between multiple cores of a multi-core processor is an urgent problem in the current computing field.
[0096] In embodiments of the present disclosure, when the data to be synchronized is tensor data, a descriptor may be included in the operand of the synchronization instruction, through which information related to the tensor data can be quickly obtained. For example, the second synchronization instruction may include a descriptor of the tensor data, and the second core (destination end) may then determine the destination address of the tensor data based on the descriptor.
[0097] Specifically, the descriptor may at least indicate shape information of the tensor data. The shape information of the tensor data may be used to determine a data address of the tensor data corresponding to the operand in the data storage space.
[0098] Various possible implementations of the shape information of tensor data will be described in detail below with reference to the accompanying drawings.
[0099] Tensors can contain a variety of data structures. Tensors can be of different dimensions. For example, a scalar can be considered a 0-dimensional tensor, a vector can be considered a 1-dimensional tensor, and a matrix can be a 2-dimensional or higher tensor. The shape of a tensor includes information such as the dimensions of the tensor and the size of each dimension of the tensor. For example, for a 3D tensor:
[0100] x3=[[[1,2,3],[4,5,6]];[[7,8,9],[10,11,12]]]
[0101] The shape or dimensions of this tensor can be expressed as X3 = (2, 2, 3), which means that the three parameters indicate that the tensor is a three-dimensional tensor with a first dimension of 2, a second dimension of 2, and a third dimension of 3. When storing tensor data in memory, the shape of the tensor data cannot be determined based on its data address (or storage area), and thus, related information such as the relationship between multiple tensor data cannot be determined, resulting in low processor access efficiency for tensor data.
[0102] In one possible implementation, a descriptor can be used to indicate the shape of N-dimensional tensor data, where N is a positive integer, such as N=1, 2, or 3, or zero. The three-dimensional tensor in the above example can be represented by a descriptor as (2, 2, 3). It should be noted that this disclosure does not limit the manner in which a descriptor indicates the shape of a tensor.
[0103] In one possible implementation, the value of N can be determined according to the dimension (also called order) of the tensor data, or it can be set according to the usage requirements of the tensor data. For example, when the value of N is 3, the tensor data is three-dimensional tensor data, and the descriptor can be used to indicate the shape of the three-dimensional tensor data in three dimensions (such as offset, size, etc.). It should be understood that those skilled in the art can set the value of N according to actual needs, and this disclosure does not limit this.
[0104] Although tensor data can be multi-dimensional, because the memory layout is always one-dimensional, there is a correspondence between tensors and storage on the memory. Tensor data is usually allocated in a continuous memory space, that is, tensor data can be expanded one-dimensionally (for example, row-major) and stored in the memory.
[0105] This relationship between a tensor and its underlying storage can be expressed through dimensions such as offset, size, and stride. A dimension's offset refers to the offset relative to a reference position in that dimension. A dimension's size refers to the size of that dimension, or the number of elements in that dimension. A dimension's stride refers to the spacing between adjacent elements in that dimension. For example, the stride of the three-dimensional tensor above is (6, 3, 1), meaning the stride of the first dimension is 6, the stride of the second dimension is 3, and the stride of the third dimension is 1.
[0106] Figure 4A Schematic diagram showing the data storage space according to the embodiment of the present disclosure. Figure 4A As shown, data storage space 41 stores two-dimensional data in a row-major manner, which can be represented by (x, y) (where the X axis is horizontal and rightward, and the Y axis is vertical and downward). The size in the X-axis direction (the size of each row, or the total number of columns) is ori_x (not shown in the figure), and the size in the Y-axis direction (the total number of rows) is ori_y (not shown in the figure). The starting address PA_start (base address) of data storage space 41 is the physical address of the first data block 42. Data block 43 is a portion of the data in data storage space 41. Its offset 45 in the X-axis direction is represented by offset_x, its offset 44 in the Y-axis direction is represented by offset_y, its size in the X-axis direction is represented by size_x, and its size in the Y-axis direction is represented by size_y.
[0107] In one possible implementation, when a descriptor is used to define data block 43, the data reference point of the descriptor can be the first data block of data storage space 41. The descriptor's reference address can be agreed to be the starting address PA_start of data storage space 41. The content of the descriptor for data block 43 can then be determined by combining the X-axis size ori_x and Y-axis size ori_y of data storage space 41, as well as the Y-axis offset offset_y, X-axis offset offset_x, X-axis size size_x, and Y-axis size size_y of data block 43.
[0108] In a possible implementation, the following formula (1) can be used to express the content of the descriptor:
[0109]
[0110] It should be understood that although in the above examples, the content of the descriptor represents a two-dimensional space, those skilled in the art can set the specific dimension represented by the content of the descriptor according to actual conditions, and this disclosure does not limit this.
[0111] In one possible implementation, the base address of the data reference point of the descriptor in the data storage space can be agreed upon. Based on the base address, the content of the descriptor of the tensor data is determined according to the positions of at least two vertices at diagonal positions in N dimensional directions relative to the data reference point.
[0112] For example, the data reference point of the descriptor can be agreed to be the reference address PA_base in the data storage space. For example, a data (e.g., data at position (2, 2)) can be selected in the data storage space 41 as the data reference point, and the physical address of the data in the data storage space can be used as the reference address PA_base. The position of the two vertices at the diagonal position relative to the data reference point can be used to determine the reference address PA_base. Figure 4A The content of the descriptor of data block 43 in the data block 43 is determined. First, the positions of at least two vertices at the diagonal positions of the data block 43 relative to the data reference point are determined. For example, the positions of the diagonal vertices from the upper left to the lower right relative to the data reference point are used, where the relative position of the upper left vertex is (x_min, y_min) and the relative position of the lower right vertex is (x_max, y_max). Then, the content of the descriptor of data block 43 can be determined based on the reference address PA_base, the relative position of the upper left vertex (x_min, y_min), and the relative position of the lower right vertex (x_max, y_max).
[0113] In a possible implementation, the following formula (2) can be used to express the content of the descriptor (the base address is PA_base):
[0114]
[0115] It should be understood that although the vertices at the upper left corner and the lower right corner are used in the above example to determine the content of the descriptor, those skilled in the art can set the specific vertices of at least two vertices at the diagonal positions according to actual needs, and this disclosure does not limit this.
[0116] In one possible implementation, the content of the tensor data descriptor can be determined based on the reference address of the descriptor's data reference point in the data storage space and the mapping relationship between the data description position and the data address of the tensor data indicated by the descriptor. The mapping relationship between the data description position and the data address can be set according to actual needs. For example, when the tensor data indicated by the descriptor is three-dimensional spatial data, the function f(x, y, z) can be used to define the mapping relationship between the data description position and the data address.
[0117] In a possible implementation, the following formula (3) can be used to express the content of the descriptor:
[0118]
[0119] In one possible implementation, the descriptor is further used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor further includes at least one address parameter representing the address of the tensor data. For example, the content of the descriptor may be the following formula (4):
[0120]
[0121] Where PA is the address parameter. The address parameter can be a logical address or a physical address. When parsing the descriptor, PA can be used as any vertex, midpoint, or preset point of the vector shape, combined with the shape parameters in the X and Y directions to obtain the corresponding data address.
[0122] In a possible implementation, the address parameter of the tensor data includes a reference address of a data reference point of the descriptor in the data storage space of the tensor data, and the reference address includes a starting address of the data storage space.
[0123] In a possible implementation, the descriptor may further include at least one address parameter representing the address of the tensor data. For example, the content of the descriptor may be the following formula (5):
[0124]
[0125] PA_start is the base address parameter and will not be described in detail.
[0126] It should be understood that those skilled in the art can set the mapping relationship between the data description location and the data address according to actual conditions, and this disclosure does not limit this.
[0127] In one possible implementation, a predetermined reference address can be set within a task. All descriptors in instructions within this task use this reference address, and the descriptor content can include shape parameters based on this reference address. This reference address can be determined by setting the environment parameters for this task. A description of the reference address and its use can be found in the above embodiments. In this implementation, the descriptor content can be mapped to data addresses more quickly.
[0128] In one possible implementation, the base address can be included in the content of each descriptor, so that the base address of each descriptor can be different. Compared with the method of using environmental parameters to set a common base address, each descriptor in this method can describe data more flexibly and use a larger data address space.
[0129] In one possible implementation, the data address of the data corresponding to the operand of the processing instruction in the data storage space can be determined based on the content of the descriptor. The data address is calculated automatically by hardware, and the calculation method of the data address varies depending on the representation of the descriptor content. This disclosure does not limit the specific method for calculating the data address.
[0130] For example, the content of the descriptor in the operand is expressed using formula (1). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y, and the size is size_x*size_y. Then, the starting data address PA1 of the tensor data indicated by the descriptor in the data storage space is (x,y) It can be determined using the following formula (6):
[0131] PA1 (x,y) =PA_start+(offset_y-1)*ori_x+offset_x (6)
[0132] The data starting address PA1 is determined according to the above formula (6) (x,y) , combined with the offsets offset_x and offset_y, and the sizes size_x and size_y of the storage area, the storage area of the tensor data indicated by the descriptor in the data storage space can be determined.
[0133] In one possible implementation, when the operand also includes a data description location for a descriptor, the data address of the data corresponding to the operand in the data storage space can be determined based on the content of the descriptor and the data description location. In this way, partial data (e.g., one or more data) in the tensor data indicated by the descriptor can be processed.
[0134] For example, the content of the descriptor in the operand is expressed using formula (2). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y respectively, and the size is size_x*size_y. The data description position for the descriptor included in the operand is (x q ,y q ), then the data address PA2 of the tensor data indicated by the descriptor in the data storage space (x,y) It can be determined using the following formula (7):
[0135] PA2 (x,y) =PA_start+(offset_y+y q -1)*ori_x+(offset_x+x q ) (7)
[0136] In one possible implementation, the descriptor can indicate data blocks. Data blocks can effectively speed up operations and improve processing efficiency in many applications. For example, in graphics processing, convolution operations often use data blocks for fast processing.
[0137] Figure 4B Schematic diagram showing data blocks in data storage space according to an embodiment of the present disclosure. Figure 4B As shown, the data storage space 46 also uses a row-first approach to store two-dimensional data, which can be represented by (x, y) (where the X axis is horizontal to the right and the Y axis is vertically downward). The size in the X-axis direction (the size of each row, or the total number of columns) is ori_x (not shown in the figure), and the size in the Y-axis direction (the total number of rows) is ori_y (not shown in the figure). Figure 4A Tensor data, Figure 4B The tensor data stored in consists of multiple data blocks.
[0138] In this case, the descriptor needs more parameters to represent these data blocks. Taking the X-axis (X dimension) as an example, the following parameters may be involved: ori_x, x.tile.size (the size of the block is 47), x.tile.stride (the step size in the block is 48, that is, the distance between the first point of the first small block and the first point of the second small block), x.tile.num (the number of blocks, Figure 4B ), x.stride (the overall step size, that is, the distance from the first point in the first row to the first point in the second row), etc. Other dimensions can similarly include corresponding parameters.
[0139] In one possible implementation, a descriptor may include a descriptor identifier and / or descriptor content. The descriptor identifier is used to distinguish the descriptor, for example, the descriptor identifier may be a number; the descriptor content may include at least one shape parameter representing the shape of the tensor data. For example, if the tensor data is three-dimensional data, and the shape parameters of two of the three dimensions of the tensor data are fixed, the descriptor content may include the shape parameter representing the other dimension of the tensor data.
[0140] In one possible implementation, the data address of the data storage space corresponding to each descriptor can be a fixed address. For example, a separate data storage space can be divided for tensor data, and the starting address of each tensor data in the data storage space corresponds one-to-one to the descriptor. In this case, the circuit or module responsible for parsing the computing instruction (such as an entity outside the computing device of the present disclosure) can determine the data address of the data corresponding to the operand in the data storage space based on the descriptor.
[0141] In one possible implementation, when the data address of the data storage space corresponding to the descriptor is a variable address, the descriptor can also be used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor can also include at least one address parameter representing the address of the tensor data. For example, the tensor data is 3-dimensional data. When the descriptor points to the address of the tensor data, the content of the descriptor may include an address parameter representing the address of the tensor data, such as the starting physical address of the tensor data, or may include multiple address parameters of the address of the tensor data, such as the starting address + address offset of the tensor data, or the address parameters of the tensor data based on each dimension. Those skilled in the art can set the address parameters according to actual needs, and this disclosure does not limit this.
[0142] In one possible implementation, the address parameter of the tensor data may include the reference address of the descriptor's data reference point in the data storage space of the tensor data. The reference address may vary depending on the data reference point. This disclosure does not limit the selection of the data reference point.
[0143] In one possible implementation, the reference address may include the starting address of the data storage space. When the data reference point of the descriptor is the first data block in the data storage space, the reference address of the descriptor is the starting address of the data storage space. When the data reference point of the descriptor is data other than the first data block in the data storage space, the reference address of the descriptor is the address of the data block in the data storage space.
[0144] In one possible implementation, the shape parameters of the tensor data include at least one of the following: the size of the data storage space in at least one direction of the N-dimensional directions, the size of the storage area in at least one direction of the N-dimensional directions, the offset of the storage area in at least one direction of the N-dimensional directions, the positions of at least two vertices at diagonal positions in the N-dimensional directions relative to the data reference point, and the mapping relationship between the data description position of the tensor data indicated by the descriptor and the data address. The data description position is the mapping position of the point or area in the tensor data indicated by the descriptor. For example, when the tensor data is 3D data, the descriptor can use a three-dimensional coordinate space (x, y, z) to represent the shape of the tensor data. The data description position of the tensor data can be the position of the point or area mapped in the three-dimensional space, represented by the three-dimensional space coordinates (x, y, z).
[0145] It should be understood that those skilled in the art can select shape parameters representing tensor data according to actual circumstances, and this disclosure does not limit this. By using descriptors in the data access process, associations between data can be established, thereby reducing the complexity of data access and improving instruction processing efficiency.
[0146] Figure 5 FIG. 5 is a structural diagram showing a combined processing device 500 according to an embodiment of the present disclosure. Figure 5 As shown in FIG, the combined processing device 500 includes a computing device 502, an interface device 504, other processing devices 506, and a storage device 508. According to different application scenarios, the computing device may include one or more computing devices 510, which may be configured as Figure 1 The multi-core processor shown is used to implement the present invention in combination with Figure 3A-3B The described operation.
[0147] In various embodiments, the computing and processing device of the present disclosure may be configured to perform user-specified operations. In exemplary applications, the computing and processing device may be implemented as a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing and processing device may be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, the computing and processing device of the present disclosure may be considered to have a homogeneous multi-core structure.
[0148] In exemplary operation, the computing processing device of the present disclosure can interact with other processing devices through interface means, to jointly complete the operation specified by the user. Depending on the difference in implementation, the other processing devices of the present disclosure may include one or more types of processors in general and / or special processors such as central processing unit (Central Processing Unit, CPU), graphics processing unit (Graphics Processing Unit, GPU), artificial intelligence processor. These processors may include but are not limited to digital signal processor (Digital Signal Processor, DSP), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As previously mentioned, only with respect to the computing processing device of the present disclosure, it can be regarded as having a homogeneous multi-core structure. However, when the computing processing device and other processing devices are considered together, the two can be regarded as forming a heterogeneous multi-core structure.
[0149] In one or more embodiments, the other processing device may serve as an interface between the computing device disclosed herein (which may be embodied as an artificial intelligence computing device such as a neural network computing device) and external data and control, performing basic control including but not limited to data transfer, starting and / or stopping the computing device, and so on. In other embodiments, the other processing device may also collaborate with the computing device to jointly complete computing tasks.
[0150] In one or more embodiments, the interface device can be used to transmit data and control instructions between the computing and processing device and other processing devices. For example, the computing and processing device can obtain input data from other processing devices via the interface device and write it to the storage device (or memory) on the computing and processing device chip. Furthermore, the computing and processing device can obtain control instructions from other processing devices via the interface device and write them to the control cache on the computing and processing device chip. Alternatively or optionally, the interface device can also read data from the storage device of the computing and processing device and transmit it to other processing devices.
[0151] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is connected to the computing processing device and the other processing device, respectively. In one or more embodiments, the storage device may be used to store data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing device.
[0152] In some embodiments, the present disclosure also discloses a chip (e.g. Figure 6 In one implementation, the chip is a system on chip (SoC) and integrates one or more components such as Figure 5 The chip can be connected to the external interface device (such as Figure 6 The external interface device 606 shown in the figure is connected to other related components. The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card or a wifi interface. In some application scenarios, other processing units (such as video codecs) and / or interface modules (such as DRAM interfaces) can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip packaging structure, which includes the above-mentioned chip. In some embodiments, the present disclosure also discloses a board card, which includes the above-mentioned chip packaging structure. The following will be combined with Figure 6 The board is described in detail.
[0153] Figure 6 FIG. 1 is a schematic diagram showing the structure of a board 600 according to an embodiment of the present disclosure. Figure 6 As shown in , the board includes a storage device 604 for storing data, which includes one or more storage units 610. The storage device can be connected to the control device 608 and the chip 602 described above and transmit data by means of, for example, a bus. Further, the board also includes an external interface device 606, which is configured for data relay or transfer function between the chip (or the chip in the chip packaging structure) and the external device 612 (such as a server or computer, etc.). For example, the data to be processed can be passed to the chip by the external device through the external interface device. For another example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms, for example, it can adopt a standard PCIE interface, etc.
[0154] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. To this end, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the working state of the chip.
[0155] According to the above combination Figure 5 and Figure 6 Based on the description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which may include one or more of the above-mentioned boards, one or more of the above-mentioned chips and / or one or more of the above-mentioned combined processing devices.
[0156] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.
[0157] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.
[0158] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this document divides them based on the consideration of logical functions, and there may be other ways of division in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.
[0159] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.
[0160] In some implementation scenarios, the above-mentioned integrated unit can be implemented in the form of a software program module. If implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the scheme of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server or a network device, etc.) to perform some or all of the steps of the method described in the embodiment of the present disclosure. The aforementioned memory may include, but is not limited to, various media that can store program code, such as a USB flash drive, a flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0161] In some other implementation scenarios, the above-mentioned integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.
[0162] The foregoing content can be better understood in accordance with the following terms:
[0163] Clause 1. A processing unit comprising at least one first core, wherein:
[0164] The first core is configured to:
[0165] In response to a first synchronization instruction associated with a second core, querying whether there is an unprocessed synchronization event associated with the second core; and
[0166] Based on the result of the query, a corresponding synchronization operation is performed.
[0167] Clause 2. The processing unit of clause 1, wherein the first core is further configured to:
[0168] When it is found that there is the unprocessed synchronization event, the data to be synchronized is transmitted to the second core; and
[0169] When the data transfer is completed, the data transfer completion signal is sent to the second core.
[0170] Clause 3. The processing unit of clause 2, wherein the first core is configured to send the data transfer end signal in a semaphore manner.
[0171] Clause 4. The processing unit of any one of clauses 1-3, wherein the first core is further configured to:
[0172] When no unprocessed synchronization event is found, waiting for a synchronization request signal from the second core; and
[0173] In response to receiving the synchronization request signal from the second core, the synchronization event is recorded.
[0174] Clause 5. The processing unit according to any one of clauses 1 to 4, wherein:
[0175] The first core is further configured to transmit data requiring synchronization to a predetermined address of the second core.
[0176] Clause 6. The processing unit according to any one of clauses 1 to 5, wherein:
[0177] The first core is further configured to: end the synchronization operation in response to receiving a response signal from the second core indicating that synchronization is complete.
[0178] Clause 7. A processing unit according to any one of clauses 1 to 6, wherein:
[0179] The first core and the second core are storage cores of different clusters in a multi-core processor architecture; or
[0180] The first core and the second core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
[0181] Clause 8. The processing unit according to any one of clauses 1 to 7, wherein:
[0182] The first synchronization instruction is a send instruction, which indicates that the sender of the to-be-synchronized data of the related synchronization event is ready.
[0183] Clause 9. A processing unit comprising at least one second core, wherein:
[0184] The second core is configured to:
[0185] In response to a second synchronization instruction associated with the first core, sending a synchronization request signal to the first core to indicate a pending synchronization event; and
[0186] In response to receiving a data transmission completion signal from the first core, the transmitted data to be synchronized is acquired.
[0187] Clause 10. The processing unit of clause 9, wherein the second core is further configured to:
[0188] In response to the data transfer end signal, determining a destination address of the data to be synchronized based on the second synchronization instruction;
[0189] reading the transmitted data to be synchronized from a predetermined address for receiving the data to be synchronized transmitted from the first core; and
[0190] The read data to be synchronized is written into the destination address.
[0191] Clause 11. The processing unit according to any one of clauses 9-10, wherein the data to be synchronized is tensor data, the second synchronization instruction includes a descriptor of the tensor data, and,
[0192] The second core is configured to determine the destination address based on the descriptor.
[0193] Clause 12. The processing unit of clause 11, wherein the descriptor comprises at least shape information of the tensor data.
[0194] Clause 13. The processing unit of any one of clauses 9-12, wherein the second core is further configured to:
[0195] In response to receiving the data transfer end signal from the first core, sending an acknowledgement signal to the first core to indicate synchronization completion.
[0196] Clause 14. The processing unit according to any one of clauses 9-13, wherein the second core is configured to send the synchronization request signal in a semaphore manner.
[0197] Clause 15. A processing unit according to any one of clauses 9 to 14, wherein:
[0198] The second core and the first core are storage cores of different clusters in a multi-core processor architecture; or
[0199] The second core and the first core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
[0200] Clause 16. A processing unit according to any one of clauses 9 to 15, wherein:
[0201] The second synchronization instruction is a receive instruction, which indicates that the receiving end of the to-be-synchronized data of the related synchronization event is ready.
[0202] Clause 17. A multi-core processor, comprising a processing unit as described in any one of Clauses 1-8 and a processing unit as described in any one of Clauses 9-16.
[0203] Item 18. A chip, characterized in that the chip encapsulates a processing unit as described in any one of Items 1-8, or a processing unit as described in any one of Items 9-16, or a multi-core processor as described in Item 17.
[0204] Clause 19. A board, comprising the chip according to Clause 18.
[0205] Clause 20. A synchronization method for a processing unit, the processing unit comprising at least one first core, the method comprising:
[0206] In response to a first synchronization instruction associated with a second core, the first core queries whether there is an unprocessed synchronization event associated with the second core; and
[0207] The first core performs corresponding synchronization operations based on the query result.
[0208] Clause 21. The synchronization method of clause 20, further comprising:
[0209] When it is found that there is the unprocessed synchronization event, the data to be synchronized is transmitted to the second core; and
[0210] When the data transfer is completed, the data transfer completion signal is sent to the second core.
[0211] Clause 22. The synchronization method according to Clause 21, wherein the data transfer end signal is sent in the form of a semaphore.
[0212] Clause 23. The synchronization method according to any one of Clauses 20-22, wherein the method further comprises:
[0213] When no unprocessed synchronization event is found, waiting for a synchronization request signal from the second core; and
[0214] In response to receiving the synchronization request signal from the second core, the synchronization event is recorded.
[0215] Clause 24. The synchronization method according to any one of clauses 20-23, wherein the method further comprises:
[0216] The first core transmits data that needs to be synchronized to a predetermined address of the second core.
[0217] Clause 25. The synchronization method according to any one of clauses 20-24, wherein the method further comprises:
[0218] The first core ends the synchronization operation in response to receiving a response signal indicating completion of synchronization from the second core.
[0219] Clause 26. The synchronization method according to any one of clauses 20-25, wherein:
[0220] The first core and the second core are storage cores of different clusters in a multi-core processor architecture; or
[0221] The first core and the second core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
[0222] Clause 27. The synchronization method according to any one of clauses 20-26, wherein:
[0223] The first synchronization instruction is a send instruction, which indicates that the sender of the to-be-synchronized data of the related synchronization event is ready.
[0224] Clause 28. A synchronization method for a processing unit, the processing unit comprising at least one second core, the method comprising:
[0225] The second core sends a synchronization request signal to the first core in response to a second synchronization instruction associated with the first core to indicate a pending synchronization event; and
[0226] In response to receiving the data transmission completion signal from the first core, the second core obtains the transmitted data to be synchronized.
[0227] Clause 29. The synchronization method of clause 28, wherein the method further comprises:
[0228] The second core determines, in response to the data transfer end signal, a destination address of the data to be synchronized based on the second synchronization instruction;
[0229] The second core reads the transmitted data to be synchronized from a predetermined address for receiving the data to be synchronized transmitted from the first core; and
[0230] The second core writes the read data to be synchronized into the destination address.
[0231] Clause 30. The synchronization method according to any one of clauses 28-29, wherein the data to be synchronized is tensor data, the second synchronization instruction includes a descriptor of the tensor data, and,
[0232] The second core determining the destination address includes determining the destination address based on the descriptor.
[0233] Clause 31. A synchronization method according to Clause 30, wherein the descriptor includes at least shape information of the tensor data.
[0234] Clause 32. The synchronization method according to any one of clauses 28 to 31, wherein the method further comprises:
[0235] In response to receiving the data transfer end signal from the first core, the second core sends a response signal to the first core to indicate synchronization completion.
[0236] Clause 33. The synchronization method according to any one of clauses 28-32, wherein the synchronization request signal is sent in the form of a semaphore.
[0237] Clause 34. The synchronization method according to any one of clauses 28-33, wherein:
[0238] The second core and the first core are storage cores of different clusters in a multi-core processor architecture; or
[0239] The second core and the first core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
[0240] Clause 35. The synchronization method according to any one of clauses 28 to 34, wherein:
[0241] The second synchronization instruction is a receive instruction, which indicates that the receiving end of the to-be-synchronized data of the related synchronization event is ready.
[0242] Clause 36. A synchronization method for a multi-core processor, wherein the multi-core processor includes a first processing unit and a second processing unit, the method comprising:
[0243] The first processing unit performs the synchronization method according to any one of clauses 20-27; and
[0244] The second processing unit executes the synchronization method according to any one of clauses 28-35.
Claims
1. A processing unit, comprising at least one first core, wherein: The first core is configured to: In response to a first synchronization instruction associated with a second core, querying whether there is an unprocessed synchronization event associated with the second core; and Based on the result of the query, executing corresponding synchronization operations; The first core is further configured to: when the unprocessed synchronization event is found, transmit the data to be synchronized to the second core; in response to receiving the synchronization request signal from the second core, record the synchronization event; The synchronization request signal indicates a receive-ready state to the first core.
2. The processing unit according to claim 1, wherein: The first core is further configured to: When the data transfer is completed, the data transfer completion signal is sent to the second core.
3. The processing unit according to claim 2, wherein: The first core is configured to send the data transfer end signal in a semaphore manner.
4. The processing unit according to any one of claims 1 to 3, wherein: The first core is further configured to: When the unprocessed synchronization event is not found, the synchronization request signal from the second core is waited for.
5. The processing unit according to any one of claims 1 to 3, wherein: The first core is further configured to transmit data requiring synchronization to a predetermined address of the second core.
6. The processing unit according to any one of claims 1 to 3, wherein: The first core is further configured to: end the synchronization operation in response to receiving a response signal from the second core indicating that synchronization is complete.
7. The processing unit according to any one of claims 1 to 3, wherein: The first core and the second core are storage cores of different clusters in a multi-core processor architecture; or The first core and the second core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
8. The processing unit according to any one of claims 1 to 3, wherein: The first synchronization instruction is a send instruction, which indicates that the sender of the to-be-synchronized data of the related synchronization event is ready.
9. A processing unit comprising at least one second core, wherein: The second core is configured to: In response to a second synchronization instruction associated with the first core, sending a synchronization request signal to the first core to indicate a pending synchronization event; and In response to receiving a data transmission end signal from the first core, acquiring the transmitted data to be synchronized; The synchronization request signal indicates a receive-ready state to the first core; the data to be synchronized is tensor data, the second synchronization instruction includes a descriptor of the tensor data, and the second core is configured to determine a destination address based on the descriptor.
10. The processing unit of claim 9, wherein the second core is further configured to: In response to the data transfer end signal, determining a destination address of the data to be synchronized based on the second synchronization instruction; reading the transmitted data to be synchronized from a predetermined address for receiving the data to be synchronized transmitted from the first core; as well as The read data to be synchronized is written into the destination address. The processing unit according to claim 9 , wherein the descriptor comprises at least shape information of the tensor data.
12. The processing unit according to any one of claims 9 to 11, wherein the second core is further configured to: In response to receiving the data transfer end signal from the first core, sending an acknowledgement signal to the first core to indicate synchronization completion. 13 . The processing unit according to claim 9 , wherein the second core is configured to send the synchronization request signal in a semaphore manner.
14. The processing unit according to any one of claims 9 to 11, wherein: The second core and the first core are storage cores of different clusters in a multi-core processor architecture; or The second core and the first core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
15. The processing unit according to any one of claims 9 to 11, wherein: The second synchronization instruction is a receive instruction, which indicates that the receiving end of the to-be-synchronized data of the related synchronization event is ready.
16. A multi-core processor, characterized in that: The method comprises the processing unit according to any one of claims 1 to 8 and the processing unit according to any one of claims 9 to 15.
17. A chip, characterized in that: The chip encapsulates a processing unit according to any one of claims 1 to 8, or a processing unit according to any one of claims 9 to 15, or a multi-core processor according to claim 16.
18. A board, characterized in that: The board includes the chip according to claim 17.
19. A synchronization method for a processing unit, the processing unit comprising at least one first core, the method comprising: The first core queries, in response to a first synchronization instruction associated with the second core, whether there is an unprocessed synchronization event associated with the second core; as well as The first core performs a corresponding synchronization operation based on the query result; When it is found that there is an unprocessed synchronization event, the data to be synchronized is transmitted to the second core; In response to receiving a synchronization request signal from the second core, recording the synchronization event; The synchronization request signal indicates a receive-ready state to the first core.
20. The synchronization method according to claim 19, wherein: The method further comprises: When the data transfer is completed, the data transfer completion signal is sent to the second core.
21. The synchronization method according to claim 20, wherein: The data transmission completion signal is sent in the form of a semaphore.
22. The synchronization method according to any one of claims 19 to 21, wherein: The method further comprises: When the unprocessed synchronization event is not found, the synchronization request signal from the second core is waited for.
23. The synchronization method according to any one of claims 19 to 21, wherein the method further comprises: The first core transmits data that needs to be synchronized to a predetermined address of the second core.
24. The synchronization method according to any one of claims 19 to 21, wherein the method further comprises: The first core ends the synchronization operation in response to receiving a response signal indicating completion of synchronization from the second core.
25. The synchronization method according to any one of claims 19 to 21, wherein: The first core and the second core are storage cores of different clusters in a multi-core processor architecture; or The first core and the second core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
26. The synchronization method according to any one of claims 19 to 21, wherein: The first synchronization instruction is a send instruction, which indicates that the sender of the to-be-synchronized data of the related synchronization event is ready.
27. A synchronization method for a processing unit, the processing unit comprising at least one second core, the method comprising: The second core sends a synchronization request signal to the first core in response to a second synchronization instruction associated with the first core to indicate a pending synchronization event; as well as The second core acquires the transmitted data to be synchronized in response to receiving the data transmission end signal from the first core; The synchronization request signal indicates a receive-ready state to the first core; the data to be synchronized is tensor data, the second synchronization instruction includes a descriptor of the tensor data, and the second core is configured to determine a destination address based on the descriptor.
28. The synchronization method according to claim 27, wherein the method further comprises: The second core determines, in response to the data transfer end signal, a destination address of the data to be synchronized based on the second synchronization instruction; The second core reads the transmitted data to be synchronized from a predetermined address for receiving the data to be synchronized transmitted from the first core; as well as The second core writes the read data to be synchronized into the destination address. The synchronization method according to claim 27 , wherein the descriptor includes at least shape information of the tensor data.
30. The synchronization method according to any one of claims 27 to 29, wherein the method further comprises: In response to receiving the data transfer end signal from the first core, the second core sends a response signal to the first core to indicate synchronization completion.
31. The synchronization method according to any one of claims 27 to 29, wherein the synchronization request signal is sent in the form of a semaphore.
32. The synchronization method according to any one of claims 27 to 29, wherein: The second core and the first core are storage cores of different clusters in a multi-core processor architecture; or The second core and the first core are respectively processor cores or storage cores in the same cluster in a multi-core processor architecture.
33. The synchronization method according to any one of claims 27 to 29, wherein: The second synchronization instruction is a receive instruction, which indicates that the receiving end of the to-be-synchronized data of the related synchronization event is ready.
34. A synchronization method for a multi-core processor, wherein the multi-core processor includes a first processing unit and a second processing unit, the method comprising: The first processing unit executes the synchronization method according to any one of claims 19 to 26; as well as The second processing unit executes the synchronization method according to any one of claims 27-33.
Citation Information
Patent Citations
Communication method and device of processor cores based on array structure
CN102446157A
Information synchronization method and system
CN105915579A
Inter-core data transmission device and method
CN110046050A