Multi-core processor, synchronization method for multi-core processor and corresponding product

By introducing synchronization instructions and synchronization controllers into multi-core processors, the problem of collaborative work between multiple cores is solved, and the flexibility and processing efficiency of task scheduling are improved.

CN114281559BActive Publication Date: 2025-09-30ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011036264.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-27
Publication Date
2025-09-30
Estimated Expiration
2041-03-11

AI Technical Summary

Technical Problem

In a multi-core processor system, there are synchronization issues in the collaborative work between multiple cores, which leads to inflexible task scheduling and low processing efficiency.

Method used

By introducing synchronization instructions and synchronization controllers, the core determines the synchronization mode and range in response to the synchronization instructions, performs synchronization operations, and realizes collaborative work among multiple cores.

Benefits of technology

It improves the flexible scheduling and processing efficiency of multi-core processor tasks and ensures the collaborative work between cores in the multi-core processor architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114281559B_ABST
    Figure CN114281559B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a multi-core processor, a method for a multi-core processor, and related products. The multi-core processor can be implemented as a computing device included in a combined processing device, and the combined processing device can also include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete the computing operations specified by the user. The combined processing device can also include a storage device, which is connected to the computing device and the other processing devices respectively and is used to store data from the computing device and the other processing devices. The multi-core processor provided by the solution of the present disclosure can effectively achieve collaborative work between multiple cores through synchronization instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of processors, and in particular to a multi-core processor, a synchronization method, a chip, and a board for a multi-core processor. Background Art

[0002] With the advancement of computer technology, various applications (such as video structuring, ad recommendations, and intelligent translation) are placing increasing demands on machine storage and computing power. Because single-core processors are no longer sufficient for these applications, various multi-core processor systems have emerged. A key issue in multi-core processor systems is the coordinated operation of multiple cores. Therefore, how to achieve inter-core coordination within a multi-core architecture is an urgent problem to be solved. Summary of the Invention

[0003] To address one or more of the above-mentioned technical issues, the present disclosure provides, in various aspects, a multi-core processor and a synchronization method for the multi-core processor, wherein coordinated operation between multiple cores is achieved through synchronization instructions and a synchronization controller. The multi-core processor and corresponding synchronization method disclosed herein can solve the problem of instruction stream synchronization between multiple cores.

[0004] In a first aspect, the present disclosure provides a multi-core processor comprising multiple cores and one or more synchronization controllers, wherein the cores are configured to: determine a synchronization mode and a synchronization range of a synchronization instruction in response to a synchronization instruction; and send a synchronization signal to a corresponding synchronization controller based on the determined synchronization mode and synchronization range to perform a synchronization operation.

[0005] In a second aspect, the present disclosure provides a chip, in which a multi-core processor as described in any embodiment of the first aspect is packaged.

[0006] In a third aspect, the present disclosure provides a board comprising the chip of any one of the embodiments of the second aspect.

[0007] In a fourth aspect, the present disclosure provides a synchronization method for a multi-core processor, wherein the multi-core processor includes multiple cores and one or more synchronization controllers, and the method includes: the core determines the synchronization mode and synchronization range of the synchronization instruction in response to the synchronization instruction; and based on the determined synchronization mode and synchronization range, sends a synchronization signal to the corresponding synchronization controller to perform a synchronization operation.

[0008] Through the multi-core processor, synchronization method, chip and board provided above, the disclosed embodiments ensure the collaborative work between multiple cores in the multi-core processor architecture through synchronization instructions and synchronization controller, which is conducive to the flexible scheduling of multi-core processor tasks and the improvement of processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts.

[0010] Figure 1 An exemplary structural diagram showing a multi-core processor architecture to which the embodiments of the present disclosure may be applied;

[0011] Figure 2 An exemplary internal architectural diagram of a processor core is shown;

[0012] Figure 3A-3C A schematic structural block diagram of a multi-core processor according to an embodiment of the present disclosure is shown;

[0013] Figure 4A-4B A schematic flow chart showing a synchronization method according to an embodiment of the present disclosure;

[0014] Figure 5 A structural diagram showing a combined processing device according to an embodiment of the present disclosure; and

[0015] Figure 6 A schematic structural diagram of a board card according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.

[0017] It should be understood that the terms "first," "second," "third," and "fourth," etc., which may be used in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0018] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0019] As used in this specification and claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0020] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0021] Figure 1 An exemplary structural diagram of a multi-core processor architecture to which embodiments of the present disclosure may be applied is shown. The multi-core processor 100 may be used to process input data such as computer vision, speech, natural language, and data mining. Figure 1 The multi-core processor 100 in the embodiment of the present invention adopts a multi-core hierarchical structure design. The multi-core processor 100 can be used as a system on a chip, which can include multiple clusters (Clusters), and each cluster includes multiple cores (Cores). In other words, the multi-core processor 100 is composed of a hierarchy of system on a chip, clusters, and cores.

[0022] At the system-on-chip level, Figure 1 As shown, the multi-core processor 100 includes an external memory controller 111 , a peripheral communication module 112 , an on-chip interconnect module 113 and a plurality of clusters 115 .

[0023] There can be multiple external memory controllers 111, and two are shown as an example in the figure. They are used to respond to access requests issued by the processor core and access external memory devices (such as DRAM) to read data from outside the chip or write data. The peripheral communication module 112 is used to receive control signals from the processing device (not shown) through the interface device (not shown) to start the multi-core processor 100 to perform tasks. The on-chip interconnect module 113 connects the external memory controller 111, the peripheral communication module 112 and the multiple clusters 115 to transmit data and control signals between the modules. The multiple clusters 115 are the computing cores of the multi-core processor 100. In the figure, four are shown on each bare chip. With the development of hardware, the multi-core processor 100 disclosed in this disclosure can also include 8, 16, 64, or even more clusters 115. The clusters 115 are used to efficiently execute deep learning algorithms.

[0024] At the cluster level, Figure 1 As shown, each cluster 115 includes multiple processor cores (IPU cores) 121 and a memory core (MEM core) 122 .

[0025] The figure shows four processor cores 121 as an example, but the present disclosure does not limit the number of processor cores 121 .

[0026] The storage core 122 is primarily used for storage and communication, namely, to store shared data or intermediate results between the processor cores 121, and to perform communication between the execution cluster 115 and the DRAM 127, between the clusters 115, and between the processor cores 121. In other embodiments, the storage core 122 has scalar operation capabilities and is used to perform scalar operations.

[0027] The memory core 122 includes a shared memory unit (SMEM) 124, a broadcast bus 123, a cluster direct memory access (CDMA) module 126, and a global direct memory access (GDMA) module 125. SMEM 124 acts as a high-performance data transfer station. Data reused between different processor cores 121 within the same cluster 115 does not need to be obtained from each processor core 121 individually through DRAM 127. Instead, it is transferred between the processor cores 121 via SMEM 124. The memory core 122 only needs to quickly distribute the reused data from SMEM 124 to multiple processor cores 121, improving inter-core communication efficiency and significantly reducing on-chip and off-chip input / output accesses. The broadcast bus 123, CDMA 126, and GDMA 125 are used for communication between processor cores 121, communication between clusters 115, and data transfer between clusters 115 and DRAM 127, respectively. These are described below.

[0028] The broadcast bus 123 facilitates high-speed communication between the processor cores 121 within the cluster 115. In this embodiment, the broadcast bus 123 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (i.e., single-core to single-core) data transmission. Multicast transmits a copy of data from the SMEM 124 to a specific number of processor cores 121. Broadcast, a special case of multicast, transmits a copy of data from the SMEM 124 to all processor cores 121.

[0029] The CDMA 126 is used to control memory access to the SMEM 124 between different clusters 115 within the same multi-core processor 100 .

[0030] The GDMA 125 cooperates with the external memory controller 111 to control memory access from the SMEM 124 of the cluster 115 to the DRAM 127 , or to read data from the DRAM 127 to the SMEM 124 .

[0031] Although Figure 1 The multi-core processor architecture is described by taking a multi-core processor 100 of a single system on chip as an example. Those skilled in the art will appreciate that a multi-core processor can also be constructed using multiple single-core or multi-core processing units, and the present disclosure is not limited in this regard.

[0032] Figure 2 FIG. 1 shows an exemplary internal architecture diagram of the processor core 121. Figure 2 As shown, the processor core 121 may include three modules: a control module 21 , a calculation module 22 and a storage module 23 .

[0033] The control module 21 coordinates and controls the operations of the computing module 22 and the storage module 23 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 211 and an instruction decode unit (IDU) 212. The instruction fetch unit 211 retrieves instructions from a processing device (not shown), while the instruction decode unit 212 decodes the retrieved instructions and sends the decoded results as control information to the computing module 22 and the storage module 23.

[0034] The operation module 22 includes a vector operation unit 221 and a matrix operation unit 222. The vector operation unit 221 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformation. The matrix operation unit 222 is responsible for the core calculations of deep learning algorithms, such as matrix multiplication and convolution.

[0035] The storage module 23 is used to store or transfer relevant data, including neuron RAM (NRAM) 231, weight RAM (WRAM) 232, input / output direct memory access module (IODMA) 233, and move direct memory access module (MVDMA) 234. NRAM 231 is used to store input and output data and intermediate results for calculation by the processor core 121; WRAM 232 is used to store the weights of the deep learning network; IODMA 233 is used to transmit data to the processor core 121 through the broadcast bus 123 (see Figure 1 ) controls memory access between NRAM 231 / WRAM 232 and DRAM 127; MVDMA 234 controls memory access between NRAM 231 / WRAM 232 and SMEM 124.

[0036] In some embodiments, GDMA 125 (see Figure 1 ) functions and IODMA 233 (see Figure 2 ) can be integrated into the same component. For ease of description, this disclosure treats GDMA 125 and IODMA 233 as different components. For those skilled in the art, as long as the functions implemented and the technical effects achieved are similar to those disclosed herein, they fall within the scope of protection of this disclosure. Furthermore, the functions of GDMA 125, IODMA 233, CDMA 126, and MVDMA 234 can also be implemented by the same component. Similarly, as long as the functions implemented and the technical effects achieved are similar to those disclosed herein, they fall within the scope of protection of this disclosure.

[0037] In a multi-core processor architecture, multiple processes can run simultaneously on multiple cores, and some processes may have certain associations. Multiple processes may cooperate with each other to complete the same task, thus forming a synchronization relationship between the processes. Different processes may enter a competitive state in order to compete for limited system resources (hardware or software resources), thus forming a mutually exclusive relationship between the processes. If two cores access the same address area at the same time, there may be a memory access conflict. In order to solve the problem of multi-core memory access conflicts, the disclosed embodiment provides an instruction stream synchronization mechanism applied to a multi-core processor. In this synchronization mechanism, a synchronization instruction is provided. When the core executes the synchronization instruction, it interacts with the corresponding synchronization controller according to the synchronization mode and synchronization range indicated by the synchronization instruction to perform the corresponding synchronization operation. Based on the aforementioned multi-core processor architecture, instruction stream synchronization can include instruction stream synchronization between cores of different clusters, and instruction stream synchronization between cores within the same cluster (for example, storage cores or processor cores). The disclosed embodiment provides a synchronization mechanism for a multi-core processor architecture, which can be applied to any of the above-mentioned instruction stream synchronization scenarios.

[0038] Figure 3A FIG. 1 shows a schematic diagram of the structure of a multi-core processor according to an embodiment of the present disclosure. Figure 3A As shown, the multi-core processor 300 includes multiple (N shown in the figure) cores 330 and a synchronization controller 310. These cores 330 can be the processor cores or storage cores described above. The synchronization controller 310 is used to interact with the cores 330 to process or control synchronization events related to these cores.

[0039] In embodiments of the present disclosure, a synchronization instruction may include various information. For example, a synchronization instruction may include an identifier for identifying a synchronization event, a synchronization core count for indicating the number of cores involved in the synchronization event, and the like. In some implementations, the synchronization instruction may be a barrier instruction. The synchronization event identifier may also be referred to as a barrier identifier.

[0040] In some implementations, the synchronization instruction has different synchronization modes. Therefore, the synchronization instruction may include information indicating the synchronization mode. These synchronization modes may include, for example, a production-side synchronization mode, a consumer-side synchronization mode, and a full synchronization mode.

[0041] A classic data synchronization problem is the producer-consumer problem, in which a producer and consumer share the same storage space during the same time period. The producer generates data into the storage space, and the consumer retrieves data. Similar to data synchronization events, instruction stream synchronization events can also adopt a producer-consumer model, but instead of data transmission, they involve instruction execution. Specifically, the producer of an instruction stream synchronization event must first execute some instructions, and the consumer of the instruction stream synchronization event cannot execute subsequent instructions until the producer has completed those instructions.

[0042] Producer-side synchronization mode means that when the first core executes a synchronization instruction, it must wait for the previous instruction in the instruction stream containing the synchronization instruction to complete execution before the other core associated with the synchronization event (assuming it is the second core) can execute subsequent instructions. In other words, when the first core executes a synchronization instruction in producer-side synchronization mode, it must wait for the previous instruction in the instruction stream containing the synchronization instruction to complete execution. During this period, the execution of subsequent instructions in the first core will not be affected.

[0043] Consumer-side synchronization mode means that when the second core executes a synchronization instruction, it blocks the issuance of subsequent instructions in the instruction stream containing the synchronization instruction. Subsequent instructions can only be executed after confirming that the other cores associated with this synchronization event (such as the first core mentioned above) have completed executing the preceding instruction. In other words, when the second core executes a synchronization instruction in consumer-side synchronization mode, it blocks the issuance of subsequent instructions in the instruction stream containing the synchronization instruction.

[0044] Full synchronization mode means that when the first core executes the synchronization instruction, it must wait for the execution of the previous instruction in the instruction stream of the synchronization instruction in the first core to complete, and it must also block the issuance of the subsequent instruction in the instruction stream of the synchronization instruction in the first core. The other core in the synchronization event (for example, the second core) operates similarly. In other words, both parties in the synchronization event must wait for the other party's instructions to complete before executing subsequent operations.

[0045] The interaction between the core 330 and the synchronization controller 310 in various synchronization modes will be described in detail later in conjunction with the method flow chart.

[0046] In some implementations, the synchronization instruction has different synchronization scopes. Therefore, the synchronization instruction may include information for indicating the synchronization scope. As mentioned above, instruction stream synchronization may include instruction stream synchronization between cores of different clusters, as well as instruction stream synchronization between cores within the same cluster (e.g., storage cores or processor cores). In the disclosed embodiment, different synchronization controllers are provided to handle synchronization events within different scopes for instruction stream synchronization within different scopes.

[0047] Figure 3B FIG2 shows a schematic diagram of the structure of a multi-core processor according to an embodiment of the present disclosure, wherein a hardware solution for synchronizing instruction streams between cores (eg, storage cores or processor cores) within the same cluster is provided. Figure 3B As shown, a multi-core processor 300A includes multiple (five shown) cores 330 and a local synchronization controller 310A. These cores 330 are located in the same cluster and may be processor cores or storage cores within the same cluster as described above. The local synchronization controller 310A is used to interact with the cores 330 to process or control synchronization events between these cores.

[0048] Figure 3C FIG. 1 shows a schematic diagram of the structure of a multi-core processor according to another embodiment of the present disclosure, wherein a hardware solution for synchronizing instruction streams between cores in different clusters is provided. Figure 3C As shown, multi-core processor 300B includes multiple (16 shown) clusters 320 and a global synchronization controller 310B. Each cluster 320 may contain multiple (5 shown) cores 330, which may be processor cores or memory cores as described above. Global synchronization controller 310B is used to interact with clusters 320 to process or control synchronization events between cores in different clusters.

[0049] Those skilled in the art will understand that depending on the number and distribution of cores 330 and the number and distribution of clusters in a multi-core processor (e.g., 300, 300A, and 300B), there may be one or more local synchronization controllers and one or more global synchronization controllers, and the embodiments of the present disclosure are not limited in this regard.

[0050] The above describes various exemplary multi-core processor structures provided by the embodiments of the present disclosure. By introducing circuit modules such as local synchronization controllers and global synchronization controllers, synchronization mechanisms can be implemented between any two cores in the same cluster of a multi-core processor and between different clusters, thereby facilitating collaborative work between multiple processes.

[0051] The interaction between the core 330 and the synchronization controller ( 310 , 310A, 310B) will be described in detail below with reference to the method flow chart.

[0052] Figure 4A The exemplary flow chart of the synchronization operation performed at the core according to the embodiment of the present disclosure is schematically shown. The core may be, for example, the one previously referenced Figure 3A-3C The method 400A shows a method flow applied to any core of the exemplary multi-core processor described herein and executed due to a synchronization event caused by a synchronization instruction.

[0053] like Figure 4AAs shown, at step S410, any core in the multi-core processor executes a synchronization instruction when executing its local instruction stream. As described above, the synchronization instruction may have different synchronization modes and synchronization ranges, and accordingly, different synchronization operations will be taken and different synchronization controllers will be interacted with.

[0054] In step S420, the synchronization mode type of the synchronization instruction can be determined according to the synchronization mode field carried in the synchronization instruction. In the embodiment of the present disclosure, the synchronization mode includes at least the following three types: production-side synchronization mode, consumer-side synchronization mode and full synchronization mode.

[0055] If the synchronization instruction is determined to be in producer synchronization mode, method 400A proceeds to step S430, where it waits for the completion of execution of the preceding instruction in the instruction stream in which the synchronization instruction resides. Producer synchronization mode means that only after the core has completed execution of the preceding instruction can other cores synchronized with it execute subsequent instructions. In producer synchronization mode, the execution of subsequent instructions by the producer core is not affected.

[0056] When the execution of the previous instruction is completed, at step S431 , in response to the completion of the execution of the previous instruction, a first synchronization signal is sent to the corresponding synchronization controller to indicate a synchronization event.

[0057] At this time, the synchronization range of the synchronization instruction can be determined based on the synchronization range field carried in the synchronization instruction, thereby determining the corresponding synchronization controller. In the embodiment of the present disclosure, for inter-core synchronization, the synchronization range includes at least the following two types: inter-core synchronization within the same cluster and inter-core synchronization between different clusters. As mentioned above Figure 3B As described above, for inter-core synchronization within the same cluster, the local synchronization controller 310A can be responsible for handling synchronization events; and for inter-core synchronization between different clusters, such as Figure 3C As described above, the global synchronization controller 310B may be responsible for handling synchronization events.

[0058] Therefore, in step S431, when the synchronization scope is determined to be intra-cluster synchronization, a first synchronization signal is sent to the corresponding local synchronization controller; when the synchronization scope is determined to be inter-cluster synchronization, a first synchronization signal is sent to the corresponding global synchronization controller.

[0059] The first synchronization signal is used to indicate a synchronization event to the synchronization controller. In this synchronization event, the core sending the first synchronization signal is acting as the producer and is ready, meaning it has completed executing the preceding instructions. The first synchronization signal can include various information as needed. For example, the first synchronization signal can include a synchronization event identifier and a synchronization core count. The synchronization core count indicates the number of cores involved in the synchronization event.

[0060] Then, for this synchronization event, the synchronization operation at the production end can be terminated. Those skilled in the art will understand that in the production end synchronization mode, while waiting for the execution of the previous instruction, the core serving as the production end can still execute other instructions normally, that is, it will not block the execution of instructions following the synchronization instruction.

[0061] Returning to step S420, if the synchronization instruction is in consumer-side synchronization mode, method 400A proceeds to step S440, where the execution of subsequent instructions in the instruction stream containing the synchronization instruction is blocked. Consumer-side synchronization mode means that the core must wait for the execution of related pre-order instructions of other cores with which it is synchronized to complete before executing subsequent instructions.

[0062] At this time, the core needs to interact with the corresponding synchronization controller, for example, issuing a synchronization request, so as to obtain the instruction execution progress of other cores involved in the synchronization event.

[0063] Optionally or additionally, before interacting with the corresponding synchronization controller, the method 400A may further include step S441 , determining whether there are idle synchronization resources available for use.

[0064] Specifically, each core can maintain a synchronization request queue (BQ), which records the synchronization requests that have been issued, and each synchronization request can be marked with a synchronization event identifier (ID). Since the capacity of the synchronization request queue is limited and the synchronization event identifiers are also limited, new synchronization events can only be processed when there are idle synchronization request queue resources and idle synchronization event identifier resources, otherwise it is necessary to wait for the resources to be released. That is, at this time in step S441, it is determined whether the synchronization request queue is full and whether there are unused synchronization event identifiers. In other words, it is determined whether there are old requests in the synchronization request queue that have not received a response with the same synchronization event identifier to be allocated. If the queue is full or there are old requests in the queue with the same identifier that have not received a response, the new request will be back-pressed until the resources are released.

[0065] When the synchronization request queue is not full and there are unused synchronization event identifiers, the method proceeds to step S442 to allocate a synchronization event identifier for the new request (ie, the request for the synchronization event associated with the current synchronization instruction) and add it to the synchronization request queue.

[0066] Next, in step S443 , a second synchronization signal is sent to the corresponding synchronization controller.

[0067] Similar to the production-side synchronization mode, the synchronization scope of the synchronization instruction can be determined based on the synchronization scope field carried in the synchronization instruction, thereby determining the corresponding synchronization controller. Specifically, in step S443, if the synchronization scope is determined to be intra-cluster synchronization, a second synchronization signal is sent to the corresponding local synchronization controller; if the synchronization scope is determined to be inter-cluster synchronization, a second synchronization signal is sent to the corresponding global synchronization controller.

[0068] The second synchronization signal is used to indicate a synchronization event to the synchronization controller. In this synchronization event, the core that sends the second synchronization signal is the consumer end and is ready, that is, it has blocked the execution of subsequent instructions and is waiting for the previous instructions of other cores in the synchronization event to complete. Therefore, the second synchronization signal can also be called a synchronization request signal. As needed, the second synchronization signal can include various information. For example, the second synchronization signal can include a synchronization event identifier, a synchronization core count, etc. The synchronization core count is used to indicate the number of cores involved in the synchronization event.

[0069] Then, since the following instructions are blocked, the core can enter a sleep state until receiving a synchronization response signal from the synchronization controller (step S444). The synchronization response signal indicates that all parties in the synchronization event are ready, that is, the synchronization is successful.

[0070] At this time, in step S446, in response to receiving a synchronization response signal from the synchronization controller, the blocking is released, that is, the subsequent instructions can be executed.

[0071] Optionally or additionally, when the synchronization response signal is received, the method further includes step S445 of clearing the corresponding synchronization request from the synchronization request queue, thereby releasing synchronization request queue resources and synchronization event identification resources.

[0072] Thus, the synchronization operation at the consumer end can be terminated for the synchronization event. Those skilled in the art will understand that in the consumer-end synchronization mode, while the previous instructions of other cores waiting for the synchronization event are being executed, the core serving as the consumer end cannot execute subsequent instructions and can only enter a sleep state, waiting to be awakened by a synchronization response signal from the synchronization controller. In this sense, the first synchronization signal sent by the producer end can also be called a wake-up signal, which is used by the synchronization controller to wake up other cores synchronized with it that are in a sleep state due to waiting for the completion of execution of the previous instruction.

[0073] Returning to step S420 again, when it is determined that the mode of the synchronization instruction is the full synchronization mode, method 400A proceeds to step S450, where it waits for the execution of the previous instruction in the instruction stream where the synchronization instruction is located to be completed and blocks the issuance of the subsequent instruction in the instruction stream where the synchronization instruction is located.

[0074] Next, in step S451 , in response to the completion of the execution of the previous instruction, a third synchronization signal is sent to the corresponding synchronization controller to indicate a synchronization event and request synchronization.

[0075] Finally, in step S452 , in response to receiving a synchronization response signal from the synchronization controller, the blocking is released.

[0076] Fully synchronous mode means that the core acts as both a producer and a consumer. Therefore, the core must wait for the execution of the preceding instruction in the instruction stream containing the synchronization instruction within its own core to complete. It also needs to block the issuance of subsequent instructions in the instruction stream containing the synchronization instruction within its own core until the preceding instruction in the other cores with which it is synchronized completes execution. Therefore, the various steps described above for producer and consumer synchronization modes, including optional or additional steps, also apply to fully synchronous mode and are not detailed here.

[0077] The third synchronization signal is used to indicate a synchronization event to the synchronization controller. In this synchronization event, the core sending the third synchronization signal is both a producer and a consumer. As a producer, it is ready, that is, it has completed the execution of the preceding instruction. As a consumer, it is also ready, that is, it has blocked the execution of subsequent instructions and is waiting for the completion of the preceding instructions of other cores in the synchronization event. The third synchronization signal can include various information as needed. For example, the third synchronization signal can include a synchronization event identifier, a synchronization core count, etc. The synchronization core count is used to indicate the number of cores involved in the synchronization event.

[0078] Reference above Figure 4A Describes the synchronization operations performed at each core of a multi-core processor during a synchronization event. Figure 4B To describe the synchronization operation at the synchronization processor.

[0079] like Figure 4B As shown, it schematically shows an exemplary flow chart of the synchronization operation performed at the synchronization controller according to an embodiment of the present disclosure. The synchronization controller can be, for example, the one previously referred to. Figure 3A-3C The described synchronization controller 310, local synchronization controller 310A or global synchronization controller 310B. Method 400B shows a method flow for processing synchronization events applied to the synchronization controller.

[0080] At step S460, the synchronization controller receives synchronization signals from the cores it is responsible for. Different synchronization controllers handle synchronization events of different scopes. A local synchronization controller handles intra-cluster synchronization events and thus receives synchronization signals related to intra-cluster synchronization events of cores within a cluster for which it is responsible. A global synchronization controller handles inter-cluster synchronization events and thus receives synchronization signals related to inter-cluster synchronization events of cores in multiple clusters for which it is responsible.

[0081] Depending on the synchronization mode of the synchronization instruction executed by the core, different synchronization signals may be used. These synchronization signals may include, for example, the first synchronization signal, the second synchronization signal, and the third synchronization signal described above. The meaning of each synchronization signal can be found in the previous description and will not be repeated here.

[0082] Next, at step S470, the synchronization controller may associate and record the matching synchronization signals based on the received synchronization signals. Specifically, the synchronization controller may match the synchronization signals based on the synchronization event identifiers indicated in the synchronization signals. For example, signals with the same synchronization event identifier may be associated and recorded.

[0083] Finally, at step S480 , in response to the matched synchronization signal satisfying a predetermined condition, a synchronization response signal is sent to the corresponding core that initiated the synchronization request.

[0084] A synchronization event may involve synchronization operations between multiple cores, and this information can be indicated, for example, in the synchronization core count of the synchronization signal. Therefore, when the synchronization controller receives synchronization signals equal to the number indicated in the synchronization core count for a certain synchronization event (identified by the synchronization event identifier), it indicates that the synchronization of the synchronization event has been completed. That is, when the number of matching synchronization signals reaches the synchronization core count indicated in the synchronization signal, it is determined that the above-mentioned predetermined condition is met. Therefore, the synchronization controller can send a synchronization response signal to the core requesting synchronization in the synchronization event (that is, the core serving as the consumer end) to indicate that the synchronization is complete.

[0085] The instruction stream synchronization mechanism applied to a multi-core processor according to an embodiment of the present disclosure has been described above with reference to the flowchart. It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited to the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present disclosure.

[0086] It should be further explained that although Figure 4A and Figure 4B The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 4A and Figure 4B At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0087] Figure 5 FIG. 5 is a structural diagram showing a combined processing device 500 according to an embodiment of the present disclosure. Figure 5 As shown in FIG, the combined processing device 500 includes a computing device 502, an interface device 504, other processing devices 506, and a storage device 508. According to different application scenarios, the computing device may include one or more computing devices 510, which may be configured as Figure 3A-3C The multi-core processor shown is used to implement the present invention in combination with Figure 4A-4B The described operation.

[0088] In various embodiments, the computing and processing device of the present disclosure may be configured to perform user-specified operations. In exemplary applications, the computing and processing device may be implemented as a multi-core artificial intelligence processor. Similarly, one or more computing devices included in the computing and processing device may be implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core. When multiple computing devices are implemented as an artificial intelligence processor core or a partial hardware structure of an artificial intelligence processor core, the computing and processing device of the present disclosure may be considered to have a homogeneous multi-core structure.

[0089] In exemplary operation, the computing processing device of the present disclosure can interact with other processing devices through interface means, to jointly complete the operation specified by the user. Depending on the difference in implementation, the other processing devices of the present disclosure may include one or more types of processors in general and / or special processors such as central processing unit (Central Processing Unit, CPU), graphics processing unit (Graphics Processing Unit, GPU), artificial intelligence processor. These processors may include but are not limited to digital signal processor (Digital Signal Processor, DSP), application specific integrated circuit (Application Specific Integrated Circuit, ASIC), field programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As previously mentioned, only with respect to the computing processing device of the present disclosure, it can be regarded as having a homogeneous multi-core structure. However, when the computing processing device and other processing devices are considered together, the two can be regarded as forming a heterogeneous multi-core structure.

[0090] In one or more embodiments, the other processing device may serve as an interface between the computing device disclosed herein (which may be embodied as an artificial intelligence computing device such as a neural network computing device) and external data and control, performing basic control including but not limited to data transfer, starting and / or stopping the computing device, and so on. In other embodiments, the other processing device may also collaborate with the computing device to jointly complete computing tasks.

[0091] In one or more embodiments, the interface device can be used to transmit data and control instructions between the computing and processing device and other processing devices. For example, the computing and processing device can obtain input data from other processing devices via the interface device and write it to the storage device (or memory) on the computing and processing device chip. Furthermore, the computing and processing device can obtain control instructions from other processing devices via the interface device and write them to the control cache on the computing and processing device chip. Alternatively or optionally, the interface device can also read data from the storage device of the computing and processing device and transmit it to other processing devices.

[0092] Additionally or optionally, the combined processing device of the present disclosure may further include a storage device. As shown in the figure, the storage device is connected to the computing processing device and the other processing device, respectively. In one or more embodiments, the storage device may be used to store data of the computing processing device and / or the other processing device. For example, the data may be data that cannot be fully stored in the internal or on-chip storage device of the computing processing device or other processing device.

[0093] In some embodiments, the present disclosure also discloses a chip (e.g. Figure 6 In one implementation, the chip is a system on chip (SoC) and integrates one or more components such as Figure 5 The chip can be connected to the external interface device (such as Figure 6 The external interface device 606 shown in the figure is connected to other related components. The related components can be, for example, a camera, a display, a mouse, a keyboard, a network card or a wifi interface. In some application scenarios, other processing units (such as video codecs) and / or interface modules (such as DRAM interfaces) can be integrated on the chip. In some embodiments, the present disclosure also discloses a chip packaging structure, which includes the above-mentioned chip. In some embodiments, the present disclosure also discloses a board card, which includes the above-mentioned chip packaging structure. The following will be combined with Figure 6 The board is described in detail.

[0094] Figure 6 FIG. 1 is a schematic diagram showing the structure of a board 600 according to an embodiment of the present disclosure. Figure 6 As shown in , the board includes a storage device 604 for storing data, which includes one or more storage units 610. The storage device can be connected to the control device 608 and the chip 602 described above and transmit data by means of, for example, a bus. Further, the board also includes an external interface device 606, which is configured for data relay or transfer function between the chip (or the chip in the chip packaging structure) and the external device 612 (such as a server or computer, etc.). For example, the data to be processed can be passed to the chip by the external device through the external interface device. For another example, the calculation result of the chip can be transmitted back to the external device via the external interface device. According to different application scenarios, the external interface device can have different interface forms, for example, it can adopt a standard PCIE interface, etc.

[0095] In one or more embodiments, the control device in the disclosed board can be configured to regulate the state of the chip. To this end, in one application scenario, the control device can include a microcontroller unit (MCU) for regulating the working state of the chip.

[0096] According to the above combination Figure 5 and Figure 6 Based on the description, those skilled in the art can understand that the present disclosure also discloses an electronic device or apparatus, which may include one or more of the above-mentioned boards, one or more of the above-mentioned chips and / or one or more of the above-mentioned combined processing devices.

[0097] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.

[0098] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.

[0099] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this document divides them based on the consideration of logical functions, and there may be other ways of division in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.

[0100] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.

[0101] In some implementation scenarios, the above-mentioned integrated unit can be implemented in the form of a software program module. If implemented in the form of a software program module and sold or used as an independent product, the integrated unit can be stored in a computer-readable memory. Based on this, when the scheme of the present disclosure is embodied in the form of a software product (such as a computer-readable storage medium), the software product can be stored in a memory, which may include several instructions to enable a computer device (such as a personal computer, a server or a network device, etc.) to perform some or all of the steps of the method described in the embodiment of the present disclosure. The aforementioned memory may include, but is not limited to, various media that can store program code, such as a USB flash drive, a flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0102] In some other implementation scenarios, the above-mentioned integrated unit can also be implemented in the form of hardware, that is, a specific hardware circuit, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as CPUs, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.

[0103] The foregoing content can be better understood in accordance with the following terms:

[0104] Clause 1. A multi-core processor comprising a plurality of cores and one or more synchronization controllers, wherein:

[0105] The core is configured to:

[0106] In response to a synchronization instruction, determining a synchronization mode and a synchronization range of the synchronization instruction; and

[0107] Based on the determined synchronization mode and synchronization range, a synchronization signal is sent to a corresponding synchronization controller to perform a synchronization operation.

[0108] Clause 2. The multi-core processor of clause 1, wherein the core is further configured to:

[0109] When the synchronization mode is the production-side synchronization mode, waiting for the execution of the previous instruction in the instruction stream where the synchronization instruction is located to be completed; and

[0110] In response to completion of execution of the previous instruction, a first synchronization signal is sent to a corresponding synchronization controller to indicate a synchronization event.

[0111] Clause 3. The multi-core processor of any one of clauses 1-2, wherein the core is further configured to:

[0112] When the synchronization mode is a consumer-side synchronization mode, blocking the transmission of subsequent instructions in the instruction stream where the synchronization instruction is located;

[0113] sending a second synchronization signal to a corresponding synchronization controller to request synchronization; and

[0114] In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

[0115] Clause 4. The multi-core processor of any one of clauses 1-3, wherein the core is further configured to:

[0116] When the synchronization mode is the full synchronization mode, waiting for the completion of execution of the previous instruction in the instruction stream where the synchronization instruction is located and blocking the issuance of the subsequent instruction in the instruction stream where the synchronization instruction is located;

[0117] In response to completion of execution of the previous instruction, sending a third synchronization signal to a corresponding synchronization controller to indicate a synchronization event and request synchronization; and

[0118] In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

[0119] Clause 5. The multi-core processor of any one of clauses 3-4, wherein the core is further configured to:

[0120] Before sending a synchronization signal to the synchronization controller, determining whether a synchronization request queue is full and whether there is an unused synchronization event identifier;

[0121] When the synchronization request queue is not full and there is an unused synchronization event identifier, adding the synchronization request associated with the synchronization instruction to the synchronization request queue and sending the synchronization signal to the synchronization controller; and

[0122] In response to receiving the synchronization response signal from the synchronization controller, the corresponding synchronization request is cleared from the synchronization request queue.

[0123] Clause 6. The multi-core processor of any one of clauses 1-5, wherein the core is further configured to:

[0124] When the synchronization range is intra-cluster synchronization, sending the synchronization signal to the corresponding local synchronization controller; and / or

[0125] When the synchronization range is inter-cluster synchronization, the synchronization signal is sent to the corresponding global synchronization controller.

[0126] Clause 7. The multi-core processor according to any one of clauses 1-6, wherein the synchronization controller is configured to:

[0127] receiving one or more synchronization signals from the core;

[0128] Associating and recording the matching synchronization signals; and

[0129] In response to the matched synchronization signal satisfying a predetermined condition, a synchronization response signal is sent to the corresponding core that initiated the synchronization request.

[0130] Clause 8. The multi-core processor of clause 7, wherein the synchronization controller is further configured to:

[0131] matching the synchronization signal based on an identification of a synchronization event indicated in the synchronization signal; and

[0132] When the number of matched synchronization signals reaches the synchronization core count indicated in the synchronization signal, it is determined that the predetermined condition is satisfied.

[0133] Clause 9. A multi-core processor according to any one of clauses 1-8, wherein the synchronization instruction is a barrier instruction, and the barrier instruction includes at least a barrier identifier for identifying a synchronization event and a synchronization core count for indicating the number of cores involved in the synchronization event.

[0134] Clause 10. A chip, characterized in that the multi-core processor as described in any one of Clauses 1-9 is encapsulated in the chip.

[0135] Clause 11. A board, characterized in that the board includes the chip described in Clause 10.

[0136] Clause 12. A synchronization method for a multi-core processor, the multi-core processor comprising a plurality of cores and one or more synchronization controllers, the method comprising:

[0137] The core determines, in response to a synchronization instruction, a synchronization mode and a synchronization range of the synchronization instruction; and

[0138] Based on the determined synchronization mode and synchronization range, a synchronization signal is sent to a corresponding synchronization controller to perform a synchronization operation.

[0139] Clause 13. The synchronization method of clause 12, wherein the method further comprises:

[0140] When the synchronization mode is a production-side synchronization mode, the core waits for the execution of a previous instruction in the instruction stream where the synchronization instruction is located to be completed; and

[0141] In response to completion of execution of the previous instruction, a first synchronization signal is sent to a corresponding synchronization controller to indicate a synchronization event.

[0142] Clause 14. The synchronization method according to any one of clauses 12-13, wherein the method further comprises:

[0143] When the synchronization mode is a consumer-side synchronization mode, the core blocks the issuance of subsequent instructions in the instruction stream where the synchronization instruction is located;

[0144] sending a second synchronization signal to a corresponding synchronization controller to request synchronization; and

[0145] In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

[0146] Clause 15. The synchronization method according to any one of clauses 12-14, wherein the method further comprises:

[0147] When the synchronization mode is the full synchronization mode, the core waits for the execution of the preceding instruction in the instruction stream where the synchronization instruction is located to be completed and blocks the issuance of the following instruction in the instruction stream where the synchronization instruction is located;

[0148] In response to completion of execution of the previous instruction, sending a third synchronization signal to a corresponding synchronization controller to indicate a synchronization event and request synchronization; and

[0149] In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

[0150] Clause 16. The synchronization method according to any one of clauses 12-15, wherein the method further comprises:

[0151] Before sending the synchronization signal to the synchronization controller, the core determines whether the synchronization request queue is full and whether there is an unused synchronization event identifier;

[0152] When the synchronization request queue is not full and there is an unused synchronization event identifier, adding the synchronization request associated with the synchronization instruction to the synchronization request queue and sending the synchronization signal to the synchronization controller; and

[0153] In response to receiving the synchronization response signal from the synchronization controller, the corresponding synchronization request is cleared from the synchronization request queue.

[0154] Clause 17. The synchronization method according to any one of clauses 12-16, wherein the method further comprises:

[0155] When the synchronization range is intra-cluster synchronization, the core sends the synchronization signal to the corresponding local synchronization controller; and / or

[0156] When the synchronization scope is inter-cluster synchronization, the core sends the synchronization signal to the corresponding global synchronization controller.

[0157] Clause 18. The synchronization method according to any one of clauses 12-17, wherein the method further comprises:

[0158] The synchronization controller receives one or more synchronization signals from the core;

[0159] Associating and recording the matching synchronization signals; and

[0160] In response to the matched synchronization signal satisfying a predetermined condition, a synchronization response signal is sent to the corresponding core that initiated the synchronization request.

[0161] Clause 19. The synchronization method of clause 18, wherein the method further comprises:

[0162] The synchronization controller matches the synchronization signal based on an identification of a synchronization event indicated in the synchronization signal; and

[0163] When the number of matched synchronization signals reaches the synchronization core count indicated in the synchronization signal, it is determined that the predetermined condition is satisfied.

[0164] Clause 20. The synchronization method according to any one of clauses 12-19, wherein the synchronization instruction is a barrier instruction, and the barrier instruction includes at least a barrier identifier for identifying a synchronization event and a synchronization core count for indicating the number of cores involved in the synchronization event.

Claims

1. A multi-core processor comprising a plurality of cores and one or more synchronization controllers, wherein the multi-core processor adopts a multi-core hierarchical structure design and comprises a plurality of clusters, each cluster comprising a plurality of the cores, wherein: The core is configured to: In response to a synchronization instruction, determining a synchronization mode and a synchronization range of the synchronization instruction, wherein the synchronization mode includes at least a production-end synchronization mode, a consumer-end synchronization mode, and a full synchronization mode, and the synchronization range includes intra-cluster synchronization and inter-cluster synchronization; as well as Based on the determined synchronization mode and synchronization range, a synchronization signal is sent to the corresponding synchronization controller to perform a synchronization operation, wherein the synchronization instruction has different synchronization ranges, and the synchronization instruction includes information for indicating the synchronization range. For instruction stream synchronization within different ranges, different synchronization controllers are provided to handle synchronization events within different ranges.

2. The multi-core processor of claim 1 , wherein the core is further configured to: When the synchronization mode is the production-side synchronization mode, waiting for the execution of the previous instruction in the instruction stream where the synchronization instruction is located to be completed; and In response to completion of execution of the previous instruction, a first synchronization signal is sent to a corresponding synchronization controller to indicate a synchronization event.

3. The multi-core processor of claim 1 , wherein the core is further configured to: When the synchronization mode is a consumer-side synchronization mode, blocking the transmission of subsequent instructions in the instruction stream where the synchronization instruction is located; sending a second synchronization signal to a corresponding synchronization controller to request synchronization; and In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

4. The multi-core processor of claim 1 , wherein the core is further configured to: When the synchronization mode is the full synchronization mode, waiting for the completion of execution of the previous instruction in the instruction stream where the synchronization instruction is located and blocking the issuance of the subsequent instruction in the instruction stream where the synchronization instruction is located; In response to completion of execution of the previous instruction, sending a third synchronization signal to a corresponding synchronization controller to indicate a synchronization event and request synchronization; and In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

5. The multi-core processor according to any one of claims 3 to 4, wherein the core is further configured to: Before sending a synchronization signal to the synchronization controller, determining whether a synchronization request queue is full and whether there is an unused synchronization event identifier; When the synchronization request queue is not full and there is an unused synchronization event identifier, adding the synchronization request associated with the synchronization instruction to the synchronization request queue and sending the synchronization signal to the synchronization controller; and In response to receiving the synchronization response signal from the synchronization controller, the corresponding synchronization request is cleared from the synchronization request queue.

6. The multi-core processor of claim 1 , wherein the core is further configured to: When the synchronization range is intra-cluster synchronization, sending the synchronization signal to the corresponding local synchronization controller; and / or When the synchronization range is inter-cluster synchronization, the synchronization signal is sent to the corresponding global synchronization controller.

7. The multi-core processor according to claim 1, wherein the synchronization controller is configured to: receiving one or more synchronization signals from the core; Associating and recording the matching synchronization signals; and In response to the matched synchronization signal satisfying a predetermined condition, a synchronization response signal is sent to the corresponding core that initiated the synchronization request.

8. The multi-core processor according to claim 7, wherein the synchronization controller is further configured to: matching the synchronization signal based on an identification of a synchronization event indicated in the synchronization signal; and When the number of matched synchronization signals reaches the synchronization core count indicated in the synchronization signal, it is determined that the predetermined condition is satisfied. 9 . The multi-core processor according to claim 1 , wherein the synchronization instruction is a barrier instruction, and the barrier instruction includes at least a barrier identifier for identifying a synchronization event and a synchronization core count for indicating the number of cores involved in the synchronization event.

10. A chip, characterized in that: The chip encapsulates the multi-core processor according to any one of claims 1 to 9.

11. A board, characterized in that: The board includes the chip according to claim 10.

12. A synchronization method for a multi-core processor, the multi-core processor comprising multiple cores and one or more synchronization controllers, the multi-core processor being designed with a multi-core hierarchical structure and comprising multiple clusters, each cluster comprising multiple cores, the method comprising: The core determines, in response to the synchronization instruction, a synchronization mode and a synchronization range of the synchronization instruction, wherein the synchronization mode includes at least a production-end synchronization mode, a consumer-end synchronization mode, and a full synchronization mode, and the synchronization range includes intra-cluster synchronization and inter-cluster synchronization; as well as Based on the determined synchronization mode and synchronization range, a synchronization signal is sent to the corresponding synchronization controller to perform a synchronization operation, wherein the synchronization instruction has different synchronization ranges, and the synchronization instruction includes information for indicating the synchronization range. For instruction stream synchronization within different ranges, different synchronization controllers are provided to handle synchronization events within different ranges.

13. The synchronization method according to claim 12, wherein the method further comprises: When the synchronization mode is a production-side synchronization mode, the core waits for the execution of a previous instruction in the instruction stream where the synchronization instruction is located to be completed; as well as In response to completion of execution of the previous instruction, a first synchronization signal is sent to a corresponding synchronization controller to indicate a synchronization event.

14. The synchronization method according to claim 12, wherein the method further comprises: When the synchronization mode is a consumer-side synchronization mode, the core blocks the issuance of subsequent instructions in the instruction stream where the synchronization instruction is located; sending a second synchronization signal to a corresponding synchronization controller to request synchronization; and In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

15. The synchronization method according to claim 12, wherein the method further comprises: When the synchronization mode is the full synchronization mode, the core waits for the execution of the preceding instruction in the instruction stream where the synchronization instruction is located to be completed and blocks the issuance of the following instruction in the instruction stream where the synchronization instruction is located; In response to completion of execution of the previous instruction, sending a third synchronization signal to a corresponding synchronization controller to indicate a synchronization event and request synchronization; and In response to receiving a synchronization acknowledgement signal from the synchronization controller, the blocking is released.

16. The synchronization method according to any one of claims 12 to 15, wherein the method further comprises: The core determines whether a synchronization request queue is full and whether there is an unused synchronization event identifier before sending a synchronization signal to the synchronization controller; When the synchronization request queue is not full and there is an unused synchronization event identifier, adding the synchronization request associated with the synchronization instruction to the synchronization request queue and sending the synchronization signal to the synchronization controller; as well as In response to receiving the synchronization response signal from the synchronization controller, the corresponding synchronization request is cleared from the synchronization request queue.

17. The synchronization method according to claim 12, wherein the method further comprises: When the synchronization range is intra-cluster synchronization, the core sends the synchronization signal to the corresponding local synchronization controller; and / or When the synchronization scope is inter-cluster synchronization, the core sends the synchronization signal to the corresponding global synchronization controller.

18. The synchronization method according to claim 12, wherein the method further comprises: The synchronization controller receives one or more synchronization signals from the core; Associate and record the matching synchronization signals; as well as In response to the matched synchronization signal satisfying a predetermined condition, a synchronization response signal is sent to the corresponding core that initiated the synchronization request.

19. The synchronization method according to claim 18, wherein the method further comprises: The synchronization controller matches the synchronization signal based on an identifier of a synchronization event indicated in the synchronization signal; as well as When the number of matched synchronization signals reaches the synchronization core count indicated in the synchronization signal, it is determined that the predetermined condition is satisfied. 20 . The synchronization method according to claim 12 , wherein the synchronization instruction is a barrier instruction, and the barrier instruction at least includes a barrier identifier for identifying a synchronization event and a synchronization core count for indicating the number of cores involved in the synchronization event.

Citation Information

Patent Citations

  • Synchronous processing method and device based on multicore system

    CN102334104A

  • Message type internal memory accessing device and accessing method thereof

    CN102609378A

  • Method and apparatus for a hierarchical synchronization barrier in a multi-node system

    US20120179896A1