Processor, method for data processing, device, and storage medium

The integration of a distributor in multi-core processors for centralized data and instruction distribution addresses bandwidth limitations, enhancing efficiency in vector calculations by reducing cache coherence issues and improving data transmission.

JP2025521252AActive Publication Date: 2025-07-08BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024572689
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-14
Filing Date
2023-06-06
Publication Date
2025-07-08
Estimated Expiration
2043-06-06

AI Technical Summary

Technical Problem

Existing multi-core processors face inefficiencies in utilizing limited bandwidth for vector calculations with many instruction repetitions and large data volumes, leading to cache coherence issues and reduced data read/write efficiency.

Method used

A distributor is integrated into the processor to centrally manage data and instruction distribution, bypassing conventional cache coherence designs, allowing for large-capacity data caching and efficient data exchange through broadcast methods.

Benefits of technology

This approach enhances data transmission efficiency, maximizes bandwidth utilization, and improves computing performance, particularly in vector calculations like neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521252000001_ABST
    Figure 2025521252000001_ABST
Patent Text Reader

Abstract

According to an embodiment of the present invention, a processor, a method for data processing, a device, and a storage medium are provided. The processor includes a plurality of processor cores, and each of the plurality of processor cores includes a data cache for reading and writing data and an instruction cache separated from the data cache for reading instructions. The processor further includes a distributor communicatively coupled to the plurality of processor cores. The distributor is configured to distribute data to be processed to a corresponding data cache of at least one of the plurality of processor cores and distribute instructions associated with the data to be processed to a corresponding instruction cache of at least one of the plurality of processor cores for execution. Thereby, the efficiency of data transmission and vector calculation can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - reference to Related Applications] This application claims the priority of a Chinese invention patent application filed on June 14, 2022, with the invention title "Processor, Method for Data Processing, Device, and Storage Medium", and the application number 202210674851.9, the entire content of which is incorporated herein by reference.

[0002] [Technical Field] Exemplary embodiments of the present invention generally relate to the field of computers, and in particular, to processors, methods for data processing, devices, and computer - readable storage media.

Background Art

[0003] With the development of information technology, various data - processing services have increasingly higher requirements for the computing power and computing resources of computing systems. Currently, it has been proposed to use multi - core processors to improve the overall computing power and computing throughput of the system by means of parallel computing. In some vector calculations with a large number of instruction repetitions and a large amount of data, how to fully utilize the limited bandwidth of multi - core processors to process such vector calculations has become a problem worthy of attention.

Summary of the Invention

[0004] In a first aspect of the present invention, a processor is provided. The processor includes a plurality of processor cores, and each of the plurality of processor cores includes a data cache for reading and writing data and an instruction cache separated from the data cache for reading instructions. The processor further includes a distributor communicatively coupled to the plurality of processor cores. The distributor is configured to distribute data to be processed to a corresponding data cache of at least one of the plurality of processor cores and to distribute instructions associated with the data to be processed to a corresponding instruction cache of at least one of the plurality of processor cores for execution.

[0005] In a second aspect of the present invention, a method for data processing is provided. The method includes distributing, by a distributor of a processor, data to be processed to a corresponding data cache of at least one of a plurality of processor cores of the processor. The distributor is communicatively coupled to the plurality of processor cores. The method further includes distributing instructions associated with the data to be processed to a corresponding instruction cache of at least one of the plurality of processor cores for execution.

[0006] In a third aspect of the present invention, an electronic device is provided. The electronic device includes at least the processor of the first aspect.

[0007] In a fourth aspect of the present invention, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method of the second aspect is implemented.

[0008] It should be understood that the content described in the summary part of the present invention is not intended to limit the main features or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will be readily understood from the following description.

Brief Description of the Drawings

[0009] With reference to the following detailed description in conjunction with the drawings, the above-described features and other features, advantages, and aspects of each embodiment of the present invention will become more apparent. In the drawings, the same or similar symbols indicate the same or similar elements, where

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in more detail with reference to the drawings. Although specific embodiments of the present invention are shown in the drawings, the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not used to limit the protection scope of the present invention.

[0011] In the description of the embodiments of the present invention, the term "including" and its similar terms are open-ended inclusion of "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may be other explicit and implicit definitions hereinafter.

[0012] It is understood that the data related to the present technical solution (including but not limited to the data itself, data acquisition, or data use) should comply with the corresponding laws and regulations and related designated requirements.

[0013] As described above, with the development of information technology, various data processing services have increasingly higher requirements for the computing capabilities and computing resources of computing systems. Currently, it has been proposed to use multi-core processors to improve the overall computing capabilities and computing throughput of the system by means of parallel computing. Generally, each processor core of a multi-core processor is an independent and complete instruction execution unit. When multiple processor cores operate in cooperation, for example, when multiple processor cores need to access the same storage address, there may be a problem of instruction or data competition.

[0014] Conventional solutions to address the above-mentioned instruction or data conflict problems involve using a mailbox policy for instructions. For example, when multiple processor cores need to execute operations synchronously, they send instructions to each other via a mailbox. In the case of data, conventional solutions use cache coherence policies such as Modified Exclusive Shared Invalidate (MESI) technology to ensure that the data in the cache is the same as the data in the main memory. Multi-core processor architectures that use the above-mentioned conventional policies include the big.LITTLE architecture and the like.

[0015] According to research, in some vector calculations with a large number of instruction repetitions and a large data volume, it has been found that single instruction multiple data (SIMD) processors of multi-core architectures are often used. On the other hand, when using a conventional multi-core processor architecture based on cache coherence, the first-level cache of the SIMD processor becomes smaller. Also, due to a large amount of calculation data and low locality, a large number of cache misses occur frequently, resulting in a decrease in data read / write efficiency. On the other hand, in such normal vector calculations, complex thread switching tasks are rarely required, and the conventional multi-core processor control mode becomes bloated and appears redundant.

[0016] In summary, in some vector calculations with a large number of instruction repetitions and a large data volume, how to make full use of the limited bandwidth of multi-core processors to process such vector calculations has become a problem worthy of attention.

[0017] According to an embodiment of the present invention, an improved solution for a processor is proposed. In this solution, a distributor is provided in the processor to distribute data and / or instructions to each processor core of the processor. By using the distributor to centrally schedule the distribution of data and / or instructions without using the conventional cache coherence design, many cache coherence problems of the multi-core processor are avoided.

[0018] On the other hand, in the case of a conventional processor based on cache coherence, the complexity of data transmission and processing is often high. As a result, the data storage capacity is limited and it is difficult to increase the clock frequency. In this solution, the distributor distributes data to the data cache of each processor core for reading and writing data. By directly transmitting data using the respective data caches of each processor core and the distributor, a large-capacity data cache can be utilized and data exchange with the outside can be reduced as much as possible.

[0019] On the other hand, in the conventional multi-core scheduling method, usually, since the processor actively initiates a data transmission request, a single processor cannot know the data requests of other processors, so it is difficult to make it compatible with the data broadcast method. That is, data cannot be transmitted by the broadcast method. This solution uses a centralized data scheduling mechanism and can easily transmit data by applying the broadcast method, further improving the data transmission efficiency. In this way, this solution can make full use of limited bandwidth resources and further improve the efficiency of vector calculation, especially the vector calculation of neural networks.

[0020] Hereinafter, some exemplary embodiments of the present invention will be described with continued reference to the drawings.

[0021] FIG. 1 shows a schematic diagram of an exemplary environment 100 in which embodiments of the present invention can be implemented. In the exemplary environment 100, the processor 101 is a multi-core processor including processor cores 120-1, 120-2, …, 120-N, where N is an integer greater than 1. For ease of explanation, hereinafter, the processor cores 120-1, 120-2, …, 120-N are collectively or individually referred to as processor core 120. Each processor core 120 may be a SIMD processor core. In some embodiments, the processor 101 may include four processor cores 120 (i.e., the value of N may be 4).

[0022] Of course, it should be understood that any specific numerical values appearing herein and in other places in this specification are exemplary unless otherwise specified. For example, in other embodiments, depending on different metrics such as the process level and line width of the flow chip, there may be a different number of processor cores 120 accordingly.

[0023] The processor 101 further includes a distributor 110. The distributor 110 is communicatively coupled to each processor core 120. That is, the distributor 110 and each processor core 120 can communicate with each other according to an appropriate data transmission protocol and / or standard. During operation, the distributor 110 can distribute data 140 and / or instructions 130 to each processor core 120. It should be noted that the distributor 110 is sometimes referred to as a "scheduler", and in this context, the two may be used interchangeably.

[0024] In some embodiments, the distributor 110 can be implemented as a hardware circuit. The hardware circuit may be integrated or embedded in the processor 101. Alternatively or additionally, the distributor 110 may be implemented in whole or in part by software modules, which may be implemented, for example, as executable instructions and stored in a memory (not shown).

[0025] The distributor 110 is configured to distribute instructions 130 and / or data 140 received by the processor 101 from other devices (not shown) in the environment 100 to each processor core 120. The device that sends the instructions 130 and / or data 140 to the distributor 110 is also referred to as the originating device of the data processing request. In some embodiments, the processor 101 receives data 140 and / or instructions 130 transmitted by a storage device or other external device in the environment 100, for example, via a bus, and distributes the data 140 and / or instructions 130 received via the distributor 110 to each processor core 120. The distribution process for the data 140 and / or instructions 130 will be described below in conjunction with FIG. 2.

[0026] It should be understood that the configuration and functions of the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present invention. For example, the processor 101 can be applied to various existing or future computing platforms or computing systems. The processor 101 can be implemented in various embedded applications (such as data processing systems such as mobile network base stations) to provide services such as large-scale vector calculations. The processor 101 may be integrated or embedded in various electronic devices or computing devices to provide various computing services. The application environment and application scenarios of the processor 101 are not limited here.

[0027] FIG. 2 shows a schematic diagram of an exemplary architecture 200 for the distribution of instruction 130 and data 140 according to some embodiments of the present invention. For ease of explanation, architecture 200 will be described with reference to environment 100 of FIG. 1.

[0028] As shown in FIG. 2, each of the plurality of processor cores 120 includes a data cache for reading and writing data, and an instruction cache separated from the data cache for reading instructions. For example, processor core 120-1 includes instruction cache 220-1 and data cache 230-1, processor core 120-2 includes instruction cache 220-2 and data cache 230-2, …, processor core 120-N includes instruction cache 220-N and data cache 230-N. For ease of explanation, hereinafter, instruction caches 220-1, 220-2, …, 220-N will be collectively or individually referred to as instruction cache 220, and data caches 230-1, 230-2, …, 230-N will be collectively or individually referred to as data cache 230.

[0029] Note that instruction cache 220 is typically not implemented as a cache. From the perspective of processor 101, instruction cache 220 is read-only. Data cache 230 can include a Vector Closely-coupled Memory (VCCM). Similar to instruction cache 220, data cache 230 is typically not implemented as a cache. However, unlike instruction cache 220, data cache 230 is readable and writable. By using VCCM as data cache 230 instead of using a cache, the complexity of the design of processor 101 can be reduced, and the processor cache capacity and clock frequency can be increased. In this way, the data transmission efficiency of processor 101 can be improved.

[0030] In some embodiments, the distributor 110 is configured to distribute the received instructions 130 and / or data 140 to at least one of the plurality of processor cores 120. For example, the distributor 110 may be configured to distribute the received instructions 130 and / or data 140 only to the first processor core 120-1. Also, for example, the distributor 110 may distribute the received instructions 130 and / or data 140 to the first processor core 120-1 and the second processor core 120-2. Alternatively, the distributor 110 may distribute the instructions 130 and / or data 140 to each of the plurality of processor cores 120. In some embodiments, the distributor 110 can receive configuration information 210. The configuration information 210 can instruct the distributor 110 to distribute the instructions 130 and / or data 140 to one or some of the plurality of processor cores 120. For example, the configuration information 210 can instruct the distributor 110 to distribute the instructions 130 and / or data 140 only to the first processor core 120-1. Also, for example, the configuration information 210 can instruct the distributor 110 to distribute the instructions 130 and / or data 140 to each of the plurality of processor cores 120. Alternatively or additionally, the distributor 110 may be pre-programmed to distribute the instructions 130 and / or data 140 to one or some, or all, of the processor cores 120.

[0031] In some embodiments, the distributor 110 can receive a data set and an instruction set to be processed by the processor 101. The configuration information 210 received by the distributor 110 can at least indicate an association between the data to be processed (also referred to as data 140) in the data set and the instruction 130 in the instruction set. For example, the association between the data 140 and the instruction 130 can indicate that the data 140 in the data set is processed according to the instruction 130 in the instruction set. The distributor 110 can distribute the data 140 and the instruction 130 at least partially depending on the association. For example, the distributor 110 distributes the data 140 to the corresponding data cache 230 of at least one of the plurality of processor cores 120, and distributes the instruction 130 associated with the data 140 to the corresponding instruction cache 220 of at least one of the processor cores 120 for execution.

[0032] In some embodiments, the distributor 110 can broadcast the same instruction 130 to at least one processor core. FIG. 3 shows a schematic diagram for broadcasting the instruction 130 to a plurality of processor cores according to some embodiments of the present invention. As shown in FIG. 3, the instruction 130 may include instruction 0, instruction 1, …, instruction M (where M is an integer greater than or equal to 1). The instruction 130 may be broadcast to the processor cores 120-1, 120-2, …, 120-J (where J is an integer greater than 1) via the distributor 110. For example, the same instructions 310-1, 310-2, …, 310-J as the instruction 130 are broadcast to the processor cores 120-1, 120-2, …, 120-J. For ease of explanation, hereinafter, the instructions 310-1, 310-2, …, 310-J are collectively or individually referred to as instruction 310. It should be understood that there may be a delay 320 between the time when the instruction 310 is received in the processor core 120 and the time when the instruction 130 is received in the distributor 110.

[0033] FIG. 3 shows that instruction 130 is broadcast to processor cores 120-1, 120-2, …, 120-J. However, in some embodiments, it should be understood that distributor 110 can broadcast instruction 130 to more or fewer of the plurality of processor cores 120. For example, distributor 110 can broadcast instruction 130 to only one processor core 120. Also, for example, distributor 110 can broadcast instruction 130 to all processor cores 120 (i.e., J is equal to N). Whether distributor 110 broadcasts instruction 130 to which processor core 120 or which some processor cores 120 may be preset or set based on received configuration information 210.

[0034] Alternatively, in some embodiments, distributor 110 can also distribute instruction 130 to each processor core 120 using other transmission methods. For example, distributor 110 can first send instruction 130 to processor core 120-1 and then send instruction 130 to processor core 120-2, and so on by analogy. Compared with such a sequential distribution method, by distributing instructions using the broadcast method, the overhead of repeatedly reading instructions can be reduced, and thus the overhead of instruction transmission can be significantly reduced.

[0035] In the above example of broadcasting the same instruction to each processor core, distributor 110 can send different data to be processed associated with the instruction to different processor cores 120 respectively. For example, distributor 110 can broadcast or send the first data among data 140 to the first processor core 120-1 for processing, and also broadcast or send the second data different from the first data among data 140 to the second processor core 120-2 for processing.

[0036] Such an arrangement is beneficial. For example, in many computing scenarios such as neural network inference, there are numerous scenarios where the same instruction is used to compute different data. In such scenarios, by broadcasting the same instruction and different data to different processor cores 120 in a broadcast manner, the overhead of instruction and data transmission can be significantly reduced, and thus the data processing efficiency can be improved.

[0037] Although FIG. 3 only shows an example of broadcasting the same instruction 130 to each processor core 120, it should be understood that alternatively or additionally, in some embodiments, the same process as in FIG. 3 can also be used to broadcast the same data to each processor core 120. For example, data 140 can be broadcast to at least one of the plurality of processor cores 120. In such an example, different instructions associated with the data 140 can be distributed to different processor cores 120. For example, by sending or broadcasting the first instruction among the instructions 130 to the processor core 120-1, the processor core 120 can process the data 130 based on the first instruction, and by sending or broadcasting the second instruction among the instructions 130 to the processor core 120-2, the processor core 120-2 can process the data 140 based on the second instruction.

[0038] A method of distributing such the same data and different instructions to each processor core 120 is applicable to many data processing scenarios, such as specific data processing processes with a small amount of data but complex processing processes. For example, in neural network computing, there are often scenarios where different calculation processes are used for the same data. The above-described method of distributing the same data and different instructions to each processor core 120 can be well applied to such scenarios. Thereby, the overhead of such data and instruction transmission processes can be significantly reduced, and thus the data processing efficiency can be improved.

[0039] Continuing to refer to FIG. 2, each processor core 120 is configured to execute instructions according to the associated execution pipelines 240-1, 240-2, …, execution pipeline 240-N. For ease of explanation, hereinafter, the execution pipelines 240-1, execution pipeline 240-2, …, execution pipeline 240-N are collectively or individually referred to as the execution pipeline 240. The execution pipeline 240 is configured to process the data in the data cache 230 by using the instructions in the instruction cache 220. For example, the execution pipeline 240 can process the data written in the data cache 230 according to the instructions in the instruction cache 220 and transmit the processing result to the data cache 230. Alternatively or additionally, each processor core 120 can transmit the processing result from the data cache 230 to the distributor 110.

[0040] In some embodiments, the distributor 110 can receive the processing results from the corresponding data cache 230 of at least one processor core 120 respectively. Additionally, the distributor 110 can send the processing results of the received data to be processed to other devices such as the originating device of the data processing request. Thereby, the distributor 110 can be responsible for the exchange between external data and the data in the data cache 230 of the processor core 120, and thereby reduce the data exchange between the processor core 120 and the outside.

[0041] In some embodiments, the distributor 110 can distribute the instructions 130 and the data 140 associated with the instructions to each processor core 120. Each processor core 120 can process the received data 140 according to the instructions 130 respectively and send the processing results to the distributor 110.

[0042] In some embodiments, the distributor 110 can adopt a periodic approach to read and write data to the data cache 230. For example, the distributor 110 can distribute the third data among the data to be processed (such as data 140) to the first processor core 120-1 in at least one processor core 120 for processing. In response to receiving the first result obtained by processing the third data from the processor core 120-1, the distributor 110 distributes the fourth data different from the third data among the data to be processed to the first processor core 120-1.

[0043] Figure 4 shows a schematic diagram of the periodic writing of data, execution of instructions, and reading of data by the processor core 120 according to some embodiments of the present invention. As shown in Figure 4, the distributor 110 transmits data 410-1 to the processor core 120. The processor core 120 writes the data 410-1 into the data cache 230 by performing a data write 430-1 on the received data 410-1. The processor core 120 performs an instruction execution 440-1 using the execution pipeline 240 shown in Figure 2 and the like according to the instructions associated with the received data 410-1. The processor core 120 further performs a data read 450-1 from the data cache 230 for the result obtained by the instruction execution 440-1. The processor core 120 transmits the read data 420-1 to the distributor 110.

[0044] In response to receiving the data 420-1 from the processor core 120, the distributor 110 transmits data 410-2 to the processor core 120. The processor core 120 further performs processes such as a data write 430-2, an instruction execution 440-2, and a data read 450-2 on the data 410-2, and transmits the data 420-2 of the processing result corresponding to the read data 410-2 to the distributor 110.

[0045] Similarly, in response to receiving the previous data processing result from the processor core 120, the distributor 110 can send data 410-K (where K is an integer greater than 1) to the processor core 120. The processor core 120 further performs processes such as data writing 430-K, instruction execution 440-K, and data reading 450-K on the data 410-K, and sends data 420-K of the processing result corresponding to the read data 410-K to the distributor 110. The instructions associated with the data 410-1, 410-2, …, 410-K may be the same instruction, and the instruction can be transmitted to the processor core 120 by the distributor 110 only once. In some embodiments, the number of times K of periodically performing data reading / writing and data processing may be preset. Alternatively or additionally, the number of times K of periodically performing data reading / writing and data processing may be set based on the configuration information 210 received by the distributor 110.

[0046] FIG. 4 shows only the process of periodically performing data reading / writing and data processing of one of the processor cores 120 therein, but it should be understood that data reading / writing and data processing can be periodically performed on other processor cores 120 according to the same process. Such a periodic approach facilitates the large-scale loading of data into the data cache 230 of the processor core 120. By adopting such a mode, the instruction is distributed only once, and the data associated with the instruction can further reduce the instruction and data transmission overhead by adopting a periodic reading / writing and processing process.

[0047] It should be understood that unless otherwise specified, the above processes such as each data distribution and instruction distribution can be executed in any appropriate order. The above embodiments of each data distribution, instruction distribution and the embodiments of each data reading / writing and processing can be implemented in combination.

[0048] As described above, various embodiments of distributing instructions and / or data to each processor core 120 using the distributor 110 in conjunction with FIGS. 2-4 have been described. By adopting the embodiments of the present invention, on the one hand, the distributor distributes data to the data caches of each processor core, and each data cache of the processor core directly transmits data with the distributor, so that each processor core can use a large-capacity data cache, and the data exchange between the processor core and the outside can be reduced as much as possible.

[0049] On the other hand, the embodiments of the present solution use a centralized data scheduling or distribution mechanism, and it is easy to apply the broadcast method to transmit data and / or instructions, and the transmission efficiency of data and instructions can be further improved. Thereby, the present solution can make full use of limited bandwidth resources and further improve the efficiency of operations such as vector calculation.

[0050] In the case of calculations such as neural network training and / or inference, the required bandwidth of the underlying computing unit is often several times or even dozens of times the bandwidth that can be provided by the outside. The solution of the present invention can make full use of limited bandwidth resources, thereby improving the computing efficiency of computing units such as neural network accelerators.

[0051] FIG. 5 shows a flowchart of a process 500 for data processing according to some embodiments of the present invention. The process 500 may be implemented in the distributor 110 of the processor 101. For ease of explanation, the process 500 will be described with reference to the environment 100 of FIG. 1.

[0052] In block 510, the distributor 110 distributes the data to be processed (such as data 140) to the corresponding data cache 230 of at least one of the plurality of processor cores 120. For example, the distributor 110 can distribute the data 140 to each of the plurality of processor cores 120 or one or more of the plurality of processor cores 120. In block 520, the distributor 110 distributes the instructions 130 associated with the data to be processed (such as data 140) to the corresponding instruction cache 220 of at least one of the processor cores 120 for execution. For example, the distributor 110 can distribute the instructions 130 to each of the plurality of processor cores 120 or one or more of the plurality of processor cores 120.

[0053] In some embodiments, the distributor 110 can distribute the instructions 130 to at least one of the processor cores 120 by broadcasting the instructions 130 to at least one of the processor cores 120. In such an example, the distributor 110 can send the first data among the data 140 to the first processor core 120-1 of at least one of the processor cores 120 for processing, and send the second data different from the first data among the data 140 to the second processor core 120-2 of at least one of the processor cores 120 for processing.

[0054] Additionally or alternatively, in some embodiments, the distributor 110 can distribute the data 140 to at least one processor core 120 by broadcasting the data 140 to the at least one processor core 120. In such an example, the distributor 110 transmits the first instruction among the instructions 130 to the first processor core 120-1 among the at least one processor core 120, so that the first processor core 120-1 can process the data 140 based on the first instruction. The distributor 110 transmits a second instruction different from the first instruction among the instructions 130 to the second processor core 120-2 among the at least one processor core 120, so that the second processor core 120-2 can process the data 140 based on the second instruction.

[0055] In some embodiments, at block 530, the distributor 110 is further configured to receive the processing results from the corresponding data caches 230 of the at least one processor core 120, respectively. The processing results are obtained by the at least one processor core 120 processing the received data 140 according to the instructions 130, respectively. For example, the distributor 110 can distribute the third data among the data 140 to the first processor core 120-1 among the at least one processor core 120 for processing. In response to receiving the first result obtained by processing the third data from the first processor core 120-1, the distributor 110 distributes the fourth data different from the third data among the data 140 to the first processor core 120-1 for processing.

[0056] In some embodiments, the distributor 110 further receives a data set and a set of instructions to be processed by the processor 101. The distributor 110 further receives configuration information. The configuration information at least indicates an association between the data to be processed (e.g., data 140) in the data set and the instruction 130 in the set of instructions. For example, the association between the data 140 and the instruction 130 indicates that the data 140 is to be processed by the instruction 130. In such an example, the distribution of the data 140 and the instruction 130 can depend at least in part on the above-described association.

[0057] Although the steps are shown in a particular order in FIG. 5, it should be understood that some or all of these steps may be performed in other orders or in parallel. For example, box 510 in FIG. 5 may be performed before or after block 520. The scope of the present invention is not limited in this regard.

[0058] FIG. 6 shows a block diagram of an electronic device 600 including a processor 101 according to one or more embodiments of the present invention. It should be understood that the electronic device 600 shown in FIG. 6 is merely exemplary and should not limit the functions and scope of the embodiments described herein.

[0059] As shown in FIG. 6, the electronic device 600 is in the form of a general-purpose electronic device or a computing device. The components of the electronic device 600 may include, but are not limited to, one or more processors 101, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. In some embodiments, the processor 101 can execute various processes based on a program stored in the memory 620. Each processor core 120 of the processor 101 improves the parallel processing ability of the electronic device 600 by executing computer-executable instructions in parallel.

[0060] The electronic device 600 typically includes multiple computer storage media. Such media may be any accessible media that can be accessed by the electronic device 600, including but not limited to volatile media and non-volatile media, removable media and non-removable media. The memory 620 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or a specific combination thereof. The storage device 630 may be removable or non-removable media and may include machine-readable media such as flash memory drives, magnetic disks, or any other media, and can be used to store information and / or data (e.g., training data for training) and can be accessible within the electronic device 600.

[0061] The electronic device 600 can further include another removable / non-removable, volatile / non-volatile storage medium. Although not shown in FIG. 6, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive may be connected to a path (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules, and these program modules are configured to execute various methods or operations of various embodiments of the present invention. For example, these program modules may be configured to implement various functions or operations of the distributor 110.

[0062] The communication unit 640 implements communication with other computing devices via a communication medium. Additionally, the functions of the components of the electronic device 600 may be implemented as a single computing cluster or multiple computing machines, which can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or other network nodes.

[0063] The input device 650 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 660 may be one or more output devices such as a display, speaker, printer, etc. The electronic device 600 can further communicate with one or more external devices (not shown) such as a storage device, display device, etc. via the communication unit 640 as needed, communicate with one or more devices that enable a user to interact with the electronic device 600, or the electronic device 600 can perform communication with one or more other electronic devices or any device for communication of computing devices (network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0064] According to an exemplary implementation of the present invention, there is provided a computer-readable storage medium storing one or more computer instructions, and the one or more computer instructions are executed by a processor to implement the above method. According to an exemplary implementation of the present invention, there is further provided a computer program product, where the computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions executed by a processor to implement the above method.

[0065] Here, each aspect of the present invention has been described with reference to the flowcharts and / or block diagrams of the method, apparatus (system), and computer program product implemented by the present invention. It should be understood that each box in the flowchart and / or block diagram and combinations of boxes in the flowchart and / or block diagram can all be implemented by computer-readable program instructions.

[0066] These computer-readable program instructions are provided to the processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to generate a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, an apparatus for implementing the functions / operations specified in one or more boxes in the flowchart and / or block diagram can be generated. These computer-readable program instructions may be stored in a computer-readable storage medium, and by operating the computer, programmable data processing apparatus, and / or other devices in a specific form, the computer-readable medium in which the instructions are stored constitutes a manufactured product including instructions for implementing each aspect of the functions / operations specified in one or more boxes in the flowchart and / or block diagram.

[0067] By loading the computer-readable program instructions into a computer, other programmable data processing apparatus, or other devices, a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to generate a process implemented by the computer, whereby the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / operations specified in one or more boxes in the flowchart and / or block diagram.

[0068] According to one or more embodiments of the present invention, Example 1 describes a processor that includes a plurality of processor cores. The plurality of processor cores each include a data cache for reading and writing data, and an instruction cache separated from the data cache for reading instructions. The processor further includes a distributor communicatively coupled to the plurality of processor cores. The distributor is configured to distribute data to be processed to a corresponding data cache of at least one of the plurality of processor cores, and to distribute instructions associated with the data to be processed to a corresponding instruction cache of at least one of the plurality of processor cores for execution.

[0069] According to one or more embodiments of the present invention, Example 2 includes the processor described in Example 1, wherein distributing instructions to at least one processor core includes broadcasting the instructions to at least one processor core.

[0070] According to one or more embodiments of the present invention, Example 3 includes the processor described in Example 2, wherein distributing data to be processed to at least one processor core includes sending first data of the data to be processed to a first processor core for processing, and sending second data of the data to be processed to a second processor core for processing. The second data is different from the first data.

[0071] According to one or more embodiments of the present invention, Example 4 includes the processor described in Example 1, wherein distributing data to be processed to at least one processor core includes broadcasting the data to be processed to at least one processor core.

[0072] According to one or more embodiments of the present invention, Example 5 includes the processor described based on Example 4, where distributing instructions to at least one processor core includes transmitting a first instruction to a first processor core so that the first processor core processes data to be processed according to the first instruction, and transmitting a second instruction to a second processor core so that the second processor core processes data to be processed according to the second instruction. The first instruction is different from the second instruction.

[0073] According to one or more embodiments of the present invention, Example 6 includes the processor described based on Example 1, where the distributor further receives processing results from corresponding data caches of at least one processor core respectively. The processing results are obtained by processing data to be processed received according to instructions by at least one processor core respectively.

[0074] According to one or more embodiments of the present invention, Example 7 includes the processor described based on Example 6, where distributing data to be processed to at least one processor core includes distributing third data among the data to be processed to a first processor core among at least one processor core for processing, and in response to receiving a first result obtained by processing the third data from the first processor core, distributing fourth data among the data to be processed to the first processor core for processing. The third data is different from the fourth data.

[0075] According to one or more embodiments of the present invention, Example 8 includes the processor described based on Example 1, where the distributor further receives a data set and an instruction set to be processed by the processor, and receives configuration information. The configuration information at least indicates an association between data to be processed in the data set and instructions in the instruction set. The distribution of data to be processed and instructions depends at least in part on the association.

[0076] According to one or more embodiments of the present invention, Example 9 describes a data processing method. The method includes distributing data to be processed to a corresponding data cache of at least one of a plurality of processor cores of a processor by a distributor of the processor. The distributor is communicatively coupled to the plurality of processor cores. The method further includes distributing an instruction associated with the data to be processed to a corresponding instruction cache of at least one of the at least one processor core for execution.

[0077] According to one or more embodiments of the present invention, Example 10 includes the method described based on Example 9, where distributing an instruction to at least one processor core includes broadcasting the instruction to the at least one processor core.

[0078] According to one or more embodiments of the present invention, Example 11 includes the method described based on Example 10, where distributing data to be processed to at least one processor core includes distributing first data among the data to be processed to a first processor core for processing and distributing second data among the data to be processed to a second processor core for processing. The first data is different from the second data.

[0079] According to one or more embodiments of the present invention, Example 12 includes the method described based on Example 9, where distributing data to be processed to at least one processor core includes broadcasting the data to be processed to the at least one processor core.

[0080] According to one or more embodiments of the present invention, Example 13 includes the method described based on Example 12, wherein distributing instructions to at least one processor core includes transmitting a first instruction to the first processor core such that the first processor core processes data to be processed according to the first instruction, and transmitting a second instruction to the second processor core such that the second processor core processes data to be processed according to the second instruction. The first instruction is different from the second instruction.

[0081] According to one or more embodiments of the present invention, Example 14 includes the method described based on Example 9. The method further includes respectively receiving processing results from corresponding data caches of at least one processor core. The processing results are respectively obtained by processing data to be processed according to instructions by at least one processor core.

[0082] According to one or more embodiments of the present invention, Example 15 includes the method described based on Example 14, wherein distributing data to be processed to at least one processor core includes distributing third data among the data to be processed to the first processor core among at least one processor core for processing, and in response to receiving a first result obtained by processing the third data from the first processor core, distributing fourth data among the data to be processed to the first processor core for processing. The third data is different from the fourth data.

[0083] According to one or more embodiments of the present invention, Example 16 includes the method described based on Example 9. The method further includes receiving a data set and an instruction set to be processed by a processor, and receiving configuration information. The configuration information at least indicates an association between data to be processed in the data set and instructions in the instruction set. The distribution of the data to be processed and the instructions depends at least in part on the association.

[0084] According to one or more embodiments of the present invention, Example 17 describes an electronic device, which includes at least the processor described in any one of Examples 1 to 8.

[0085] According to one or more embodiments of the present invention, Example 18 describes a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method described in any one of Examples 9 to 16 is implemented.

[0086] Flowcharts and block diagrams in the drawings show the implementable architectures, functions, and operations of multiple implementable systems, methods, and computer program products according to the present invention. In this regard, each box in the flowchart or block diagram can represent one module, program segment, or part of an instruction, and the module, program segment, or part of an instruction includes one or more executable instructions for implementing the specified logical function. In some implementations as an alternative, the functions represented in the boxes may occur in an order different from that shown in the drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or depending on the functions involved, may be executed in the reverse order. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a special-purpose hardware-based system for executing the specified function or operation, or may be implemented by a combination of special-purpose hardware and computer instructions.

[0087] The above describes each implementation of the present invention. However, the above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The choice of terms used in this specification is intended to best interpret the principles of each implementation, actual applications, or improvements to the technology in the market, or to enable those skilled in the art to understand each implementation form disclosed in this specification.

Claims

1. A plurality of processor cores each comprising a data cache for reading and writing data, and an instruction cache separated from the data cache for reading instructions, a distributor communicatively coupled to the plurality of processor cores, configured to distribute data to be processed to a corresponding data cache of at least one of the plurality of processor cores, and to distribute instructions associated with the data to be processed to a corresponding instruction cache of the at least one processor core for execution, a processor.

2. Distributing the instructions to the at least one processor core includes broadcasting the instructions to the at least one processor core, The processor according to claim 1.

3. Distributing the data to be processed to the at least one processor core includes sending first data of the data to be processed to a first processor core for processing, sending second data different from the first data of the data to be processed to a second processor core for processing, The processor according to claim 2.

4. Distributing the data to be processed to the at least one processor core includes broadcasting the data to be processed to the at least one processor core, The processor according to claim 1.

5. Distributing the instructions to the at least one processor core includes sending a first instruction to a first processor core such that the first processor core processes the data to be processed based on the first instruction, sending a second instruction different from the first instruction to a second processor core such that the second processor core processes the data to be processed based on the second instruction, The processor according to claim 4.

6. The distributor is configured to receive processing results from corresponding data caches of the at least one processor core, the processing results being obtained by the at least one processor core processing the data to be processed received based on the instructions, The processor according to claim 1.

7. Distributing the data to be processed to the at least one processor core includes distributing third data among the data to be processed to a first processor core among the at least one processor core for processing, and in response to receiving a first result obtained by processing the third data from the first processor core, distributing fourth data different from the third data among the data to be processed to the first processor core for processing. The processor according to claim 6.

8. The distributor is configured to receive a data set and an instruction set processed by the processor, and configured to receive configuration information indicating at least an association between the data to be processed in the data set and the instructions in the instruction set, wherein the distribution of the data to be processed and the instructions is arranged to depend at least in part on the association. The processor according to claim 1.

9. Distributing, by a distributor of a processor, data to be processed to a corresponding data cache of at least one processor core among a plurality of processor cores of the processor, the distributor being communicatively coupled to the plurality of processor cores, and distributing an instruction associated with the data to be processed to a corresponding instruction cache of the at least one processor core for execution. A data processing method.

10. Distributing the instructions to the at least one processor core includes broadcasting the instructions to the at least one processor core. The data processing method according to claim 9.

11. Distributing the data to be processed to the at least one processor core includes distributing first data among the data to be processed to a first processor core for processing, and distributing second data different from the first data among the data to be processed to a second processor core for processing. The data processing method according to claim 10.

12. Distributing the data to be processed to the at least one processor core includes broadcasting the data to be processed to the at least one processor core. The data processing method according to claim 9.

13. Allocating the command to the at least one processor core includes sending a first command to a first processor core so that the first processor core processes the data to be processed based on the first command; and sending a second command different from the first command to a second processor core so that the second processor core processes the data to be processed based on the second command. The data processing method according to claim 12.

14. Receiving a processing result from the corresponding data cache of the at least one processor core, the processing result being further obtained by the at least one processor core processing the data to be processed received based on the command. The data processing method according to claim 9.

15. Allocating the data to be processed to the at least one processor core includes allocating third data among the data to be processed to a first processor core among the at least one processor core for processing; and in response to receiving a first result obtained by processing the third data from the first processor core, allocating fourth data different from the third data among the data to be processed to the first processor core for processing. The data processing method according to claim 14.

16. Further including receiving a data set and an instruction set to be processed by the processor, and receiving configuration information at least indicating an association between the data to be processed in the data set and the instruction in the instruction set, wherein the allocation of the data to be processed and the instruction depends at least in part on the association. The data processing method according to claim 9.

17. An electronic device comprising at least the processor according to any one of claims 1 to 8. An electronic device.

18. A computer program is stored, and the computer program is executed by a processor to implement the method according to any one of claims 9 to 16. A computer-readable storage medium.

Citation Information

Patent Citations

  • Apparatus, method, and system for improving power performance efficiency by combining a first core type and a second core type.

    JP2013532331A

  • Method and processor for data processing

    JP2017509985A

  • Thermal mitigation for multi-core processors

    JP2018501546A

  • Mechanism for partitioning shared local memory

    JP2021099786A

  • Executing multiple programs simultaneously on a processor core

    US20180225124A1