Task scheduling method for parallel processor chip, computing device, computer readable storage medium and computer program product

By forming the effective processors in the parallel processor chip into logical processor clusters and scheduling the calculation task with the logical processor cluster as granularity, the problem of low yield of parallel processor chips due to manufacturing defects is solved, and the effect of improving chip yield and reducing costs is achieved.

CN118277101BActive Publication Date: 2025-05-06北京壁仞科技开发有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410658287.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-05-06
Estimated Expiration
2044-05-24

AI Technical Summary

Technical Problem

Due to the large number of processors, parallel processor chips are easily affected by manufacturing defects, resulting in low chip yield and excessive cost of a single intact chip.

Method used

The yield of the chip is improved by forming the effective processor in the parallel processor chip into a logical processor cluster and scheduling the computing task with the logical processor cluster as a granularity.

Benefits of technology

Without changing the programming model, the yield of the chip is improved, and the dynamically divided logic processor clusters make more processors available, reducing the cost of a single chip.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118277101B_ABST
    Figure CN118277101B_ABST
Patent Text Reader

Abstract

The present disclosure provides a task scheduling method for a parallel processor chip, a computing device, a computer-readable storage medium, and a computer program product. The method includes: determining the validity status of each processor in a plurality of processors of the parallel processor chip; dividing the valid processors in the plurality of processors into a plurality of logical processor clusters based on the validity status of each processor; and assigning a computing task to at least one logical processor cluster in the plurality of logical processor clusters so that the computing task is routed to a processor corresponding to the at least one logical processor cluster based on a mapping relationship between the plurality of logical processor clusters and the plurality of processors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of processors, and more specifically, to a task scheduling method for a parallel processor chip, a computing device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the advancement of integrated circuit manufacturing technology, parallel processor chips that integrate multiple processors (also called processing cores) are currently widely used to accelerate the processing of various computing tasks. The number of processors that can be integrated on a single parallel processor chip can range from a few to dozens, or even more. Parallel processor chips that integrate dozens or even more processors on a single chip are also called multi-core parallel processor chips. The processors in a parallel processor chip all use the same design and are huge in number, so they can be used to accelerate the operation of parallel computing tasks.

[0003] In the chip manufacturing process, manufacturing defects are inevitably introduced, causing circuit failure. Due to the large number of processors in a parallel processor chip, there is a greater probability of failure due to manufacturing defects. This defect is more prominent for multi-core parallel processor chips. If left untreated, defective processors will cause the entire chip to fail, resulting in a low chip yield and a high cost for a single intact chip. Summary of the invention

[0004] In response to the above problems, the present disclosure provides a solution for task scheduling by organizing the effective processors in a parallel processor chip into logical processor clusters, thereby enabling computing task scheduling with processor clusters as the granularity without changing the programming model, thereby improving the chip yield.

[0005] According to one aspect of the present disclosure, a task scheduling method for a parallel processor chip is provided. The method includes: determining the validity status of each processor among a plurality of processors of the parallel processor chip; dividing the valid processors among the plurality of processors into a plurality of logical processor clusters based on the validity status of each processor; and assigning a computing task to at least one logical processor cluster among the plurality of logical processor clusters so that the computing task is routed to a processor corresponding to the at least one logical processor cluster based on a mapping relationship between the plurality of logical processor clusters and the plurality of processors.

[0006] In some implementations, determining the validity status of each processor of the plurality of processors of the parallel processor chip includes: performing a DFT test on the plurality of processors to determine whether each processor is a valid processor or an invalid processor; and setting the result of the DFT test in a status register to indicate the validity status of each processor.

[0007] In some implementations, dividing effective processors among the multiple processors into multiple logical processor clusters based on the validity status of each processor includes: assigning sequentially increasing processor logical numbers to effective processors among the multiple processors based on the validity status of each processor; and dividing the effective processors into the multiple logical processor clusters based on the processor logical numbers and a predetermined number of processors per cluster.

[0008] In some implementations, allocating a computing task to at least one of the multiple logical processor clusters includes: allocating the at least one logical processor cluster to the computing task based on a requirement of the computing task for parallel processing capability; and generating a computing task barrier signal corresponding to the at least one logical processor cluster based on a mapping relationship between the multiple logical processor clusters and the multiple processors, wherein the computing task barrier signal indicates the processor corresponding to the at least one logical processor cluster.

[0009] In some implementations, the method further includes: routing the computing task to the corresponding processor through a plurality of internal interconnections of the parallel processor chip according to the computing task barrier signal.

[0010] In some implementations, the multiple internal interconnections include multiple physical processor clusters respectively connected to the multiple processors, and routing the computing task to the corresponding processor through the multiple internal interconnections of the parallel processor chip according to the computing task barrier signal includes: a first internal interconnection obtains the lowest N bits of the computing task barrier signal, and determines whether the computing task is assigned to one or more processors in the physical processor cluster corresponding to the first internal interconnection, where N is the number of processors included in the physical processor cluster; if it is determined that the computing task is assigned to one or more processors in the physical processor cluster corresponding to the first internal interconnection, routing the computing task to the one or more processors; obtaining the lowest N bits of the computing task barrier signal after being right-shifted by N bits in sequence, and passing them to other internal interconnections after the first internal interconnection to determine whether the computing task is assigned to one or more processors in the physical processor cluster corresponding to the other internal interconnections.

[0011] According to another aspect of the present disclosure, a computing device is provided, comprising: at least one processor; and at least one memory, wherein the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, wherein the instructions, when executed by the at least one processor, cause the computing device to perform the steps of the method described above.

[0012] According to yet another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program code is stored. When the computer program code is executed, the method described above is executed.

[0013] According to yet another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program performs the method described above when executed by a machine. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The present disclosure will be better understood and other objects, details, features and advantages of the present disclosure will become more apparent through the description of specific embodiments of the present disclosure given with reference to the following drawings.

[0015] Figure 1 A schematic diagram of the structure of an exemplary parallel processor chip is shown.

[0016] Figure 2 An exemplary flow chart of a task scheduling method for a parallel processor chip according to an embodiment of the present invention is shown.

[0017] Figure 3 The embodiment of the present invention shows Figure 1 Schematic diagram of a logical processor cluster obtained by dividing multiple processors.

[0018] Figure 4 An exemplary flow chart of a computing task routing process according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0019] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0020] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "one embodiment" and "some embodiments" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0021] Figure 1 FIG. 1 shows a schematic diagram of the structure of an exemplary parallel processor chip 100. Figure 1As shown in FIG, the parallel processor chip 100 may include a plurality of processors 110 and a task scheduler 120 for performing task scheduling for the plurality of processors 110. According to the current programming model, the task scheduler 120 performs computing task scheduling on the processors 110 in the form of a processor cluster. That is, a plurality of processors 110 form a processor cluster, and the task scheduler 120 assigns a computing task to the processor cluster. Here, in order to distinguish it from the virtual processor cluster described below, this fixed-form-divided processor cluster is also referred to as a physical processor cluster, such as Figure 1 As shown in the reference numeral 130. For example, Figure 1 16 processors 110 (i.e., processor 110-0, processor 110-1, ..., processor 110-15) are shown by way of example, wherein every four processors 110 constitute a physical processor cluster 130 (e.g., processors 110-0 to 110-3 constitute physical processor cluster 130-0, processors 110-4 to 110-7 constitute physical processor cluster 130-1, processors 110-8 to 110-11 constitute physical processor cluster 130-2, and processors 110-12 to 110-15 constitute physical processor cluster 130-3). Those skilled in the art can understand that Figure 1 The number of processors 110 included in the parallel processor chip 100, the number of physical processor clusters 130, and the number of processors 110 included in each physical processor cluster 130 (as well as the number of internal interconnects 140 described below) are merely exemplary, and they can be designed and implemented differently for different system requirements. In addition, in some implementations, when the number of processors 110 included in the parallel processor chip 100 has been determined, the number of physical processor clusters 130 included and the number of processors 110 in each physical processor cluster 130 (i.e., the number of processors per cluster described below) can be flexibly configured in the programming model according to different applications.

[0022] In some instances, the parallel processor chip 100 may be an AI chip, a GPU (Graphic Processing Unit), a general-purpose GPU (GPGPU), etc., wherein the processor 110 may be a general-purpose computing unit CU (Computing Unit), and the physical processor cluster 130 may also be referred to as a streaming processor cluster (Streaming Processing Cluster, SPC).

[0023] The task scheduler 120 can be connected through multiple internal interconnects 140 (such as Figure 1The internal interconnects 140 - 0 , 140 - 1 , 140 - 2 , and 140 - 3 shown in FIG. 1 are directly connected to each processor 110 in the plurality of physical processor clusters 130 , respectively, to uniformly schedule and manage these processors 110 .

[0024] In some implementations, the task scheduler 120 may be connected to each internal interconnect 140 respectively, and when scheduling a task, the task scheduler 120 may only send a task scheduling signal to the internal interconnect 140 corresponding to the physical processor cluster 130 as the task scheduling target, and the internal interconnect 140 may route the computing task assigned by the task scheduling signal to all processors 110 of the physical processor cluster 130. For example, when the task scheduler 120 is to schedule a task to the physical processor cluster 130-1, the task scheduler 120 may only send a task scheduling signal to the internal interconnect 140-1, and the internal interconnect 140-1 may route the computing task to all its processors 110-4 to 110-7.

[0025] In some other implementations, multiple internal interconnects 140 may communicate with each other. In this case, the task scheduler 120 may be connected to only one internal interconnect 140, and when scheduling a task, it may only send a task scheduling signal to the internal interconnect 140 connected to it, and the internal interconnect 140 may forward the task scheduling signal to another internal interconnect 140 corresponding to the physical processor cluster 130 as the task scheduling target. For example, Figure 1 As shown in , it is assumed that the task scheduler 120 is only connected to the first internal interconnect 140-0, and the first internal interconnect 140-0 is connected to the internal interconnect 140-1, the internal interconnect 140-2 and the internal interconnect 140-3 in sequence. In this case, the task scheduler 120 can directly send a task scheduling signal to the internal interconnect 140-0, and the internal interconnect 140-0 determines whether the physical processor cluster 130-0 connected thereto is the target of the task scheduling signal by parsing the task scheduling signal. If the physical processor cluster 130-0 is the target of the task scheduling signal, the internal interconnect 140-0 can route the computing task assigned by the task scheduling signal to all processors 110-0 to 110-3 in the physical processor cluster 130-0. If the physical processor cluster 130-0 is not the target of the task scheduling signal, the internal interconnect 140-0 can route the task scheduling signal to the next internal interconnect 140-1, and the next internal interconnect 140-1 performs operations similar to the internal interconnect 140-0. And so on, until the task scheduling signal reaches the physical processor cluster 130 that is its target.

[0026] In this case, if one or more processors 110 fail during the manufacturing process of the parallel processor chip 100 (this can be determined by testing), then since the task scheduler 120 schedules computing tasks at the granularity of physical processor clusters, all physical processor clusters 130 containing failed processors will not be able to operate normally, thereby greatly reducing the number of processors actually available.

[0027] In view of this situation, a hardware fault-tolerant mechanism called a harvest scheme is currently generally used to enable the entire chip to work normally even when a certain number of failed processors are included, so as to improve the yield of the parallel processor chip and reduce the manufacturing cost. In the related art, there are two harvest schemes for parallel processor chips. One harvest scheme is to directly shield the entire physical processor cluster containing the failed processor when the task scheduler 120 performs computing task scheduling, and only assign computing tasks to the physical processor cluster in which all processors are intact (valid). However, for this scheme, if most or even all physical processor clusters contain failed processors (even if only one failed processor is included), there will be very few or even no physical processor clusters available for scheduling by the task scheduler 120, so that the entire parallel processor chip still cannot work normally, and the improvement of the chip yield is very limited.

[0028] Another harvesting scheme is to configure redundant processors 110 in each physical processor cluster 130, so that the number of valid processors in each physical processor cluster 130 can still meet the number of processors in the cluster defined by the programming model. For example, assuming that 4 processors 110 need to be included in each physical processor cluster 130, 1 or 2 processors 110 can be additionally configured for each physical processor cluster 130 as redundant processors. When 1 or 2 of the 4 processors 110 are failed processors, the additionally configured 1 or 2 processors 110 can be used to replace the failed processor to form a physical processor cluster 130 containing 4 valid processors. However, in this scheme, the additionally configured redundant processors will occupy additional chip area, resulting in a waste of resources, and the number of redundant processors is also difficult to determine. Too many redundant processors will further increase the chip area, and too few redundant processors may still not guarantee the availability of the entire physical processor cluster 130.

[0029] In response to the above problems, the present disclosure provides a task scheduling solution for a parallel processor chip, wherein all effective processors in the chip are divided into multiple logical processor clusters as a processor pool, and the task scheduler 120 allocates computing tasks at the granularity of logical processor clusters, thereby improving the chip yield without changing the programming model.

[0030] Figure 2 FIG. 2 is an exemplary flow chart of a task scheduling method 200 for a parallel processor chip 100 according to an embodiment of the present invention. The task scheduling method 200 may be executed by the parallel processor chip 100 , more specifically, by the task scheduler 120 .

[0031] like Figure 2 As shown in FIG. 2 , at block 210 , the task scheduler 120 may determine the validity status of each processor 110 of the parallel processor chip 100 , ie, a valid processor or an invalid processor.

[0032] As mentioned above, processor failures caused during the manufacturing process of the parallel processor chip 100 can be determined by testing means. Currently, commonly used testing means include, for example, DFT testing (Design for Test), that is, by inserting various hardware logics for improving the testability of the chip (including controllability and observability) into the original design of the chip, and by initiating the test of the chip and receiving the test results through special DFT testing software, the chip becomes easy to test. DFT testing can be performed at the end of the chip manufacturing process and before packaging. However, the present invention is not limited to this, and even after the chip is packaged and delivered, DFT testing or other similar tests can still be performed on it to determine whether each processor 110 is a valid processor.

[0033] In this case, block 210 may specifically include performing a DFT test on the plurality of processors 110 of the parallel processor chip 100 to determine whether each processor 110 is a valid processor or an invalid processor, and setting the result of the DFT test in a status register to indicate the validity status of each processor 110 .

[0034] Here, the task scheduler 120 may be configured with a status register harvest_reg dedicated to the validity status of the plurality of processors 110, the size of the status register harvest_reg being equal to the number of processors 110 in the parallel processor chip 100, and each bit indicating the validity status of the corresponding processor 110. Figure 1 In the case where 16 processors 110 are included as shown, the status register harvest_reg may include 16 bits, represented as harvest_reg[15:0], wherein the lowest bit corresponds to processor 110-0, and the highest bit corresponds to processor 110-15. A bit value of 0 for a certain bit indicates that the corresponding processor 110 is an invalid processor, and a bit value of 1 indicates that the processor 110 is a valid processor.

[0035] At block 220 , the task scheduler 120 may partition the active processors in the plurality of processors 110 into a plurality of logical processor clusters based on the availability status of each processor.

[0036] Figure 3 The embodiment of the present invention shows Figure 1 Schematic diagram of a logical processor cluster 130' obtained by dividing multiple processors 110. Figure 3 and Figure 1 The only difference is that the physical processor cluster 130 is replaced by the newly divided logical processor cluster 130'. In addition, in order to indicate the validity status of each processor 110, the invalid processor 110 is shown in a dotted frame. Figure 3 In the example shown, assuming that processors 110 - 0 , 110 - 1 , 110 - 8 , and 110 - 12 are invalid processors, the state register harvest_reg[15:0]=1110 1110 1111 1100.

[0037] More specifically, in box 220, dividing the effective processors among the multiple processors 110 into multiple logical processor clusters 130' can specifically include: assigning sequentially increasing processor logical numbers to the effective processors among the multiple processors 110 based on the validity status of each processor 110, and dividing the effective processors into multiple logical processor clusters 130' based on the processor logical numbers and a predetermined number of processors per cluster.

[0038] For example, the task scheduler 120 may traverse the state register harvest_reg in order from low to high according to the processor physical numbers (i.e., the physical numbers 110-0, 110-1, ... 110-15 of the multiple processors 110), and assign the processor logical number 0 to the first processor 110 whose validity state is 1 (i.e., a valid processor), assign the processor logical number 1 to the second processor 110 whose validity state is 1, and so on, until the entire state register harvest_reg is traversed. Figure 3 The processor validity status shown will result in the following Table 1:

[0039] Table 1

[0040]

[0041] For a processor 110 whose validity status is 0 (ie, an invalid processor), no processor logical number may be assigned to it, or other marks may be used to indicate that it is an invalid processor.

[0042] Then, the task scheduler 120 may divide the logical processor clusters 130' according to the processor logical numbers (rather than the processor physical numbers) and the number of processors in each cluster. For example, still assuming that the number of processors in each cluster is 4, that is, each processor cluster contains 4 processors, then the valid processors in Table 1 may be divided into the following Tables 2 and 3 according to the processor logical numbers: Figure 3 The logical processor clusters 130'-0, 130'-1 and 130'-2 are shown in FIG.

[0043] Table 2

[0044]

[0045] Here, dividing the logical processor clusters in the order of increasing processor logical numbers is exemplary, and the present invention is not limited thereto, but the logical processor clusters may be divided in other ways. For example, the effective processors may be divided into logical processor clusters in an alternating or random manner, as long as the task scheduler 120 maintains a mapping relationship table between the processor logical numbers, the processor physical numbers, and the logical processor clusters as shown in Table 2 above.

[0046] Next, at block 230 , the task scheduler 120 may assign a computing task to at least one logical processor cluster 130 ′ among the plurality of logical processor clusters 130 ′ so that the computing task is routed to the processor 110 corresponding to the at least one logical processor cluster 130 ′.

[0047] In some embodiments, the task scheduler 120 may allocate at least one logical processor cluster to a computing task according to a requirement of the computing task for parallel processing capability.

[0048] In some implementations, when determining the number of processors per cluster for an application in the programming model, the requirements of each computing task for parallel processing capability are simultaneously converted into the form of processor clusters, that is, for each computing task, the task scheduler 120 can know how many processor clusters need to be configured for it through the demand information of the computing task when scheduling it. Alternatively, the task scheduler 120 can also convert the computing task into the processor cluster level according to the requirements of the parallel processing capability of a smaller granularity of the computing task (such as the requirements of the processor level), and the present invention does not limit this.

[0049] Of course, other factors need to be considered when allocating a computing task to which logical processor cluster or clusters, such as the availability of the logical processor cluster (for example, whether it is idle), etc. The specific allocation strategy can be determined by polling, randomly determined from the available processor clusters, or determined from the available processor clusters according to a certain rule (such as the processor cluster with the smallest workload), etc. This article does not impose any restrictions on this.

[0050] Then, the task scheduler 120 can generate a computing task barrier signal corresponding to the determined at least one logical processor cluster 130' (such as logical processor cluster 130'-0) based on the mapping relationship between multiple logical processor clusters 130 (such as logical processor clusters 130'-0 to 130'-2) and multiple processors 110 (such as processors 110-0 to 110-15) of the parallel processor chip 100, wherein the computing task barrier signal indicates the processor 110 corresponding to the at least one logical processor cluster 130'.

[0051] For example, assuming that the logical processor cluster assigned by the task scheduler 120 to the computing task is the logical processor cluster 130'-0, based on the mapping relationship between the logical processor cluster 130' and the processor 110 as shown in Table 2 above, the processor physical numbers corresponding to the logical processor cluster 130'-0 are determined to be processor 110-2, processor 110-3, processor 110-4, and processor 110-5. In this case, the computing task barrier signal generated by the task scheduler 120 can be a 16-bit bit string 0000 0000 0011 1100, wherein the bits with a bit value of 1 correspond to processor 110-2, processor 110-3, processor 110-4, and processor 110-5 (in order from low to high).

[0052] In this way, the effective processors of the parallel processor chip can be dynamically divided into logical processor clusters, so that computing tasks of the chip can be scheduled at the granularity of logical processor clusters without changing the programming model.

[0053] Furthermore, in some embodiments, the method 200 may also include routing the computing task to the corresponding processor 110 through a plurality of internal interconnections 140 of the parallel processor chip 100 according to the computing task barrier signal.

[0054] Figure 4 FIG. 2 shows an exemplary flow chart of a computing task routing process 240 according to an embodiment of the present invention. Figure 1 and Figure 3 As shown, the parallel processor chip 100 includes a plurality of internal interconnects 140 (eg, a first internal interconnect 140 - 0 , a second internal interconnect 140 - 1 , a third internal interconnect 140 - 2 , and a fourth internal interconnect 140 - 3 , wherein each internal interconnect 140 is connected to a corresponding physical processor cluster 130 ).

[0055] like Figure 4As shown in FIG. 2 , in block 241, the first internal interconnect 140-0 may obtain the lowest N bits of the computing task barrier signal, and in block 242, determine whether the computing task is assigned to one or more processors 110 in the physical processor cluster 130-0 corresponding to the first internal interconnect 140-0. Here, N is the number of processors 110 included in the physical processor cluster 130 (i.e., the number of processors per cluster), for example, Figure 1 and Figure 3 In the example, N=4.

[0056] For example, assuming that the computing task barrier signal generated by the task scheduler 120 in box 230 is a 16-bit bit string 0000 0000 0011 1100, the first internal interconnect 140-0 can obtain the lowest 4 bits 1100 of the computing task barrier signal, and determine that the computing task is assigned to two processors in the physical processor cluster 130-0 corresponding to the first internal interconnect 140-0, namely, processor 110-2 and processor 110-3, based on the bits with bit values ​​of 1 in the lowest 4 bits.

[0057] If it is determined that the computing task is assigned to one or more processors 110 (eg, processor 110 - 2 and processor 110 - 3 ) in the physical processor cluster 130 - 0 corresponding to the first internal interconnect 140 - 0 , then in block 243 , the first internal interconnect 140 - 0 may route the computing task to processor 110 - 2 and processor 110 - 3 .

[0058] Next, in block 244, the first internal interconnect 140-0 may obtain the lowest N bits of the computing task barrier signal after the computing task barrier signal is right-shifted by N bits in sequence, and transmit the bits to other internal interconnects after the first internal interconnect 140-0 (such as the second internal interconnect 140-1, the third internal interconnect 140-2, and the fourth internal interconnect 140-3). Here, regardless of whether the judgment of block 242 is "yes" or "no", the operation of block 244 is ultimately performed, that is, it is necessary to determine whether the computing task is allocated to one or more processors 110 in other physical processor clusters 130 according to the computing task barrier signal.

[0059] For example, in the case where the computing task barrier signal is a 16-bit bit string 0000 0000 0011 1100, the first internal interconnect 140-0 shifts the computing task barrier signal right by 4 bits to obtain the lowest 4 bits of 0011 (i.e., the 5th to 8th bits of the bit string 00000000 0011 1100 from low to high), and transmits the 4 bits of 0011 to the second internal interconnect 140-1; the first internal interconnect 140-0 continues to shift the computing task barrier signal right by 4 bits to obtain the lowest 4 bits of 0000 (i.e., the 9th to 12th bits of the bit string 00000000 0011 1100 from low to high), and transmits the 4 bits of 0000 to the third internal interconnect 140-2; the first internal interconnect 140-0 continues to shift the computing task barrier signal right by 4 bits to obtain the lowest 4 bits of 0000 (i.e., the 9th to 12th bits of the bit string 00000000 0011 1100 from low to high), and transmits the 4 bits of 0000 to the third internal interconnect 140-2. 1100 from low to high 13th to 16th bits), and transfers the 4 bits 0000 to the fourth internal interconnect 140-3.

[0060] Next, the other internal interconnects 140 may perform operations similar to those of the first internal interconnect 140-0, i.e., the operations described in the above blocks 242 to 244, and determine whether the computing task is assigned to one or more processors 110 in the physical processor cluster 130 corresponding to the other internal interconnect 140 based on the N-bit signal received from the first internal interconnect 140-0, and when it is determined that the computing task is assigned to the processor 110 in the corresponding physical processor cluster 130, route the computing task to the corresponding processor.

[0061] For example, the second internal interconnect 140-1 can determine, based on the received 4-bit signal 0011, that the computing task is assigned to two processors in the physical processor cluster 130-1 corresponding to the second internal interconnect 140-1, namely, processor 110-4 and processor 110-5, and route the computing task to processor 110-4 and processor 110-5; the third internal interconnect 140-2 and the fourth internal interconnect 140-3 can determine, based on the received 4-bit signal 0000, that the computing task is not assigned to the processor 110 in the physical processor clusters 130-2 and 130-3 corresponding to the third internal interconnect 140-2 and the fourth internal interconnect 140-3.

[0062] By using the solution of the present invention, all valid processors in a parallel processor chip are used as a processor pool and divided into multiple logical processor clusters, and the task scheduler allocates computing tasks based on the granularity of the logical processor cluster, so that more available processor clusters can be obtained when the number of invalid processors is the same, and the chip yield rate is improved without changing the programming model. In addition, by saving the mapping relationship between the logical processor cluster and the processor in the task scheduler, the task scheduler can realize the task allocation based on the granularity of the logical processor cluster without complex logic.

[0063] Those skilled in the art will appreciate that the parallel processor chip 100 shown in the figure is merely illustrative and may include more or fewer components.

[0064] The parallel processor chip 100 according to the present disclosure and the task scheduling method 200 executed in the parallel processor chip 100 are described above in conjunction with the accompanying drawings. However, those skilled in the art will appreciate that the parallel processor chip 100 does not necessarily include all the components shown in the drawings, and may only include some of the components or more components necessary to perform the functions described in the present disclosure, and the connection method of these components is not limited to the form shown in the drawings, and the method 200 may also include more steps not shown in the drawings.

[0065] The present invention can be implemented as a method, a computing device (such as a parallel processor chip 100 or a task scheduler 120 of a parallel processor chip 100), a computer-readable storage medium and / or a computer program product. A computer-readable storage medium stores a computer program code, which is used to execute the method of the present disclosure when it is executed. The computer program product includes a computer program, which executes the method of the present disclosure when it is executed. The computing device and / or the computing device may include at least one processor and at least one memory coupled to the at least one processor, and the memory may store instructions for execution by the at least one processor. When the instruction is executed by the at least one processor, the computing device and / or the control device may execute the above method.

[0066] In one or more exemplary designs, the functions described in the present disclosure may be implemented using hardware, software, firmware, or any combination thereof. For example, if implemented using software, the functions may be stored as one or more instructions or codes on a computer-readable medium, or transmitted as one or more instructions or codes on a computer-readable medium.

[0067] Those skilled in the art should also understand that the various illustrative logical blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments of the present disclosure may be implemented as electronic hardware, computer software, or a combination of the two.

[0068] The above description of the present disclosure is intended to enable any person of ordinary skill in the art to implement or use the present disclosure. Various modifications of the present disclosure are obvious to those of ordinary skill in the art, and the general principles defined herein may also be applied to other variations without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.

Claims

1. A task scheduling method for a parallel processor chip, comprising: determining a validity state of each of a plurality of processors of the parallel processor chip; dividing active processors of the plurality of processors into a plurality of logical processor clusters based on the availability status of each processor; as well as allocating computing tasks to at least one logical processor cluster among the plurality of logical processor clusters so that the computing tasks are routed to processors corresponding to the at least one logical processor cluster based on a mapping relationship between the plurality of logical processor clusters and the plurality of processors, The allocating the computing task to at least one logical processor cluster among the plurality of logical processor clusters comprises: Allocating the at least one logical processor cluster to the computing task according to the requirement of the computing task for parallel processing capability; and generating a computing task barrier signal corresponding to the at least one logical processor cluster based on a mapping relationship between the plurality of logical processor clusters and the plurality of processors, wherein the computing task barrier signal indicates the processor corresponding to the at least one logical processor cluster; And the method also includes: Routing the computing task to the corresponding processor through multiple internal interconnections of the parallel processor chip according to the computing task barrier signal, wherein the multiple internal interconnections are respectively connected to multiple physical processor clusters of the multiple processors, and routing the computing task to the corresponding processor through multiple internal interconnections of the parallel processor chip according to the computing task barrier signal comprises: The first internal interconnect obtains the lowest N bits of the computing task barrier signal, and determines whether the computing task is allocated to one or more processors in a physical processor cluster corresponding to the first internal interconnect, where N is the number of processors included in the physical processor cluster; If it is determined that the computing task is assigned to one or more processors in a physical processor cluster corresponding to the first internal interconnect, routing the computing task to the one or more processors; The lowest N bits of the computing task barrier signal after being right-shifted by N bits are obtained and transmitted to other internal interconnections after the first internal interconnection to determine whether the computing task is allocated to one or more processors in the physical processor cluster corresponding to the other internal interconnections.

2. The method of claim 1 , wherein determining the validity status of each processor of the plurality of processors of the parallel processor chip comprises: performing DFT testing on the plurality of processors to determine whether each processor is a valid processor or an invalid processor; as well as The results of the DFT test are set in a status register to indicate the validity status of each processor.

3. The method of claim 1 , wherein partitioning active processors of the plurality of processors into a plurality of logical processor clusters based on the effectiveness status of each processor comprises: assigning sequentially increasing processor logical numbers to valid processors of the plurality of processors based on the validity status of each processor; as well as The valid processors are divided into the plurality of logical processor clusters based on the processor logical numbers and the predetermined number of processors per cluster.

4. The method of claim 3, further comprising: A mapping relationship between the processor logical numbers of the plurality of processors, the physical numbers of the plurality of processors and the logical processor cluster is established and saved.

5. A computing device comprising: at least one processor; as well as At least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the computing device to perform the steps of the method according to any one of claims 1 to 4.

6. A computer-readable storage medium having a computer program code stored thereon, wherein the computer program code executes the method according to any one of claims 1 to 4 when executed.

7. A computer program product, comprising a computer program, which, when executed by a machine, performs the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Reliable computing with a many-core processor

    CN101278264A

  • Data processing method and device based on many-core system, electronic equipment and medium

    CN118034924A