Instruction scheduling method and device and electronic equipment

By splitting multiple threads into two groups and executing alternately, the problem of register read competition risk is solved, the instruction execution performance is improved and the compilation difficulty is reduced.

CN120123004APending Publication Date: 2025-06-10SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242944.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art tends to experience read competition risks when reading multiple operands from registers, resulting in degradation in instruction execution performance.

Method used

By splitting multiple threads into a first thread set and a second thread set, and following a specified clock cycle interval, the target thread group is alternately determined from the two, the target thread group is executed, and the operand is read from the register resource corresponding to the target thread group.

Benefits of technology

It effectively eliminates read conflicts when reading multiple operands, improves the overall performance of instruction execution, reduces the compiler's compile difficulty, and avoids limitations on the execution efficiency of functional units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123004A_ABST
    Figure CN120123004A_ABST
Patent Text Reader

Abstract

The invention discloses an instruction scheduling control method and device and electronic equipment, and the method comprises the steps: splitting a plurality of threads into a first thread set and a second thread set; the first thread set and the second thread set correspond to a first register resource and a second register resource respectively; according to a specified clock cycle interval, alternately determining a target thread group from the first thread set and the second thread set, and executing the target thread group; the target thread group comprises target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units; reading an operand of the target thread group from a register resource corresponding to the target thread group; and sending the target thread group and the corresponding operand to the corresponding functional unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, but is not limited to, the field of computer technology, and in particular, to an instruction scheduling method, apparatus, and electronic device. Background Art

[0002] An execution instruction of a central processing unit (CPU), a graphics processing unit (GPU), or other processors may include two or more operands, and a read competition risk may occur during the process of reading operands from registers. Therefore, how to eliminate the read conflicts caused by reading multiple operands and the impact on the instruction execution performance has become an urgent problem to be solved. Summary of the Invention

[0003] In view of this, the present disclosure provides at least an instruction scheduling method, apparatus, and electronic device.

[0004] The technical solution of the present disclosure is implemented as follows:

[0005] On the one hand, the present disclosure provides an instruction scheduling method, which includes:

[0006] Splitting a plurality of threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to a first register resource and a second register resource;

[0007] Determining target thread groups alternately from the first thread set and the second thread set at specified clock cycle intervals and executing the target thread groups; the target thread groups include target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units;

[0008] Reading the operands of the target thread groups from the register resources corresponding to the target thread groups;

[0009] Sending the target thread groups and their corresponding operands to the corresponding functional units.

[0010] On the other hand, the present disclosure further provides an instruction scheduling apparatus, including:

[0011] A splitting module, configured to split a plurality of threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to a first register resource and a second register resource;

[0012] A determination module, configured to alternately determine a target thread group from a first thread set and a second thread set at specified clock period intervals and execute the target thread group; the target thread group includes target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units;

[0013] A register reading module, configured to read the operands of the target thread group from the register resources corresponding to the target thread group;

[0014] A sending module, configured to send the target thread group and its corresponding operands to the corresponding functional unit.

[0015] In another aspect, the present disclosure also provides an electronic device, including: a control unit, at least one functional unit, a first register resource, and a second register resource; wherein,

[0016] The control unit:

[0017] Split multiple threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to the first register resource and the second register resource;

[0018] At specified clock period intervals, alternately determine a target thread group from the first thread set and the second thread set and execute the target thread group; the target thread group includes target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units;

[0019] Read the operands of the target thread group from the register resources corresponding to the target thread group;

[0020] Send the target thread group and its corresponding operands to the corresponding functional unit.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present disclosure. Description of the Drawings

[0022] The drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solution of the present disclosure.

[0023] Figure 1 It is a schematic diagram of splitting register resources in the related art;

[0024] Figure 2 It is a schematic diagram of splitting register resources according to the low-order data area and high-order data area of the register in the related art;

[0025] Figure 3Schematic diagram of the implementation process of an instruction scheduling method provided by the present disclosure;

[0026] Figure 4 Schematic diagram of the register resource grouping in the instruction scheduling method provided by the present disclosure;

[0027] Figure 5 Schematic diagram of the implementation process of an embodiment of the instruction scheduling method provided by the present disclosure;

[0028] Figure 6 Schematic diagram of the composition structure of an instruction scheduling device provided by the present disclosure;

[0029] Figure 7 Schematic diagram of the hardware entity of an electronic device provided by the present disclosure. Detailed implementation manners

[0030] In order to make the objectives, technical solutions, and advantages of the present disclosure clearer, the technical solutions of the present disclosure will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0031] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0032] The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs. The terms used herein are only for the purpose of describing the present disclosure and are not intended to limit the present disclosure.

[0034] In the related art, in order to solve the read conflict problem generated when reading registers with multiple operands, the following two solutions are proposed:

[0035] The first solution: First, divide the register resources into two register queues. For example, as Figure 1the first queue 110 and the second queue 120 therein; then, during the instruction compilation stage, multiple operands in the same instruction are allocated to different register queues. For example, two operands are respectively allocated to the first queue 110 and the second queue 120; finally, after the thread corresponding to the instruction is dispatched from the thread queue, the two operands in the instruction are respectively read from the two register queues. For example, as Figure 1 shown, the first warp 130 corresponding to instruction A and the second warp 140 corresponding to instruction B are respectively dispatched in two clock cycles T0 and T1. Among them, instruction A includes operand A0 and operand A1, instruction B includes operand B0 and operand B1, and operand A0 and operand B0 are stored in the first queue 110, and operand A1 and operand B1 are stored in the second queue 120. In this way, as shown in the operand read timing sequence in Table 1, in clock cycle T0, the first warp 130 reads operand A0 and A1 from the first queue 110 and the second queue 120 respectively; in clock cycle T1, the second warp 140 reads operand B0 and B1 from the first queue 110 and the second queue 120 respectively.

[0036] Clock cycle T0 T1 Issue instruction Instruction A Instruction B First queue Read operand A0 Read operand B0 Second queue Read operand A1 Read operand B1

[0037] Table 1: Operand Read Timing Sequence of Instruction A and Instruction B Implemented Based on the First Scheme

[0038] In this first scheme, it is necessary to evenly allocate multiple operands to different register queues during the compilation stage, resulting in a significant increase in the compilation difficulty, higher hardware requirements for the compiler, and an increase in the compilation duration.

[0039] The second scheme: First, the available register resources are divided into two register queues according to the low-order data area and high-order data area of the registers, as Figure 2The first queue 210 corresponding to the low-order data area and the second queue 220 corresponding to the high-order data area shown in the figure; then, the executed instructions are cut into two groups and respectively issued within two warps. For example, instruction A is cut into instruction A_low and instruction A_high, and instruction B is cut into instruction B_low and instruction B_high. Among them, instruction A_low includes operand A0_low and operand A1_low, instruction A_high includes operand A0_high and operand A1_high, instruction B_low includes operand B0_low and operand B1_low, and instruction B_high includes operand B0_high and operand B1_high; after that, after the threads corresponding to the instructions are issued from the thread queue, the clock frequency for reading operands from the register is increased. For example, two operand reading operations are performed within one clock cycle. As shown in Table 2, the clock cycle T0 is cut into register reading cycles t0 and t1. In this way, the first warp 230 corresponding to instruction A is issued in clock cycles T0 and T1, and the second warp 240 corresponding to instruction B is issued in instruction issue cycles T2 and T3. In this way, the timing of reading operands from the first queue 210 and the second queue 220 is as shown in Table 2 below:

[0040]

[0041] Table 2: Issue timing of instruction A and instruction B implemented based on the second scheme

[0042] In this second scheme, by increasing the clock frequency of operand reading to read multiple operands, the working frequency that the entire computing unit can reach is greatly limited, resulting in the performance of the entire computing unit being restricted.

[0043] Based on this, the present disclosure provides an instruction scheduling method, which can be executed by an electronic device. The electronic device can be various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), etc., or can also be implemented as a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0044] Next, in combination with the accompanying drawings of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely.

[0045] Figure 3 It is a schematic implementation flow diagram of an instruction scheduling method provided by the present disclosure, asFigure 3 As shown in Figure 3 , the method includes the following steps S301 to S304:

[0046] Step S301: Split multiple threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to a first register resource and a second register resource.

[0047] Here, a thread set can be a set including any number of threads; where a thread is the smallest unit that the operating system running on the processor can perform operation scheduling on, and each thread has its own register resource. In some embodiments, the thread set can be a set of threads in the same thread queue, that is, the first thread set and the second thread set are sets of multiple threads in different thread queues. In some embodiments, the thread set can be a set obtained by dividing multiple threads according to thread numbers.

[0048] A register refers to a group of high-speed storage units in a processor for storing data and instruction information. During the execution of an instruction, a register can be used to save the operands in the instruction, various intermediate calculation results, address information, control information, etc. Classified by function, registers can include general-purpose registers for storing temporary data and instruction operands, a program counter for storing the address of the currently executing instruction, a stack pointer for pointing to the position of the top of the stack where the thread is currently located, and so on. In some embodiments, the first register resource and the second register resource are general-purpose register resources.

[0049] Since the threads included in the first thread set and the second thread set are different, the first register resource and the second register resource are different register resources.

[0050] Step S302: Alternately determine a target thread group from the first thread set and the second thread set at a specified clock cycle interval and execute the target thread group; the target thread group includes target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units.

[0051] Here, a clock cycle is the basic time unit for the processor to execute instructions, access memory, transfer data, etc. The length of the clock cycle determines the working frequency of the processor, that is, the number of clock cycles that can be executed per second.

[0052] The specified clock cycle interval can be any suitable clock cycle interval. For example, 1, 2, or more clock cycles can be used as the specified clock cycle interval. In implementation, if the specified clock cycle interval is 2, the target thread group can be determined from the first thread set in the first clock cycle and the target thread group can be executed. After an interval of 2 clock cycles (i.e., in the third clock cycle), a new target thread group can be determined from the second thread set and the new target thread group can be executed. After an interval of 2 clock cycles (i.e., in the fifth clock cycle), a new target thread group can be determined from the first thread set again and the new target thread group can be executed, and so on.

[0053] The target thread group includes target threads corresponding to at least two target instructions, that is, the target thread group includes at least two types of threads, and each type of thread is used to execute a target instruction. In some implementation manners, the target instruction can be a complete instruction in a computer program or a partial instruction obtained by splitting a complete instruction in the computer program. For example, when the complete instruction corresponding to the target instruction is a simple and indivisible instruction such as loading a register or storing in memory, the target instruction is the corresponding complete instruction; when the complete instruction corresponding to the target instruction is a complex and divisible instruction such as a floating-point calculation instruction or a string calculation instruction, the target instruction can be a partial instruction of the corresponding complete instruction. It can be seen that when the target instruction is a complete instruction, the target instructions in two adjacent target thread groups determined from the same thread set (for example, the first thread set or the second thread set) are different instructions; when the target instruction is a partial instruction of a complete instruction, the target instructions in two adjacent target thread groups determined from the same thread set may belong to the same complete instruction (for example, belong to the same floating-point calculation instruction).

[0054] At least two target instructions corresponding to the target thread group correspond to different functional units, that is, the functional units for executing different target instructions in the same target thread group are different. Among them, the functional unit can be any type of data processing unit in the processor for executing data processing. Different functional units refer to different types of data processing units determined after classifying the functional units according to any suitable classification method. For example, classified according to the operation types executed by multiple data processing units, the multiple data processing units can be divided into arithmetic operation units (for example, basic arithmetic operation units such as addition, subtraction, multiplication, and division), logical operation units (for example, logical operation units such as AND, OR, NOT, and XOR), and special function operation units (for example, special function operation units such as trigonometric functions, exponents, and logarithms); another example is that classified according to the data types processed by multiple data processing units, the multiple data processing units can be divided into integer operation units (for example, the integer part in the arithmetic logic unit), floating-point operation units (Float Point Unit, FPU), etc.

[0055] In some embodiments, determining the target thread group from the thread set may be to determine at least two threads as target threads from multiple threads according to the thread numbers of multiple threads in the first thread set or the second thread set; among them, the at least two target threads correspond to different target instructions.

[0056] In some embodiments, after determining the target thread and its corresponding target instruction, it is necessary to parse the target instruction so that the control program can understand and execute the operation or task corresponding to the target instruction. At this time, if the functional units of the instructions corresponding to two target threads with adjacent thread numbers in the target thread group are the same, then according to the execution order of the instructions corresponding to the two target threads, remove the target thread with the later execution order from the target thread group, and re-determine whether the removed thread can be determined as a target thread after a specified clock cycle.

[0057] In some embodiments, when a thread reads an instruction, it can determine the instruction read by the thread according to the thread number and the functional unit corresponding to the instruction to be processed, so that the functional units of the instructions corresponding to threads with adjacent thread numbers are different. In this way, when determining the target thread group from the thread set, the target thread group can be determined in a sequential manner.

[0058] In some embodiments, when a thread reads an instruction, it can also read the instruction in any other suitable way. In this way, when determining the target thread group from the thread set, the target thread group can be determined from multiple threads in a non-sequential manner according to the functional unit and thread number of the instruction corresponding to each thread.

[0059] Execute the target thread group, that is, launch the target thread group from the thread queue to start executing the target thread group. In implementation, the control unit (CU) in the processor is used to launch the target thread group from the thread queue.

[0060] Step S303, read the operands of the target thread group from the register resources corresponding to the target thread group.

[0061] Here, after starting to execute the target thread group, read the operands of the target instruction from the register resources corresponding to the target thread.

[0062] In some embodiments, the area for storing data in each register can be divided into multiple partitions. For example, the register can be divided into a high - data area for storing the high - order data of the operand and a low - data area for storing the low - order data of the operand. In this way, when reading the operands from the register resources corresponding to the target thread, the high - order data and low - order data corresponding to the operands can be read in the order from the low - data area to the high - data area or the reverse order; for another example, the register can be divided into a high - data area for storing the high - order data of the operand, a middle - data area for storing the middle - order data of the operand, and a low - data area for storing the low - order data of the operand. In this way, when reading the operands from the register resources corresponding to the target thread, the high - order data, middle - order data, and low - order data corresponding to the operands can be read in the order from the low - data area, middle - data area to the high - data area or the reverse order.

[0063] Step S304, send the target thread group and its corresponding operands to the corresponding functional unit.

[0064] Here, schedule the target instruction carried by the target thread group and its corresponding operands to the corresponding functional unit to execute the corresponding target instruction using the functional unit. For example, use the floating - point arithmetic unit to perform floating - point arithmetic, use the special - function arithmetic unit to perform special - function arithmetic, etc.

[0065] In the instruction scheduling method provided by the present disclosure, first, a plurality of threads are split into a first thread set and a second thread set, and the first thread set and the second thread set respectively correspond to a first register resource and a second register resource; then, at specified clock cycle intervals, target thread groups are alternately determined from the first thread set and the second thread set and the target thread groups are executed, where a target thread group includes target threads corresponding to at least two target instructions, and the at least two target instructions correspond to different functional units; thereafter, operands of the target thread group are read from the register resources corresponding to the target thread group; and finally, the target thread group and its corresponding operands are sent to the corresponding functional units. In this way, on the one hand, the multiple threads are grouped and the target thread groups in the two thread sets are alternately executed, so that there will be no register read conflicts when reading the operands corresponding to the target thread groups from the registers, thereby improving the overall performance of instruction execution; on the other hand, compared with the related art solution of using a compiler to compile operands into different register groups, the present disclosure can reduce the compilation difficulty of the compiler. At the same time, compared with the related art solution of increasing the operand reading frequency, the present disclosure does not need to limit the execution efficiency of the functional units, thus achieving the effect of improving the instruction execution efficiency; on the further hand, since the target thread group includes at least two target instructions, and the at least two target instructions correspond to different functional units, in this way, the instruction concurrency efficiency can be improved, and the at least two target instructions can be simultaneously executed by different functional units to improve the instruction execution efficiency and reduce the latency.

[0066] In some embodiments, the instruction scheduling method further includes the following step S305:

[0067] Step S305, based on the register bit identifier of each register, the first register resource and the second register resource are respectively divided into at least two register groups; different register groups have different register read interfaces.

[0068] Here, the register bit identifier of a register refers to the identification information of each data bit in the register.

[0069] In some embodiments, the high-order data areas in the registers have the same register bit identifier, and the low-order data areas have the same register bit identifier. In this way, when the first register resource and the second register resource are respectively divided into at least two register groups, the high-order data areas with the same register bit identifier in all the registers of the register resources can be divided into the same register group, and the low-order data areas with the same register bit identifier can be divided into the same register group.

[0070] In some embodiments, different data bits in a register have different register bit identifiers. In this way, when the first register resource and the second register resource are respectively divided into at least two register groups, the data bits with corresponding register bit identifiers in all the registers of the register resources can be divided into the same register group. For example, when the register bit width is 16 bits and the register bit identifier corresponding to each register is from 0 to 15, the data bits with register bit identifiers from 0 to 7 in all the registers can be divided into the same register group, and the data bits with register bit identifiers from 8 to 15 can be divided into the same register group.

[0071] In some embodiments, the number of divided register groups can be determined based on the number of operands in the instruction to be processed. For example, when the maximum number of operands in an instruction is 2, the first register resource and the second register resource can be respectively divided into 2 register groups; when the maximum number of operands in an instruction is 4, the first register resource and the second register resource can be respectively divided into 4 register groups.

[0072] Here, the register reading interface can be an interface in the form for reading the stored content in the register, and it can implement the register reading function through address lines, data lines, and control signals.

[0073] Through the above division of the register resources, multiple register groups and their corresponding register reading interfaces can be obtained. For example, when the first register resource and the second register resource are respectively divided into two register groups, the first register resource and the second register resource correspond to four register groups. In this way, data can be read from 4 register groups based on 4 register reading interfaces within one clock cycle, that is, 4 groups of data can be read out simultaneously; similarly, if the first register resource and the second register resource are respectively divided into 4 register groups, the first register resource and the second register resource correspond to 8 register groups, and in this way, 8 groups of data can be read out simultaneously within one clock cycle.

[0074] In an embodiment of the present disclosure, as Figure 4As shown, the available register resources 400 include a first register resource 410 and a second register resource 420. Among them, the first register resource 410 includes multiple registers, namely, register 411, register 412, register 413, register 414, etc., and the second register resource 420 includes multiple registers, namely, register 421, register 422, register 423, register 424, etc. At the same time, each register includes a low-order data area represented by a blank area and a high-order data area represented by a shaded area. Based on the method provided in the present disclosure, all the registers in the register resources can be divided into a low-order register group and a high-order register group according to the low-order data area and the high-order data area, so as to obtain the low-order register group and the high-order register group corresponding to the first register 410, and the low-order register group and the high-order register group corresponding to the second register resource 420.

[0075] In this way, the operation of reading the operand of the target thread group from the register resources corresponding to the target thread group, that is, the above step S303, can be implemented as the following steps S3031 to S3032:

[0076] Step S3031, determine at least two register groups corresponding to the target thread group.

[0077] Here, when the target thread group is multiple target threads determined from the first thread set, at least two register groups corresponding to the target thread group are at least two register groups obtained after dividing the first register resource; similarly, when the target thread group is multiple target threads determined from the second thread set, at least two register groups corresponding to the target thread group are at least two register groups obtained after dividing the second register resource.

[0078] Step S3032, based on the corresponding register reading interface, read the operand of the target thread group from the at least two register groups respectively.

[0079] Here, after determining at least two register groups corresponding to the target thread group, based on the register reading interfaces corresponding to the at least two register groups, read the operands of the instructions executed by the target thread from the at least two register groups respectively.

[0080] In some embodiments, the register group reading order corresponding to each register group can be determined in any manner.

[0081] In some embodiments, the reading order of the at least two register groups may be determined based on the arrangement order of the register bit identifiers corresponding to the at least two register groups. For example, in the case where the at least two register groups are a high-order register group and a low-order register group divided according to the high-order data area and the low-order data area of the register, data may be cyclically read from the register in the order from the low-order register group to the high-order register group; for another example, continuing with the above example, in the case where the at least two register groups are a high-order register group and a low-order register group, data may be cyclically read from the register resources in the order of the low-order register group, the low-order register group, the high-order register group, and the high-order register group.

[0082] In the embodiments of the present disclosure, by dividing the first register resource and the second register resource into at least two register groups respectively based on the register bit identifier of each register, and setting corresponding register reading interfaces for different register groups, the reading interfaces of the register resources corresponding to the first thread set and the second thread set are increased, so that more register data (for example, more operands) can be read in one clock cycle, thereby improving the register reading efficiency and the instruction execution efficiency.

[0083] In some embodiments, the step of dividing the first register resource and the second register resource into at least two register groups respectively based on the register bit identifier of each register, that is, the above step S305, may be implemented as the following steps S3051 to S3053:

[0084] Step S3051, obtain the preset operand quantity information; the operand quantity information represents the quantity of operands included in the instruction corresponding to the thread.

[0085] Here, the instruction corresponding to the thread refers to the instruction executed by the thread.

[0086] In some embodiments, when a complete instruction in a computer program is a relatively simple instruction, the complete instruction may be directly used as an instruction that can be executed by one thread; when a complete instruction is a relatively complex and divisible instruction, the complete instruction may be divided into at least two relatively simple instructions for the thread to execute the relatively simple instructions. Among them, the division of the relatively complex complete instruction may be implemented according to any suitable instruction division standard; for example, based on the different functions that the complete instruction can achieve, the complete instruction may be divided into multiple partial instructions for executing different functions; for another example, according to the execution order of the complete instruction, the complete instruction may be divided into multiple partial instructions corresponding to different execution stages; and so on.

[0087] The preset operand quantity information can be any type of information related to the operand quantity corresponding to a thread.

[0088] In some embodiments, the preset operand quantity information can be statistical information on the operand quantities corresponding to multiple threads. For example, based on the operand quantities corresponding to multiple threads, the proportions of threads with one operand, two operands, three operands, etc. in the total number of threads can be determined respectively, and this proportion information can be used as the preset operand quantity information.

[0089] In some embodiments, the preset operand quantity information can be the maximum quantity information of the operands that one thread can correspond to in a processor that executes the instruction scheduling method provided by the present disclosure. During implementation, this maximum quantity information can be determined according to the instruction set architecture, instruction encoding method, and / or processor design of the processor.

[0090] In some embodiments, the preset operand quantity information can be the average quantity information of the operands corresponding to multiple threads in a processor that executes the instruction scheduling method provided by the present disclosure. During implementation, this average quantity information can be determined according to the instruction set architecture, instruction encoding method, and / or processor design of the processor.

[0091] Step S3052: Based on the quantity information and the register bit identifier of each register, divide each register into at least two register partitions.

[0092] Here, the division quantity of each register is determined according to the preset operand quantity information, and each register is divided into at least two register partitions according to the determined division quantity and the register bit identifier of each register.

[0093] In some embodiments, when the quantity information represents the maximum quantity information of the operands that one thread can correspond to, the division quantity of each register can be determined according to this maximum quantity information; among them, the larger the maximum quantity, the more the determined division quantity. For example, when the maximum quantity of the operands is 2 or 3, the division quantity can be determined to be 2, that is, each register is divided into 2 register partitions; when the maximum quantity of the operands is 4, the division quantity of each register can be determined to be 4, that is, each register is divided into 4 register partitions. In this way, as described above, the more register partitions each register corresponds to, the more operands can be read in each clock cycle, so that the operand reading operation can be completed more efficiently.

[0094] In addition, considering that if the registers are divided according to the maximum number of operands corresponding to one thread, and the instructions with the largest number of operands account for a relatively small proportion among multiple threads, it may cause waste of register resources. Therefore, in some other embodiments of the present disclosure, the division quantity can be determined based on the proportion information of threads corresponding to different numbers of operands in the total number of threads. In some embodiments, the division quantity can be determined based on the number of operands corresponding to the thread with the largest proportion in the proportion information. During implementation, the number of operands corresponding to the thread with the largest proportion can be used as the division quantity. For example, when the number of operands corresponding to the thread with the largest proportion is 2, the division quantity can be determined as 2; when the number of operands corresponding to the thread with the largest proportion is 3, the division quantity can be determined as 3. In some embodiments, the division quantity can be determined based on whether the proportion of threads with the number of operands greater than or equal to 2 is greater than a proportion threshold. For example, if the proportion of threads with 2 or 3 operands is greater than the proportion threshold (e.g., 20%), and the proportion of threads with 4 or more operands is less than the proportion threshold, the division quantity is determined as 2 or 3. Another example is that if the proportion of threads with 4 operands is greater than the proportion threshold (e.g., 20%), the division quantity is determined as 4. In this way, by determining the division quantity based on the proportion information, when the proportion of the thread with the largest number of operands is relatively small, the division quantity can be determined according to the number of operands corresponding to the thread with a relatively large proportion.

[0095] During implementation, each register can be evenly divided into at least two register partitions based on the register bit identifier of each register.

[0096] Step S3053, based on the at least two divided register partitions, respectively determine at least two register groups in the first register resource and the second register resource.

[0097] Here, after each register in the first register resource and the second register resource is divided into at least two register partitions, the register partitions with corresponding register bit identifiers can be used as the same register group, so that at least two register groups corresponding to the first register resource and at least two register groups corresponding to the second register resource can be obtained.

[0098] In some embodiments, the target thread group includes a first target thread and a second target thread; the execution order of the target instruction corresponding to the first target thread is before the execution order of the target instruction corresponding to the second target thread.

[0099] Here, the target instructions corresponding to the first target thread and the second target thread can both be computer instructions of any type, and the corresponding functional units are different.

[0100] The execution order refers to the logical sequence of multiple instructions. In one embodiment, the execution order of the target instructions corresponding to the first target thread is before the execution order of the target instructions corresponding to the second target thread, which may be determined by the arrangement order of the target instructions corresponding to the first target thread and the target instructions corresponding to the second target thread in the computer program or specific jump instructions, etc.

[0101] Thus, the operation of reading the operands of the target thread group from the at least two register groups respectively based on the corresponding register read interface, that is, the above step S3032, can be implemented as the following step S3033:

[0102] Step S3033, for each register group, read the operands corresponding to the first target thread from the register group based on the register read interface corresponding to the register group; in response to the completion of reading the operands corresponding to the first target thread, read the operands corresponding to the second target thread from the register group.

[0103] Here, the number of operands corresponding to the first target thread and the second target thread is not limited.

[0104] When reading register data from each register group, based on the execution order of the target instructions, read the operands of the target instructions that are executed first from the register group, and start reading the operands of the next target instruction after the operands of the target instruction that is executed first are read. Thus, based on the read interface of the current register group, read the operands corresponding to the first target thread one by one first, and after the operands corresponding to the first target thread are read, start reading the operands corresponding to the second target thread one by one from the current register group.

[0105] In the embodiments of the present disclosure, determining the order of reading the operands corresponding to the target thread from the register group according to the execution order of the target instructions can make the reading order of the operands match the execution order of the instructions, thereby reducing the operand reading delay and improving the instruction execution efficiency.

[0106] In some embodiments, the target thread group instructions include a third target thread and a fourth target thread; the execution order of the target instructions corresponding to the third target thread is before the execution order of the target instructions corresponding to the fourth target thread, and start reading the operands corresponding to the fourth target thread at a specified clock cycle after the target thread group is issued.

[0107] Here, the target instructions corresponding to the third target thread and the fourth target thread can both be any type of computer instructions, and the corresponding functional units are different.

[0108] In one embodiment, the execution order of the target instruction corresponding to the third target thread is before the execution order of the target instruction corresponding to the fourth target thread, which may be determined by the arrangement order of the target instruction corresponding to the third target thread and the target instruction corresponding to the fourth target thread in the computer program or a specific jump instruction, etc.

[0109] In some embodiments, after the target thread group is dispatched from the thread queue according to a preset reading rule, the operands corresponding to the third target thread are read first, and the operands corresponding to the fourth target thread are read starting from a specified clock cycle after the target thread group is dispatched. In implementation, the preset reading rule may be the default reading rule of the system or the reading rule set by the user. The specified clock cycle after the target thread group is dispatched refers to the clock cycle determined according to the preset reading rule. For example, when the specified clock cycle is the second clock cycle after the target thread group is dispatched, if the target thread group is dispatched at the t0 clock cycle, the operands corresponding to the fourth target thread are read at the t2 clock cycle; if the target thread is dispatched at the t2 cycle, the operands corresponding to the fourth target thread are read at the t4 clock cycle; and so on. In some embodiments, the specified clock cycle may be determined according to the operand type, functional unit, etc. of the target instruction corresponding to the fourth target thread. For example, when the target instruction corresponding to the fourth target thread is an instruction executed by a special function arithmetic unit, the second, sixth, and tenth clock cycles after the target thread group is dispatched may be used as the specified clock cycle; for another example, when the operands of the target instruction corresponding to the fourth target thread are floating-point numbers, the third, seventh, and eleventh clock cycles after the target thread group is dispatched may be used as the specified clock cycle; and so on.

[0110] Thus, the step of reading the operands of the target thread group from the at least two register groups respectively based on the corresponding register reading interface, that is, step S3032 above, may be implemented as the following step S3034 or step S3035:

[0111] Step S3034: For each register group, in response to reaching the specified clock cycle and the operands corresponding to the third target thread not being completely read, based on the register reading interface corresponding to the register group, read the unread operands of the third target thread from the register group; in response to the operands of the third target thread being completely read, based on the register reading interface corresponding to the register group, read the operands corresponding to the fourth target thread from the register group.

[0112] Here, for each register group, if the operand corresponding to the third target thread in the current register group has not been completely read when the specified clock cycle after the target thread group is issued is reached, continue to read the operand corresponding to the third target thread from the current register group, and move the clock cycle for reading the operand corresponding to the fourth target thread backward until the operand corresponding to the third target thread is completely read.

[0113] Step S3035, for each register group, in response to reaching the specified clock cycle and the operand corresponding to the third target thread being completely read, read the operand corresponding to the fourth target thread from the register group based on the register read interface corresponding to the register group.

[0114] Here, for each register group, if the operand corresponding to the third target thread in the current register group has been completely read when the specified clock cycle after the target thread group is issued is reached, read the operand corresponding to the fourth target thread from the current register group according to a preset reading rule.

[0115] In an embodiment of the present disclosure, specifying an operand reading cycle in advance for a fourth target thread whose execution order is later enables the control program in the processor to confirm whether the operand corresponding to the third target thread in the current register group has been completely read according to the pre-specified reading cycle, without the need to confirm in real time whether the operand corresponding to the third target thread in the current register group has been completely read, thereby simplifying the actions of the control program; on the other hand, if the operand corresponding to the third target thread has not been completely read when the pre-specified operand reading cycle for the fourth target thread is reached, continue to read the operand corresponding to the third target thread, which can make the execution order of the instructions match the reading order of the operands, thereby improving the execution efficiency of the instructions.

[0116] In some embodiments, the priority corresponding to the target thread can be determined based on the type of the functional unit of the target instruction corresponding to the target thread. For example, in the case where the target thread group instruction includes a seventh target thread and an eighth target thread, the priorities corresponding to the seventh target thread and the eighth target thread can be determined respectively based on the types of the functional units of the target instructions corresponding to the seventh target thread and the eighth target thread. In this way, the order of reading the operands corresponding to the seventh target thread and the eighth target thread from the register group can be determined based on the priorities corresponding to the seventh target thread and the eighth target thread, and the operand corresponding to the target thread with a lower priority is read starting from the specified clock cycle after the target thread group is issued.

[0117] For example, when the functional unit of the target instruction corresponding to the seventh target thread is a floating-point or integer arithmetic unit, and the functional unit of the target instruction corresponding to the eighth target thread is a special function arithmetic unit, the priority corresponding to the seventh target thread is higher than the priority corresponding to the eighth target thread. In this way, the operands corresponding to the seventh target thread are read at the start of the clock cycle when the target thread group is issued, and the operands corresponding to the eighth target thread are read at a specified clock cycle after the target thread group is issued. When the specified clock cycle is the second clock cycle after the target thread group is issued, if the target thread group is issued at the t0 clock cycle, the operands corresponding to the eighth target thread are read at the t2 clock cycle; if the target thread is issued at the t2 cycle, the operands corresponding to the eighth target thread are read at the t4 clock cycle; and so on.

[0118] Correspondingly, for each register group, in response to reaching the specified clock cycle after the target thread group is issued, if the operands corresponding to the seventh target thread with a higher priority have not been completely read, based on the register reading interface corresponding to the register group, continue to read the operands corresponding to the seventh target thread from the register group; in response to the completion of reading the operands corresponding to the seventh target thread, based on the register reading interface corresponding to the register group, read the operands corresponding to the eighth target thread from the register group. Continuing with the above example, if the target thread group is issued at the t0 clock cycle and the operands corresponding to the seventh target thread have not been completely read at the t2 clock cycle, move the operand reading cycle of the eighth target thread backward by one or more clock cycles (for example, to the t3 clock cycle or the t4 clock cycle) until the operands corresponding to the seventh target thread are completely read; if the target thread is issued at the t2 cycle and the operands corresponding to the seventh target thread have not been completely read at the t4 clock cycle, move the operand reading cycle of the eighth target thread backward by one or more clock cycles (for example, to the t5 clock cycle or the t6 clock cycle) until the operands corresponding to the seventh target thread are completely read; and so on.

[0119] In the embodiment of the present disclosure, by setting the priority for the corresponding target thread according to the type of the functional unit corresponding to the target instruction and reading the operands corresponding to the target thread according to the priority, the operands corresponding to the functional units specified by the user, with a relatively large amount of computation and / or relatively high importance, can be preferentially read, thereby improving the intelligence and flexibility of instruction execution.

[0120] In some embodiments, before the step of alternately determining the target thread group from the first thread set and the second thread set according to the specified clock cycle interval and executing the target thread group, that is, before the above step S302, the instruction scheduling method provided by the present disclosure further includes the following step S306:

[0121] Step S306: Based on the functional units corresponding to the multiple instructions, distribute the multiple instructions to the corresponding multiple threads in the first thread set and the second thread set respectively, so that in the first thread set and the second thread set, the functional units of the instructions corresponding to adjacent threads are different.

[0122] Here, distributing the multiple instructions to the corresponding multiple threads in the first thread set and the second thread set respectively means that the control program of the processor loads the multiple instructions into the working memories of the corresponding multiple threads respectively, so that the threads can access and execute the corresponding instructions.

[0123] In implementation, the multiple instructions can be distributed to the corresponding multiple threads respectively based on the functional units corresponding to the multiple instructions and the thread numbers of the multiple threads, so that in the same thread set, the functional units of the instructions corresponding to the threads with adjacent thread numbers are different.

[0124] In the embodiments of the present disclosure, when distributing instructions to threads, the functional units of the instructions corresponding to adjacent threads in the same thread set are different, so that the control program can respectively determine the target thread groups from the first thread set and the second thread set according to the arrangement order of the threads (i.e., the thread numbers).

[0125] In some embodiments, there is no dependency relationship between multiple operands of at least two target instructions corresponding to the target thread group.

[0126] In some embodiments, if the output or intermediate execution variable / result of a previously executed instruction is used as the input of a subsequently executed instruction, it is considered that there is a dependency relationship between multiple operands of the previously executed instruction and the subsequently executed instruction. In some embodiments, if there is a strict order requirement for the execution of two instructions based on the system configuration, it is considered that there is a dependency relationship between multiple operands of the two instructions.

[0127] Since at least two target instructions corresponding to the target thread group are simultaneously issued and executed, therefore, if there is a dependency relationship between multiple operands of the at least two target instructions, it may be because the previously executed target instruction has not been executed yet (i.e., there is no corresponding output, or the output has not been written into the register), resulting in the subsequently executed target instruction being unable to read the correct operand as input, and further resulting in abnormal situations such as incorrect instruction execution or interruption. Therefore, there is no dependency relationship between multiple operands of at least two target instructions corresponding to the target thread group, so that the at least two target instructions corresponding to the target thread group can be normally executed.

[0128] In some embodiments, the instruction scheduling method of the present disclosure further includes the following step S307:

[0129] Step S307, in response to the at least two target instructions including a fifth target instruction and a sixth target instruction, remove the thread corresponding to the sixth target instruction from the target thread group to update the target thread group;

[0130] Wherein, the fifth target instruction and the sixth target instruction correspond to the same functional unit, or there is a dependency relationship between multiple operands in the fifth target instruction and the sixth target instruction; the instruction execution order of the fifth target instruction is before the instruction execution order of the sixth target instruction.

[0131] Here, when determining the target thread group from the first thread set or the second thread set, read the fifth target instruction and the sixth target instruction from the current thread set, and the instruction execution order of the fifth target instruction is before the instruction execution order of the sixth target instruction.

[0132] In the case where the fifth target instruction and the sixth target instruction correspond to the same functional unit, if the fifth target instruction and the sixth target instruction are executed simultaneously in the same target thread group, it will cause the functional units corresponding to the two target instructions to need to execute the two target instructions simultaneously, thereby reducing the instruction execution efficiency. Therefore, in the case where the fifth target instruction and the sixth target instruction correspond to the same functional unit, remove the thread corresponding to the sixth target instruction with the later execution order from the target thread group to update the target thread group. In this way, since the target instructions included in the updated target thread group are instructions executed by different functional units, the instruction execution efficiency can be improved.

[0133] In the case where there is a dependency relationship between multiple operands in the fifth target instruction and the sixth target instruction, if the fifth target instruction and the sixth target instruction are executed simultaneously in the same target thread group, as described above, it may cause problems such as instruction execution interruption or error. Therefore, in the case where there is a dependency relationship between multiple operands in the fifth target instruction and the sixth target instruction, remove the thread corresponding to the sixth target instruction with the later execution order from the target thread group to update the target thread group. In this way, since there is no dependency relationship between the operands in at least one target instruction included in the updated target thread group, the stability of instruction execution and system security can be improved.

[0134] Next, in conjunction with Figure 5 , the implementation process of an embodiment of the instruction scheduling method provided according to the present disclosure will be described. As Figure 5 shown, this embodiment includes the following steps S501 to S505:

[0135] Step S501, select a target thread group from the first thread set; then, execute step S502;

[0136] Here, the thread to be processed is split into a first thread set and a second thread set, and target thread groups are alternately determined from the first thread set and the second thread set at specified clock cycle intervals.

[0137] Step S502, parse the two target instructions included in the target thread group; then, execute step S503;

[0138] Here, parsing the two target instructions means analyzing and interpreting the two target instructions so that the control program can understand and execute the operations or tasks they represent.

[0139] Step S503, here, determine whether the two target instructions belong to the same functional unit or whether there is a dependency relationship; if so, execute step S504; if not, execute step S505;

[0140] Here, if the two target instructions belong to the same functional unit, or there is a dependency relationship between the operands of the two target instructions, step S504 is executed.

[0141] Step S504, read the operands of the target instruction with the earlier execution order, and send the earlier target instruction and its operands to the corresponding functional unit;

[0142] Here, remove the thread corresponding to the instruction with the later execution order from the target thread group, and only execute the thread corresponding to the target instruction with the earlier execution order as the target thread.

[0143] Step S505, sequentially read the operands of the two target instructions, and send the two target instructions and their operands to the corresponding functional units respectively.

[0144] Here, the threads corresponding to the two target instructions are launched and executed as target threads.

[0145] Continuing with the above embodiment, taking the target thread group selected from the first thread set corresponding to the first instruction and the second instruction, and the target thread group selected from the second thread set corresponding to the third instruction and the fourth instruction as an example, the timing of reading operands from the register resources is described. Among them, the register grouping method is Figure 4 the register grouping method shown in, that is, the first register resource and the second register resource are respectively divided into a low-order register group and a high-order register group.

[0146] First, when the first instruction includes operands A0 and A1, the second instruction includes operands B0 and B1, the third instruction includes operands C0 and C1, the fourth instruction includes operands D0 and D1, and it is pre-specified to start reading the operands of the target instructions (i.e., the second instruction and the fourth instruction) with later execution order in the target thread group at the 2nd clock cycle after the target thread group is issued, the operand reading timing of the first to fourth instructions can be as shown in Table 3 below.

[0147] Among them, in clock cycle T0, the first instruction and the second instruction are determined as the target thread group from the first thread set, and the first instruction and the second instruction are issued; meanwhile, the operand A0 of the first instruction is read. Since the first register resource is divided into a low register group and a high register group, first, based on the reading interface of the low register group of the first register resource, the low data of the operand A0 is read, that is, A0_low;

[0148] In clock cycle T1, based on the reading interface of the high register group of the first register resource, the high data of the operand A0 is read, that is, A0_high; meanwhile, based on the reading interface of the low register group of the first register resource, the low data of the operand A1 of the first instruction is read, that is, A1_low;

[0149] In clock cycle T2, the operand reading cycle pre-specified for the second instruction is reached, and the low data of the operands of the first instruction in the low register group of the first register resource have all been read. Therefore, there is no conflict in the operand reading time of the first instruction and the second instruction in clock cycle T2, and based on the reading interface of the low register group of the first register resource, the low data of the operand B0 of the second instruction can be read, that is, B0_low; meanwhile, based on the reading interface of the high register group of the first register resource, the high data of the operand A1 of the first instruction is read, that is, A1_high;

[0150] Meanwhile, in clock cycle T2, the third instruction and the fourth instruction are determined as the target thread group from the second thread set, and the third instruction and the fourth instruction are issued; meanwhile, based on the reading interface of the low register group of the second register resource, the low data of the operand C0 of the third instruction is read, that is, C0_low;

[0151] In clock cycle T3, based on the read interface of the lower register group of the first register resource, the lower data of operand B1 of the second instruction is read, i.e., B1_low; based on the read interface of the upper register group of the first register resource, the upper data of operand B0_ of the second instruction is read, i.e., B0_high; based on the read interface of the upper register group of the second register resource, the upper data of operand C0_ of the third instruction is read, i.e., C0_high; based on the read interface of the lower register group of the second register resource, the lower data of operand C1 of the third instruction is read, i.e., C1_low;

[0152] In clock cycle T4, in the lower register group of the first register resource, the operand of the second instruction has been read completely; continue to read the upper data of operand B1 of the second instruction based on the read interface of the upper register group of the first register resource, i.e., B1_high; read the lower data of operand D0 of the fourth instruction based on the read interface of the lower register group of the second register resource, i.e., D0_low; meanwhile, read the upper data of operand C1 of the third instruction based on the read interface of the upper register group of the second register resource, i.e., C1_high;

[0153] In clock cycle T5, read the lower data of operand D1 of the fourth instruction based on the read interface of the lower register group of the second register resource, i.e., D1_low; read the upper data of operand D0_ of the fourth instruction based on the read interface of the upper register group of the second register resource, i.e., D0_high;

[0154] In clock cycle T6, read the upper data of operand D1_ of the fourth instruction based on the read interface of the upper register group of the second register resource, i.e., D1_high.

[0155]

[0156]

[0157] Table 3: Timing of operand reading in the case of no conflict

[0158] Secondly, when the first instruction contains operands A0, A1 and A2, the second instruction contains operands B0 and B1, the third instruction contains operands C0 and C1, the fourth instruction contains operands D0 and D1, and it is pre-specified to start reading the operands of the target instructions (i.e., the second instruction and the fourth instruction) with later execution order in the target thread group in the second clock cycle after the target thread group is issued, the timing of operand reading for the first instruction to the fourth instruction can be as shown in Table 4 below.

[0159] For the first instruction and the second instruction determined from the first thread, as can be seen from Table 4, the operand reading timings in clock cycles T0 and T1 are the same as those in Table 3; in clock cycle T2, according to the preset reading rule, the lower - order data of the operand B0 of the second instruction should start to be read. However, since there is still an operand A2 of the first instruction to be read, that is, there are still operands of the previously executed instructions in the lower - order register group of the first register resource that have not been read completely. Therefore, continue to read the lower - order data of the operand A2 of the first instruction based on the reading interface of the lower - order register group of the first register resource, and delay the reading time of the operand B0 of the second instruction by one clock cycle, that is, start reading the operand B0 of the second instruction from clock cycle T3. The operand reading timings of the second instruction in clock cycles T3, T4, and T5 shown in Table 4 respectively correspond to the operand reading timings of the second instruction in clock cycles T2, T3, and T4 in Table 3, which will not be elaborated here.

[0160] Since the number of operands of the third instruction and the fourth instruction determined from the second thread set remains unchanged, as can be seen from Table 4, the operand reading timings of the third instruction and the fourth instruction are the same as those in Table 3, which will not be elaborated here.

[0161]

[0162] Table 4: Operand Reading Timings in Case of Conflicts

[0163] As can be seen from the above embodiments, in the instruction scheduling method of the present disclosure, on the one hand, by grouping multiple threads and alternately executing the target thread groups in two thread sets, register read conflicts do not occur when reading the operands corresponding to the target thread groups from the registers, thereby improving the instruction execution efficiency. On the other hand, compared with the related art solution of using a compiler to compile operands into different register groups, the present disclosure can reduce the compilation difficulty of the compiler. At the same time, compared with the related art solution of increasing the operand read frequency, the present disclosure does not need to limit the execution efficiency of functional units, thus achieving the effect of improving the instruction execution efficiency. On the other hand, since at least two target instructions in the target thread group correspond to different functional units, the at least two target instructions can be executed simultaneously by different functional units, thereby further improving the instruction execution efficiency and reducing the latency. On the other hand, by dividing the first register resource and the second register resource into at least two register groups respectively and determining the corresponding register read interfaces for each register group, more register data (for example, more operands) can be read within one clock scheduling cycle, thereby improving the register read efficiency and the instruction execution efficiency. Finally, in the case where at least two target instructions executed by the target thread group correspond to the same functional unit, or there are dependencies between multiple operands in at least two target instructions, the target thread with a later execution order can be removed from the target thread group, thereby reducing the risk of operand read conflicts and enabling the simultaneously issued target threads to be executed by different functional units, improving the execution efficiency of the functional units, and further improving the instruction execution efficiency.

[0164] Based on the foregoing embodiments, the present disclosure provides an instruction scheduling device. The device includes each unit included and each module included in each unit, and can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits. During implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0165] Figure 6 FIG. is a schematic structural diagram of a composition of an instruction scheduling device provided by the present disclosure, as Figure 6 shown, the instruction scheduling device 600 includes: a splitting module 610, a determining module 620, a register reading module 630, and a sending module 640, where:

[0166] A splitting module 610 for splitting multiple threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to a first register resource and a second register resource;

[0167] A determining module 620 for alternately determining a target thread group from the first thread set and the second thread set at specified clock cycle intervals and executing the target thread group; the target thread group includes target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units;

[0168] A register reading module 630 for reading the operands of the target thread group from the register resource corresponding to the target thread group;

[0169] A sending module 640 for sending the target thread group and its corresponding operands to the corresponding functional unit.

[0170] In some embodiments, the apparatus 600 further includes:

[0171] A register partitioning module for respectively partitioning the first register resource and the second register resource into at least two register groups based on the register bit identifier of each register; different register groups have different register reading interfaces;

[0172] The register reading module 630 is used for:

[0173] Determining at least two register groups corresponding to the target thread group;

[0174] Based on the corresponding register reading interfaces, respectively reading the operands of the target thread group from the at least two register groups.

[0175] In some embodiments, the register partitioning module is used for:

[0176] Obtaining operand quantity information; the operand quantity information represents the quantity of operands included in the instructions corresponding to the threads;

[0177] Based on the quantity information and the register bit identifier of each register, partitioning each register into at least two register partitions;

[0178] Based on the at least two register partitions obtained by partitioning, respectively determining at least two register groups in the first register resource and the second register resource.

[0179] In some embodiments, the target thread group includes a first target thread and a second target thread; the execution order of the target instructions corresponding to the first target thread is before the execution order of the target instructions corresponding to the second target thread;

[0180] The register reading module 630 is configured to, for each register group, based on the register reading interface corresponding to the register group, read the operands corresponding to the first target thread from the register group; in response to the reading of the operands corresponding to the first target thread being completed, read the operands corresponding to the second target thread from the register group.

[0181] In some embodiments, the target thread group instructions include a third target thread and a fourth target thread; the execution order of the target instructions corresponding to the third target thread is before the execution order of the target instructions corresponding to the fourth target thread, and the reading of the operands corresponding to the fourth target thread starts at a specified clock cycle after the target thread group is issued;

[0182] The register reading module 630 is configured to perform one of the following:

[0183] For each register group, in response to reaching the specified clock cycle and the operands corresponding to the third target thread not being completely read, based on the register reading interface corresponding to the register group, read the unread operands corresponding to the third target thread from the register group; in response to the reading of the operands corresponding to the third target thread being completed, based on the register reading interface corresponding to the register group, read the operands corresponding to the fourth target thread from the register group;

[0184] For each register group, in response to reaching the specified clock cycle and the operands corresponding to the third target thread being completely read, based on the register reading interface corresponding to the register group, read the operands corresponding to the fourth target thread from the register group.

[0185] In some embodiments, the apparatus 600 further includes:

[0186] An instruction distribution module, configured to, based on the functional units corresponding to multiple instructions, distribute the multiple instructions to the corresponding multiple threads in the first thread set and the second thread set respectively, so that in the first thread set and the second thread set, the functional units of the instructions corresponding to adjacent threads are different.

[0187] In some embodiments, there is no dependency relationship between multiple operands in at least two target instructions corresponding to the target thread group.

[0188] In some embodiments, the apparatus 600 further includes:

[0189] An update module, in response to the fifth target instruction and the sixth target instruction being included in the at least two target instructions, removes the thread corresponding to the sixth target instruction from the target thread group to update the target thread group;

[0190] Wherein, the fifth target instruction and the sixth target instruction correspond to the same functional unit, or there is a dependency relationship between multiple operands in the fifth target instruction and the sixth target instruction; the instruction execution order of the fifth target instruction is before the instruction execution order of the sixth target instruction.

[0191] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0192] Based on the foregoing embodiments, the present disclosure also provides an electronic device. As Figure 7 shown, the electronic device 700 includes a control unit 710, at least one functional unit 720, a first register resource 730, and a second register resource 740; wherein,

[0193] The control unit 710 is configured to:

[0194] Split multiple threads into a first thread set and a second thread set; the first thread set and the second thread set respectively correspond to the first register resource 730 and the second register resource 740;

[0195] Alternately determine a target thread group from the first thread set and the second thread set at a specified clock cycle interval and execute the target thread group; the target thread group includes target threads corresponding to at least two target instructions; the at least two target instructions correspond to different functional units 720;

[0196] Read the operands of the target thread group from the register resource corresponding to the target thread group;

[0197] Send the target thread group and its corresponding operands to the corresponding functional unit 720.

[0198] In some embodiments, the electronic device 700 further includes a bus 750 to enable the control unit 710, at least one functional unit 720, the first register resource 730, and the second register resource 740 to communicate through the bus 750.

[0199] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For the technical details not disclosed in the device embodiments of the present disclosure, please refer to the description of the method embodiments of the present disclosure for understanding.

[0200] It should be noted that in the embodiments of the present disclosure, if the above instruction scheduling method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs, etc., which can store program codes. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0201] The embodiments of the present disclosure provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0202] The embodiments of the present disclosure provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0203] The embodiments of the present disclosure provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0204] The embodiments of the present disclosure provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0205] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and the similarities or resemblances among them can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0206] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures, or characteristics related to the embodiment are included in at least one embodiment of the present disclosure. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present disclosure, the order numbers of the above steps / processes do not mean the order of execution. The order of execution of each step / process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The serial numbers of the embodiments of the present disclosure above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0207] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0208] In several embodiments provided by the present disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings between the components shown or discussed, or direct couplings, or communication connections can be through some interfaces. The indirect couplings or communication connections between devices or units can be electrical, mechanical or other forms.

[0209] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0210] In addition, each functional unit in the embodiments of the present disclosure may all be integrated into one processing unit, or each unit may be separately used as one unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0211] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks or optical discs and other various media that can store program codes.

[0212] Alternatively, if the above-mentioned integrated units of the present disclosure are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence or the part that contributes to the related art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present disclosure. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks or optical discs and other various media that can store program codes.

[0213] The above is only the implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should all be covered by the protection scope of the present disclosure.

Claims

1. An instruction scheduling method, comprising: Splitting the plurality of threads into a first thread set and a second thread set; The first thread set and the second thread set correspond to a first register resource and a second register resource, respectively; According to a specified clock cycle interval, alternately determine a target thread group from the first thread set and the second thread set and execute the target thread group; the target thread group includes at least two target threads corresponding to target instructions; the at least two target instructions correspond to different functional units; Reading an operand of the target thread group from a register resource corresponding to the target thread group; The target thread group and its corresponding operands are sent to the corresponding functional unit.

2. The method according to claim 1, further comprising: Based on the register bit identifier of each register, respectively divide the first register resource and the second register resource into at least two register groups; Different register groups have different register read interfaces; The step of reading the operand of the target thread group from the register resource corresponding to the target thread group includes: Determine at least two register groups corresponding to the target thread group; Based on the corresponding register read interface, operands of the target thread group are read from the at least two register groups respectively.

3. The method according to claim 2, wherein the first register resource and the second register resource are divided into at least two register groups based on the register bit identification of each register, comprising: Get the number of operands information; The operand quantity information represents the quantity of operands contained in the instruction corresponding to the thread; Based on the quantity information and the register bit identifier of each register, dividing each register into at least two register partitions; At least two register groups in the first register resource and the second register resource are determined respectively based on the at least two divided register partitions.

4. The method according to claim 2, wherein the target thread group comprises a first target thread and a second target thread; the execution order of the target instruction corresponding to the first target thread is before the execution order of the target instruction corresponding to the second target thread; The step of reading the operands of the target thread group from the at least two register groups based on the corresponding register reading interface comprises: For each register group, based on a register read interface corresponding to the register group, read an operand corresponding to the first target thread from the register group; In response to the operands corresponding to the first target thread being read completely, the operands corresponding to the second target thread are read from the register group.

5. The method according to claim 2, wherein the target thread group instruction comprises a third target thread and a fourth target thread; the execution order of the target instruction corresponding to the third target thread is before the execution order of the target instruction corresponding to the fourth target thread, and the operand corresponding to the fourth target thread is read at a specified clock cycle after the target thread group is issued; The step of reading the operands of the target thread group from the at least two register groups based on the corresponding register reading interface comprises one of the following: For each register group, in response to reaching the specified clock cycle and the operands corresponding to the third target thread have not been read completely, reading the unread operands of the third target thread from the register group based on the register read interface corresponding to the register group; In response to completion of reading the operand of the third target thread, based on a register read interface corresponding to the register group, reading the operand corresponding to the fourth target thread from the register group; For each register group, in response to reaching the designated clock cycle and completing reading of operands corresponding to the third target thread, based on a register read interface corresponding to the register group, read operands corresponding to the fourth target thread from the register group.

6. The method according to any one of claims 1 to 5, before alternately determining a target thread group from the first thread set and the second thread set according to a specified clock cycle interval and executing the target thread group, further comprising: Based on the functional units corresponding to the multiple instructions, the multiple instructions are distributed to the multiple threads corresponding to the first thread set and the second thread set respectively, so that the functional units of the instructions corresponding to adjacent threads in the first thread set and the second thread set are different.

7. The method according to any one of claims 1 to 5, wherein there is no dependency between multiple operands in at least two target instructions corresponding to the target thread group.

8. The method according to claim 1, further comprising: In response to the at least two target instructions including a fifth target instruction and a sixth target instruction, removing a thread corresponding to the sixth target instruction from the target thread group to update the target thread group; Among them, the fifth target instruction and the sixth target instruction correspond to the same functional unit, or there is a dependency relationship between multiple operands in the fifth target instruction and the sixth target instruction; the instruction execution order of the fifth target instruction is before the instruction execution order of the sixth target instruction.

9. An instruction scheduling device, comprising: A splitting module, used for splitting the multiple threads into a first thread set and a second thread set; The first thread set and the second thread set correspond to a first register resource and a second register resource, respectively; a determination module, configured to alternately determine a target thread group from the first thread set and the second thread set and execute the target thread group according to a specified clock cycle interval; the target thread group includes at least two target threads corresponding to target instructions; the at least two target instructions correspond to different functional units; A register reading module, used for reading the operand of the target thread group from the register resource corresponding to the target thread group; The sending module is used to send the target thread group and its corresponding operands to the corresponding functional unit.

10. An electronic device, comprising: A control unit, at least one functional unit, a first register resource and a second register resource; wherein, The control unit: Splitting the plurality of threads into a first thread set and a second thread set; the first thread set and the second thread set correspond to a first register resource and a second register resource, respectively; According to a specified clock cycle interval, alternately determine a target thread group from the first thread set and the second thread set and execute the target thread group; the target thread group includes at least two target threads corresponding to target instructions; the at least two target instructions correspond to different functional units; Reading an operand of the target thread group from a register resource corresponding to the target thread group; The target thread group and its corresponding operands are sent to the corresponding functional unit.