Instruction distribution method and device, electronic equipment and computer readable storage medium

By dynamically adjusting the distribution configuration information of the scheduling queue and optimizing the instruction distribution strategy, the problem of uneven scheduling queues in the processor core is solved, thereby improving the balance of instruction distribution and chip performance.

CN116048621BActive Publication Date: 2026-03-03HYGON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

In existing technologies, the uneven distribution of instructions among the scheduling queues of processor cores leads to uneven pipeline loads, affecting chip performance.

Method used

By obtaining the current number of tokens and instructions in the scheduling queue, the distribution configuration information is dynamically adjusted, the instruction distribution strategy is optimized, and the balance of instructions between the hybrid scheduling queues and the balance of the total number of cached micro-instructions is ensured.

Benefits of technology

This improves the balance of the same type of instructions among the mixed scheduling queues and the balance of the total number of cached microinstructions among the various scheduling queues, avoiding pipeline stalls and improving chip performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048621B_ABST
    Figure CN116048621B_ABST
Patent Text Reader

Abstract

An instruction distribution method, device, electronic equipment and computer readable storage medium. The instruction distribution method comprises: obtaining scheduling queue information of each of a plurality of scheduling queues, at least one of the plurality of scheduling queues being configured to store both first instructions of a first type and second instructions of a second type; based on the number of instructions of the first instructions, adjusting first distribution configuration information for the first instructions, and distributing a plurality of first instructions to at least part of the plurality of scheduling queues according to the first distribution configuration information; and updating the current token number after distributing the plurality of first instructions, based on the updated current token number, adjusting second distribution configuration information for the second instructions, and distributing a plurality of second instructions to at least part of the plurality of scheduling queues according to the second distribution configuration information. The method can balance the load of the plurality of scheduling queues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to an instruction distribution method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] The Central Processing Unit (CPU) is the core of a computer system, responsible for computation and control. It is the final execution unit for information processing and program execution. Since its inception, the CPU has undergone tremendous development in logical structure, operating efficiency, and functional scope. The processor core is the most crucial part of the CPU; all calculations and data processing are performed by the core. Summary of the Invention

[0003] At least one embodiment of this disclosure provides an instruction distribution method, comprising: obtaining scheduling queue information for each of a plurality of scheduling queues, wherein at least one of the plurality of scheduling queues is configured to store both first instructions of a first type and second instructions of a second type, the scheduling queue information including at least a current token count corresponding to each scheduling queue and the instruction count of the first instructions at the current time, the current token count indicating the maximum number of instructions that the scheduling queue can receive at the current time; dynamically adjusting first distribution configuration information for the first instructions based on the instruction count of the first instructions, and distributing a plurality of first instructions to at least a portion of the plurality of scheduling queues according to the first distribution configuration information; and updating the current token count after distributing the plurality of first instructions, dynamically adjusting second distribution configuration information for the second instructions based on the updated current token count, and distributing a plurality of second instructions to at least a portion of the plurality of scheduling queues according to the second distribution configuration information.

[0004] At least one embodiment of this disclosure provides an instruction distribution apparatus, comprising: an acquisition unit configured to acquire scheduling queue information for each of a plurality of scheduling queues, wherein at least one of the plurality of scheduling queues is configured to store both first instructions of a first type and second instructions of a second type, the scheduling queue information including at least a current token count corresponding to each scheduling queue and the instruction count of the first instructions at the current time, the current token count indicating the maximum number of instructions that the scheduling queue can receive at the current time; a first adjustment distribution unit configured to dynamically adjust first distribution configuration information for the first instructions based on the instruction count of the first instructions, and to distribute a plurality of first instructions to at least a portion of the plurality of scheduling queues according to the first distribution configuration information; and a second adjustment distribution unit configured to update the current token count after distributing the plurality of first instructions, dynamically adjust second distribution configuration information for the second instructions based on the updated current token count, and to distribute a plurality of second instructions to at least a portion of the plurality of scheduling queues according to the second distribution configuration information. At least one embodiment of this disclosure provides an electronic device, including a processor; a memory including one or more computer program instructions; wherein the one or more computer program instructions are stored in the memory and, when executed by the processor, implement the instructions for the instruction distribution method provided in any embodiment of this disclosure.

[0005] At least one embodiment of this disclosure provides a computer-readable storage medium that non-temporarily stores computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the instruction dispatch method provided in any embodiment of this disclosure is implemented.

[0006] The instruction dispatch method provided in at least one embodiment of this disclosure can improve the balance of the number of instructions of the same type among the corresponding hybrid scheduling queues, and improve the balance of the total number of micro-instructions cached among the scheduling queues. Attached Figure Description

[0007] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.

[0008] Figure 1A A schematic diagram of a processor core microarchitecture is shown;

[0009] Figure 1B A kind of Figure 1A A flowchart of the instruction distribution method in the middle distribution unit;

[0010] Figure 2 A flowchart of an instruction dispatch method provided by at least one embodiment of the present disclosure is shown;

[0011] Figure 3 At least one embodiment of the present disclosure is shown. Figure 2 Flowchart of the method for dynamically adjusting the second distribution configuration information used for the second instruction in step S30;

[0012] Figure 4A At least one embodiment of the present disclosure is shown. Figure 3 Flowchart of step S31;

[0013] Figure 4B At least one embodiment of the present disclosure is shown. Figure 3 Flowchart of step S32;

[0014] Figure 5 This illustration schematically shows at least one embodiment of the present disclosure. Figure 3 Flowcharts of steps S31 and S32;

[0015] Figure 6 A schematic diagram of an instruction distribution unit provided in at least one embodiment of the present disclosure is shown;

[0016] Figure 7 A schematic block diagram of an instruction distribution apparatus provided in at least one embodiment of the present disclosure is shown;

[0017] Figure 8A A schematic block diagram of an electronic device provided in at least one embodiment of the present disclosure is shown;

[0018] Figure 8B A schematic block diagram of another electronic device provided in at least one embodiment of the present disclosure is shown; and

[0019] Figure 9 A schematic diagram of a computer-readable storage medium provided in at least one embodiment of the present disclosure is shown. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0021] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0022] Figure 1A A schematic diagram of a microarchitecture 100 of a processor core microarchitecture is shown.

[0023] like Figure 1A As shown, the microarchitecture 100 of the processor core includes a Level 1 (L1) instruction cache unit 101, an instruction fetch unit 102, a decode unit 103, a dispatch unit 104, a renaming unit 105, physical registers 106, a schedule queue (SQ), an arithmetic and logical unit (ALU), a virtual address generation unit (AGU), a memory access unit 107, and an L1 data cache unit 108.

[0024] like Figure 1A As shown, the scheduling queues include SQ0, SQ1, SQ2, and SQ3. SQ0, SQ1, and SQ2 are all hybrid queues; a hybrid queue is a queue used to store different types of instructions. For example, in... Figure 1A In the example, each of SQ0, SQ1, and SQ2 is configured to store both ALU instructions and AGU instructions. SQ3 is a single queue, which stores instructions of the same type. For example, SQ3 is configured to store ALU instructions.

[0025] like Figure 1A As shown, the microarchitecture 100 of this processor core includes multiple ALUs, namely ALU20, ALU21, ALU22, and ALU23; the microarchitecture 100 of this processor core includes multiple AGUs, namely AGU10, AGU11, and AGU12. Both ALUs and AGUs are execution units. Figure 1AIn the example, each hybrid queue corresponds to a set of execution units, which includes the execution units corresponding to the different types of instructions stored in the hybrid queue. For example, if SQ0 stores ALU instructions and AGU instructions, then SQ0 corresponds to the execution unit group consisting of AGU10 and ALU20. Each individual queue corresponds to the execution unit corresponding to the instructions stored in that individual queue. For example, if SQ3 is configured to store ALU instructions, then SQ3 can correspond to ALU23, which processes ALU instructions.

[0026] The ALU is used to execute addition, subtraction, multiplication, division, AND, OR, NOT instructions, while the AGU is used to perform virtual address calculations for memory access.

[0027] The decoding unit 103 decodes the instructions from the instruction fetch unit 102, generates microinstructions, and sends them to the dispatch unit 104.

[0028] The distribution unit 104 distributes microinstructions to different scheduling queues according to the microinstruction category.

[0029] For example, if a microinstruction is a fixed-point computation operation, it can be provided to one of the SQ0, SQ1, SQ2, and SQ3 containing an ALU pipe according to the ALU distribution mechanism, so that the corresponding ALU can perform the fixed-point computation operation.

[0030] For example, if a microinstruction is a memory access operation, it can be provided to one of the scheduling queues SQ0, SQ1, and SQ2 containing AGU channels (pipes) according to the AGU distribution mechanism, so that the corresponding AGU can calculate the virtual address. If it is a write operation, in addition to providing the microinstruction to one of the scheduling queues SQ0, SQ1, and SQ2 containing AGU channels, it is also provided to one of the scheduling queues SQ0, SQ1, SQ2, and SQ3 containing ALU channels, so that the corresponding ALU can generate the source operand for the write operation.

[0031] The renaming unit 105 renames the source register and the destination register before the microinstruction is written to the scheduling queue.

[0032] Each SQ queues and schedules all received microinstructions out of order. The ALU channel selects executable ALU-type microinstructions for issuance, and the AGU channel selects executable AGU microinstructions for issuance. The issued microinstructions are read from the physical register file and then executed by the corresponding execution unit.

[0033] The ALU is responsible for executing fixed-point computation microinstructions and fixed-point write microinstructions. For fixed-point computation microinstructions, the execution result is written back to fixed-point physical register 106. For fixed-point write microinstructions, the execution result is sent to memory access unit 107. The AGU is responsible for calculating the virtual address of memory access microinstructions, and the execution result (i.e., the virtual address) is sent to memory access unit 107.

[0034] Memory access unit 107 interacts with the ALU and AGU and performs subsequent operations. For example, memory access unit 107 receives write data generated by the ALU and executes subsequent write microinstructions. For example, memory access unit 107 receives a virtual address generated by the AGU and then interacts with the L1 data cache 108. For example, memory access unit 107 sends the result (read data) of a read microinstruction to the corresponding physical register.

[0035] Figure 1B A kind of Figure 1A A flowchart of the instruction distribution method of the middle distribution unit 104.

[0036] like Figure 1B As shown, the instruction distribution method of the distribution unit 104 may include steps S201 to S209.

[0037] Step S201: Obtain microinstructions from the decoding unit.

[0038] Step S202: According to the static checking rules, select P microinstructions that are distributed together in this round. P is an integer greater than or equal to 1. The P microinstructions distributed together in this round are considered as a group of distributed microinstructions.

[0039] Step S203: Calculate the number of tokens required to issue P microinstructions.

[0040] The number of tokens represents the depth of the SQ microinstruction cache. Assuming each SQ has Y tokens, this means each SQ can receive a maximum of Y microinstructions.

[0041] For example, each time a microinstruction is dispatched to a SQ, it represents the consumption of one token, and the token count is reduced by one. Whenever an instruction is fully executed in the execution unit (ALU or AGU) corresponding to the SQ, the execution unit produces one token, and the token count of the corresponding SQ in the next round is increased by one.

[0042] Step S204: Dynamically check whether the number of tokens in the multiple scheduling queues in the processor core at the current moment meets the token requirement of the P microinstructions. For example, dynamically check whether the total number of tokens in the multiple scheduling queues used to store ALU instructions at the current moment is greater than or equal to the number of ALU instructions in the P microinstructions; and dynamically check whether the total number of tokens in the multiple scheduling queues used to store AGU instructions at the current moment is greater than or equal to the number of AGU instructions in the P microinstructions.

[0043] For example, the distribution unit receives the execution unit (e.g., Figure 1B The tokens produced by ALU20~ALU23 and AGU10~AGU12 in the scheduling queue are updated in a timely manner.

[0044] If the number of tokens in multiple scheduling queues at the current moment does not meet the token requirement of the P microinstructions, then step S205 is executed.

[0045] If the number of tokens in multiple scheduling queues at the current moment meets the token requirement of the P microinstructions, then proceed to step S206.

[0046] Step S205: Cancel the current round of distribution and wait for the execution unit to generate tokens until the token count required for the P microinstructions is met.

[0047] Step S206: Dispatch the microinstructions with fixed dispatch types from the P microinstructions. For example, the multiplication instruction has a fixed dispatch type, and the multiplication instruction can only be dispatched to SQ1. Similarly, the division instruction also has a fixed dispatch type, and the division instruction can only be dispatched to SQ2.

[0048] Step S207: Determine the instruction type of the microinstructions other than the fixed dispatch type in the P microinstructions.

[0049] Step S208: If the microinstruction is an ALU instruction (e.g., an instruction for fixed-point computation), then according to the ALU dispatch strategy, the microinstruction is dispatched to one of SQ0, SQ1, SQ2 and SQ3, which contain ALU pipes.

[0050] For example, such as Figure 1B As shown, the ALU distribution strategy can be to distribute ALU instructions in the distribution microinstruction group sequentially according to a fixed priority of SQ0, SQ1, SQ2, and SQ3. When a scheduling queue reaches its maximum number of instructions allocated in this round or has no remaining tokens, no further instructions will be allocated to that scheduling queue. The ALU microinstructions in SQ0, SQ1, SQ2, and SQ3 are executed by ALU20 to ALU23, which correspond to SQ0, SQ1, SQ2, and SQ3, respectively.

[0051] Step S209: If the microinstruction is an AGU instruction (e.g., a memory access operation instruction), then according to the AGU dispatch strategy, the microinstruction is dispatched to one of SQ0, SQ1 and SQ2 that contains an AGU pipe.

[0052] For example, such as Figure 1B As shown, the AGU distribution strategy can be to allocate AGU instructions in the same distribution microinstruction group sequentially according to the fixed priorities of SQ0, SQ1, and SQ2. When a scheduling queue reaches the maximum number of instructions to be allocated in this round or there are no remaining tokens in the scheduling queue, no further allocations will be made to that scheduling queue. The AGU microinstructions in SQ0, SQ1, and SQ2 are executed by AGU10 to AGU12 corresponding to SQ0, SQ1, and SQ2, respectively.

[0053] Different dispatch strategies will determine the number of instructions received by each scheduling queue in this round, which will affect the balance of the number of instructions of the same type among the scheduling queues and the total number of micro-instructions cached by each scheduling queue. If the load of each scheduling queue is unbalanced, it will cause some low-load pipelines to stall, ultimately affecting the performance of the chip.

[0054] At least one embodiment of this disclosure provides an instruction distribution method, instruction distribution apparatus, electronic device, and computer-readable storage medium. The instruction distribution method includes: acquiring scheduling queue information for each of a plurality of scheduling queues, at least one of the scheduling queues being configured to store both first instructions of a first type and second instructions of a second type; the scheduling queue information including at least a current token count corresponding to each scheduling queue and the number of first instructions at the current moment, the current token count indicating the maximum number of instructions a scheduling queue can receive at the current moment; dynamically adjusting first distribution configuration information for the first instructions based on the number of first instructions, and distributing a plurality of first instructions to at least a portion of the plurality of scheduling queues according to the first distribution configuration information; and updating the current token count after distributing the plurality of first instructions, dynamically adjusting second distribution configuration information for the second instructions based on the updated current token count, and distributing a plurality of second instructions to at least a portion of the plurality of scheduling queues according to the second distribution configuration information. This instruction distribution method can improve the balance of the number of instructions of the same type among corresponding mixed scheduling queues and improve the balance of the total number of micro-instructions cached among the scheduling queues.

[0055] Figure 2 A flowchart of an instruction distribution method provided by at least one embodiment of the present disclosure is shown.

[0056] like Figure 2 As shown, the method may include steps S10 to S30.

[0057] Step S10: Obtain the scheduling queue information of each of the multiple scheduling queues. At least one of the multiple scheduling queues is configured to store both the first instruction of the first type and the second instruction of the second type. The scheduling queue information includes at least the current token count and the number of first instructions at the current moment for each scheduling queue. The current token count indicates the maximum number of instructions that the scheduling queue can receive at the current moment.

[0058] Step S20: Based on the number of instructions of the first instruction, dynamically adjust the first distribution configuration information for the first instruction, and distribute the first instruction to at least a portion of the multiple scheduling queues according to the first distribution configuration information.

[0059] Step S30: After distributing multiple first instructions, update the current token count, dynamically adjust the second distribution configuration information for the second instructions based on the updated current token count, and distribute multiple second instructions to at least a portion of the multiple scheduling queues according to the second distribution configuration information.

[0060] Figure 2 The instruction dispatch method shown can, for example, be provided by... Figure 1A The distribution unit 104 in the middle is executed.

[0061] In embodiments of this disclosure, the instructions may be, for example, Figure 1A The microinstructions generated by the decoding unit 103. In the following description, embodiments of this disclosure are illustrated using the example of multiple instructions to be distributed, where the microinstructions generated by the decoding unit 103 are examples.

[0062] For step S10, for example, at least some of the multiple scheduling queues may store both the first type of microinstructions and the second type of microinstructions.

[0063] For example, the first type of microinstruction is an AGU microinstruction, and the second type of microinstruction is an ALU microinstruction; or the first type of microinstruction is an ALU microinstruction, and the second type of microinstruction is an AGU instruction, etc. This disclosure does not limit the first type and the second type.

[0064] For example, in Figure 1A and Figure 1B In the scenario shown, each of SQ0, SQ1, and SQ2 can store both ALU microinstructions and AGU microinstructions, while SQ3 can only store ALU microinstructions. In the embodiments of this disclosure, the scheduling queue that can store both the first type of microinstructions and the second type of microinstructions is referred to as a hybrid scheduling queue.

[0065] For example, the scheduling queue information for each scheduling queue may include the current number of tokens in the scheduling queue, the number of first instructions stored, the number of second instructions stored, and the ID of the scheduling queue.

[0066] The current token count indicates the maximum number of instructions a scheduling queue can receive at the current moment. For example, assuming each scheduling queue has a rated token count of Y, it means that each scheduling queue can receive a maximum of Y microinstructions. Each time a microinstruction is dispatched to a scheduling queue, it consumes one token for that scheduling queue, and the token count of that scheduling queue is decremented by one. Whenever an instruction is fully executed in an execution unit, the execution unit produces a token, and the token count of the corresponding scheduling queue is incremented by one.

[0067] For example, if the current token count of SQ0 is 3, it means that SQ0 can receive a maximum of 3 micro-instructions at the current moment.

[0068] For step S20, the first distribution configuration information may include, for example, the priority of each of the multiple scheduling queues storing the first instruction, and whether the scheduling queue is masked.

[0069] For example, a first combination of multiple scheduling queues is configured to store instructions of a first type, and a second combination of multiple scheduling queues is configured to store instructions of a second type. For example, in Figure 1A In the example, the first scheduling queue combination includes SQ0, SQ1, and SQ2, used to store AGU instructions; the second scheduling queue combination includes SQ0, SQ1, SQ2, and SQ3, used to store ALU instructions. In this embodiment, the first distribution configuration information may include, for example, the priority of each scheduling queue (i.e., SQ0, SQ1, and SQ2) in the first scheduling queue combination.

[0070] For example, the first distribution configuration information of multiple scheduling queues is updated in real time based on the number of instructions of the first instruction stored in each of the multiple scheduling queues.

[0071] In step S20, multiple first instructions are distributed to at least a portion of multiple scheduling queues according to the first distribution configuration information, including: distributing multiple first instructions in combination to the first scheduling queues according to the first distribution configuration information.

[0072] For example, based on the first distribution configuration information, one of a plurality of first instructions is sequentially distributed to at least a portion of each of the first scheduling queue combinations. In this embodiment, only one microinstruction is distributed to the scheduling queue at a time to prevent an excessive number of microinstructions from being allocated to the same scheduling queue at once, thus avoiding new imbalances.

[0073] In some embodiments of this disclosure, the first distribution configuration information includes the priorities of each scheduling queue in the first scheduling queue group. For example, according to the priorities of each scheduling queue in the first scheduling queue group, one of a plurality of first instructions is sequentially distributed to each of at least a portion of the first scheduling queue group.

[0074] For step S30, after distributing multiple first instructions, the current token count of each of the multiple scheduling queues is updated.

[0075] In some embodiments of this disclosure, the second distribution configuration information may include, for example, the priorities of multiple scheduling queues and whether the multiple scheduling queues are masked.

[0076] In step S30, according to the second distribution configuration information, multiple second instructions are distributed to at least a portion of the multiple scheduling queues, including: according to the second distribution configuration information, multiple second instructions are distributed to the second scheduling queues in combination.

[0077] For example, based on the second distribution configuration information, one of a plurality of second instructions is sequentially distributed to at least a portion of each of the second scheduling queue combinations. In this embodiment, only one microinstruction is distributed to the scheduling queue at a time to prevent an excessive number of microinstructions from being allocated to the same scheduling queue at once, thus avoiding new imbalances.

[0078] In some embodiments of this disclosure, the second distribution configuration information may include, for example, the priorities of multiple scheduling queues in the second scheduling queue combination. For example, according to the priorities of each scheduling queue in the second scheduling queue combination, one of a plurality of second instructions is sequentially distributed to at least a portion of each of the second scheduling queue combination.

[0079] In some embodiments of this disclosure, the distribution method for sequentially distributing one of multiple first instructions to each of at least a portion of the first scheduling queue combination in the embodiment of step S20, and sequentially distributing one of multiple second instructions to each of at least a portion of the second scheduling queue combination in the embodiment of step S30, is similar. Therefore, unless otherwise specified below, the first distribution configuration information and the second distribution configuration information are collectively referred to as distribution configuration information, and the first scheduling queue combination and the second scheduling queue combination are collectively referred to as scheduling queue combination. That is, the scheduling queue combination in the following text can refer to either the first scheduling queue combination or the second scheduling queue combination, and the distribution configuration information can refer to either the first distribution configuration information or the second distribution configuration information.

[0080] In embodiments of this disclosure, the dynamic adjustment in steps S20 and S30 refers to changing and selecting the distribution configuration information according to specific circumstances, relative to always using a single distribution configuration information.

[0081] For example, in response to acquiring multiple microinstructions to be distributed, one of the multiple microinstructions is sequentially distributed to each of at least a portion of the scheduling queue combination according to the order in which the microinstructions are distributed to the scheduling queue combination as indicated by the distribution configuration information (e.g., in descending priority order). If, after one round of distribution to at least a portion of the scheduling queue combination, there are still undistributed instructions, one of the undistributed instructions is again sequentially distributed to each of at least a portion of the scheduling queue combination according to the order in which the microinstructions are distributed, until all instructions are distributed.

[0082] For example, if the sorting order of multiple micro-instructions to be distributed is micro-instruction 1, micro-instruction 2, and micro-instruction 3, and the distribution configuration information indicates that the order of distribution of micro-instructions to the scheduling queue combination is scheduling queue 2 and scheduling queue 1, then micro-instruction 1 is distributed to scheduling queue 2, micro-instruction 2 is distributed to scheduling queue 1, and micro-instruction 3 is distributed to scheduling queue 2.

[0083] In other embodiments of this disclosure, for step S20, distributing multiple first instructions to at least a portion of the multiple scheduling queues according to the first distribution configuration information includes: distributing N of the multiple first instructions sequentially to each of the first scheduling queue combinations according to the first distribution configuration information. For step S30, distributing multiple second instructions to at least a portion of the multiple scheduling queues according to the second distribution configuration information includes: distributing M of the multiple second instructions sequentially to each of the second scheduling queue combinations according to the second distribution configuration information. N and M are integers greater than or equal to 1, and less than or equal to the maximum number of instructions distributed in each round of the scheduling queue. This embodiment is suitable for scenarios with a large number of micro-instructions, improving distribution efficiency.

[0084] For example, in response to acquiring multiple AGU micro-instructions to be distributed, according to the order in which the AGU micro-instructions are distributed to SQ0, SQ1, and SQ2 as indicated by the distribution configuration information (e.g., SQ0, SQ1, and SQ2 in descending order of priority), N instructions are sequentially distributed to at least a portion of SQ0, SQ1, and SQ2. If, after one round of distribution to at least a portion of SQ0, SQ1, and SQ2, there are still undistributed instructions, then again, following the order in which the micro-instructions are distributed to SQ0, SQ1, and SQ2, N undistributed instructions are sequentially distributed to at least a portion of SQ0, SQ1, and SQ2, until all instructions are distributed.

[0085] For example, if the multiple instructions to be distributed consist of 1000 AGU instructions (i.e., the first instruction) and 500 ALU instructions, and the first distribution configuration information indicates that the order of distributing AGU micro-instructions to the scheduling queue combination is SQ2, SQ1, SQ0, then N micro-instructions out of the 1000 micro-instructions are distributed sequentially to SQ2, SQ1, and SQ0 according to the order of the 1000 AGU micro-instructions. After distributing the 1000 AGU instructions, the current token count of each scheduling queue is updated, and the second distribution configuration information used for ALU instructions is readjusted. For example, if the second distribution configuration information indicates that the order of distributing AGU micro-instructions to the scheduling queue combination is SQ2, SQ3, SQ0, and SQ1, then M micro-instructions out of the 500 micro-instructions are distributed sequentially to SQ2, SQ1, and SQ0 according to the order of the 500 AGU micro-instructions.

[0086] In some embodiments of this disclosure, M and N can be equal, meaning the number of first instructions distributed to the scheduling queues in the first scheduling queue combination each time and the number of second instructions distributed to the scheduling queues in the second scheduling queue combination each time are equal. N and M can also be unequal, meaning the number of first instructions distributed to the scheduling queues in the first scheduling queue combination each time and the number of second instructions distributed to the scheduling queues in the second scheduling queue combination each time are different.

[0087] In some embodiments of this disclosure, the first scheduling queue combination is a subset of the second scheduling queue combination. For example, the first scheduling queue combination includes SQ0, SQ1, and SQ2; the second scheduling queue combination includes SQ0, SQ1, SQ2, and SQ3, where SQ0, SQ1, and SQ2 are subsets of SQ0, SQ1, SQ2, and SQ3.

[0088] In this embodiment, the first instruction is preferentially allocated to the first scheduling queue combination, which is a subset, and the second instruction is allocated to the second scheduling queue combination. This is beneficial to achieve the balance of the second instruction in the second scheduling queue combination and the balance of the total number of instructions among the scheduling queues, while ensuring the balance of the allocation of the first instruction in the first scheduling queue combination.

[0089] like Figure 2 As shown, in addition to steps S10 to 30, the instruction distribution method may also include step S40.

[0090] Step S40: Each scheduling queue sends the instructions in each scheduling queue to the execution unit corresponding to each scheduling queue, so that the execution unit corresponding to each scheduling queue executes the instructions.

[0091] In some embodiments of this disclosure, the execution unit may be, for example, an ALU or an AGU. For example, in Figure 1AIn the architecture, the execution units are ALU20, ALU21, ALU22 and ALU23, and AGU10, AGU11 and AGU12.

[0092] like Figure 1A As shown, SQ0 corresponds to ALU20 and AGU10, SQ1 corresponds to ALU21 and AGU11, SQ2 corresponds to ALU22 and AGU12, and SQ3 corresponds to ALU23. SQ0 distributes the ALU instructions and AGU instructions in SQ0 to ALU20 and AGU10 respectively, SQ1 distributes the ALU instructions and AGU instructions in SQ1 to ALU21 and AGU11 respectively, SQ2 distributes the ALU instructions and AGU instructions in SQ2 to ALU22 and AGU12 respectively, and SQ3 distributes the ALU instructions in SQ3 to ALU23.

[0093] In some embodiments of this disclosure, step S20, dynamically adjusting the first distribution configuration information for the first instruction based on the number of instructions in the first instruction, includes: adjusting the priority of the scheduling queue with a larger number of instructions in the first scheduling queue combination to a lower priority than the scheduling queue with a smaller number of instructions in the first scheduling queue combination. That is, the scheduling queue with fewer instructions in the first instruction has a higher priority.

[0094] For example, the first scheduling queue combination includes scheduling queue A and scheduling queue B. If the number of first instructions in scheduling queue A is greater than the number of first instructions in scheduling queue B, then the priority of scheduling queue A is lower than the priority of scheduling queue B.

[0095] For example, in Figure 1A and Figure 1B In the example shown, the number of ALU instructions in SQ0 is greater than the number of ALU instructions in SQ1, so the priority of SQ1 is higher than the priority of SQ0.

[0096] This method prioritizes scheduling queues with fewer first instructions over scheduling queues with more first instructions, thus prioritizing the allocation of first instructions to scheduling queues with fewer first instructions and ensuring the balance of first instructions in the first scheduling queue combination.

[0097] In some embodiments of this disclosure, step S20, distributing multiple first instructions to at least some of the multiple scheduling queues according to the first distribution configuration information, includes: distributing multiple first instructions sequentially to the scheduling queues in the first scheduling queue combination whose current token count is not 0 according to the priority of each scheduling queue in the first scheduling queue combination.

[0098] For example, in Figure 1A and Figure 1BIn the example shown, the priority order based on the number of first instructions in the first scheduling queue combination SQ0, SQ1 and SQ2 is SQ1, SQ0 and SQ2. Since the current token count of SQ2 is 0, multiple first instructions are distributed to SQ1 and SQ0 in sequence according to the priority of SQ1 and SQ0. The multiple first instructions are stored in the AGU pipes in SQ1 and SQ0.

[0099] In some embodiments of this disclosure, the second distribution configuration information includes the priority of each scheduling queue in the second scheduling queue combination.

[0100] In some embodiments of this disclosure, for the scheduling queues in the second scheduling queue combination, the current token count is updated after multiple first instructions are distributed, and the scheduling queue with a larger updated current token count has a higher priority than the scheduling queue with a smaller updated current token count.

[0101] For example, in Figure 1A and Figure 1B In the example shown, the second scheduling queue combination includes SQ0, SQ1, SQ2, and SQ3. Since SQ0, SQ1, and SQ2 are a mixed queue, after the first instruction is distributed to SQ0, SQ1, and SQ2, the current token counts of SQ0, SQ1, SQ2, and SQ3 are updated. For example, after the update, the current token counts of SQ0, SQ1, SQ2, and SQ3 are 4, 1, 2, and 3, respectively. Therefore, the priority order of each scheduling queue in the second scheduling queue combination from highest to lowest is: SQ0, SQ3, SQ2, and SQ1.

[0102] The following text Figures 3-5 An embodiment of dynamically adjusting second distribution configuration information is shown, according to another embodiment of this disclosure.

[0103] Figure 3 At least one embodiment of the present disclosure is shown. Figure 2 The flowchart of the method for dynamically adjusting the second distribution configuration information used for the second instruction in step S30.

[0104] like Figure 3 As shown, the method may include steps S31 to S33. In this embodiment, the scheduling queue information also includes the number of instructions corresponding to the second instruction in each scheduling queue in the second scheduling queue combination at the current time.

[0105] Step S31: Based on the number of instructions of the second instruction of each scheduling queue in the second scheduling queue combination, select the target scheduling queue from the second scheduling queue combination that needs to adjust the updated current token count.

[0106] Step S32: Adjust the updated current token count of the target scheduling queue to the adjustment value according to the second type of instruction count.

[0107] Step S33: Determine the priority of each scheduling queue in the second scheduling queue combination based on the adjustment value.

[0108] Before distributing the second instruction, this method filters out the target scheduling queues that need to dynamically adjust the number of tokens based on the number of second instructions in each scheduling queue in the second scheduling queue combination. This appropriately reduces the number of tokens in scheduling queues with more second instructions. Then, the second instruction is distributed based on the updated number of tokens in the scheduling queues. This ensures a balanced distribution of the load among the mixed scheduling queues while also addressing the issue of balanced distribution of the second instruction in the second scheduling queue combination.

[0109] For step S31, for example, the scheduling queue in the second scheduling queue combination where the number of instructions of the second instruction is greater than a preset threshold is taken as the target scheduling queue.

[0110] Figure 4A At least one embodiment of the present disclosure is shown. Figure 3 The flowchart of step S31.

[0111] like Figure 4A As shown, step S31 may include steps S311 to S314.

[0112] Step S311: Obtain a token to adjust the threshold.

[0113] Step S312: Obtain the average number of instructions of the second type in the second scheduling queue combination.

[0114] Step S313: Obtain the difference between each of the second scheduling queue combinations and the mean.

[0115] Step S314: For each of the second scheduling queue combinations, select the scheduling queue with a difference greater than the token adjustment threshold as the target adjustment queue.

[0116] For step S311, for example, receiving user input for a token adjustment threshold. In other embodiments of this disclosure, the token adjustment threshold may be preset by those skilled in the art.

[0117] The token adjustment threshold can be any natural number such as 0, 1, or 2. This disclosure does not limit the token adjustment threshold.

[0118] For step S312, for example in Figure 1A and Figure 1BIn the example shown, the second scheduling queue combination includes SQ0, SQ1, SQ2, and SQ3. For example, if the number of instructions of the second type stored in SQ0, SQ1, SQ2, and SQ3 (i.e., the number of instructions of the second type) are 12, 16, 14, and 15 respectively, then the mean can be equal to 14, which is the result of (12+16+14+15) / 4 rounded down as the mean.

[0119] For step S313, for example, the differences between the number of instructions of the second instruction of SQ0, SQ1, SQ2 and SQ3 and the average value 14 are -2, 2, 0 and 1 respectively.

[0120] For step S314: the token adjustment threshold can be 0 for example. Then the difference between SQ3 and SQ1 is greater than the token adjustment threshold. Therefore, SQ3 and SQ1 are the target adjustment queues.

[0121] In this embodiment, the scheduling queue that stores a significantly higher number of instructions than the average number of second instructions is selected as the target adjustment queue.

[0122] Figure 4B At least one embodiment of the present disclosure is shown. Figure 3 The flowchart of step S32.

[0123] like Figure 4B As shown, step S32 may include steps S321 to S323.

[0124] Step S321: Compare the difference in the target scheduling queue with the updated current token count in the target scheduling queue.

[0125] Step S322: In response to the updated current token count of the target scheduling queue being greater than the difference, the adjustment value is the updated current token count of the target scheduling queue minus the difference.

[0126] Step S323: In response to the updated current token count of the target scheduling queue being less than or equal to the difference, the adjustment value is 0.

[0127] For step S321, the updated current token count of the target scheduling queue is the current token count obtained by updating the token count of the target scheduling queue after distributing multiple first instructions. For example, if step S313 calculates a difference of 2 for the target scheduling queue and the updated current token count of the target scheduling queue is 4, then the updated current token count of the destination scheduling queue, 4, is greater than the difference of 2. As another example, if step S313 calculates a difference of 2 for the target scheduling queue and the updated current token count of the target scheduling queue is 1, then the updated current token count of the destination scheduling queue, 1, is less than the difference of 2.

[0128] For step S322, for example, in response to the updated current token count 4 of the target scheduling queue being greater than the difference 2, the adjustment value is the result of 4 minus 2, that is, the adjustment value is 2.

[0129] For step S323, for example, in response to the updated current token count 1 of the target scheduling queue being less than the difference 2, the adjustment value is 0.

[0130] In this embodiment, if the difference in the target scheduling queue is greater than the current token count, the adjustment value of the target scheduling queue is the current token count minus the difference; if the difference in the target scheduling queue is less than or equal to the current token count, the adjustment value of the target scheduling queue is 0. Therefore, this embodiment appropriately reduces the current token count of the target scheduling queue with a large number of instructions in the second instruction.

[0131] Figure 5 This illustration schematically shows at least one embodiment of the present disclosure. Figure 3 The flowchart of steps S31 and S32.

[0132] like Figure 5 As shown, the method may include steps S501 to S507.

[0133] Step S501: Calculate the average number of ALU instructions (hereinafter referred to as "ALU instruction count") in SQ0, SQ1, SQ2 and SQ3.

[0134] For example, if the number of ALU instructions in SQ0, SQ1, SQ2 and SQ3 are alu_num_0, alu_num_1, alu_num_2 and alu_num_3 respectively, then the mean AVG = (alu_num_0 + alu_num_1 + alu_num_2 + alu_num_3) / 4.

[0135] This step S501 is similar to execution. Figure 4A Step S312 in the process.

[0136] Step S502: Calculate the difference between the number of ALU instructions and the mean AVG in each SQ (i.e., SQ0, SQ1, SQ2 and SQ3).

[0137] This step S501 is similar to execution. Figure 4A Step S313 in the process.

[0138] For example, the differences between the number of ALU instructions and the mean AVG in SQ0, SQ1, SQ2 and SQ3 are D_value_0, D_value_1, D_value_2 and D_value_3, respectively.

[0139] Step S503: Determine whether each difference D_value_i (i = 0, 1, 2, 3) is greater than the token adjustment threshold.

[0140] For scheduling queues where the difference D_value_i is less than or equal to the token adjustment threshold, step S504 is executed. For example, if D_value_0 and D_value_2 are greater than the token adjustment threshold, then step S504 is executed for SQ0 and SQ2.

[0141] For scheduling queues where the difference D_value_i is greater than the token adjustment threshold, step S505 is executed. For example, if D_value_1 and D_value_3 are greater than the token adjustment threshold, then step S505 is executed for SQ1 and SQ3.

[0142] Step S504: Do not adjust the current token count of SQ. That is, do not adjust the current token count of scheduling queues whose difference D_value_i is less than or equal to the token adjustment threshold. For example, do not adjust the current token count of SQ0 and SQ2.

[0143] Step S505: Determine whether the current number of tokens is greater than the difference D_value_i.

[0144] For scheduling queues whose current token count is less than or equal to the difference D_value_i, proceed to step S506. For example, if the current token count of SQ1 is less than or equal to the difference D_value_1, then proceed to step S506 for SQ1.

[0145] For scheduling queues where the current token count is greater than the difference D_value_i, proceed to step S507. For example, if the current token count of SQ3 is greater than the difference D_value_3, then proceed to step S507 for SQ3.

[0146] Step S506: Adjust the current token count of SQ to 0. That is, the adjustment value of the scheduling queue whose current token count is greater than the difference D_value_i is 0.

[0147] This step S506 is similar to execution. Figure 4B Step S323 in the example. For example, the current token count of SQ1 is adjusted to 0.

[0148] Step S507: The current token count of SQ minus the difference is taken as the current token count of SQ. That is, the adjustment value of the scheduling queue whose current token count is greater than the difference D_value_i is equal to the current token count minus the difference.

[0149] This step S507 is similar to execution. Figure 4B Step S322 in the example. For example, the adjusted token count for SQ3 is the current token count minus D_value_3.

[0150] for Figure 3 Step S33: Determine the priority of each scheduling queue in the second scheduling queue combination based on the adjustment value.

[0151] In some embodiments of this disclosure, the current token count is adjusted based on an adjustment value, and the priority of each scheduling queue in the second scheduling queue combination is determined based on the adjusted current token count.

[0152] For example, in Figure 1A and Figure 1B In the example shown, after the first instruction is distributed to SQ0, SQ1, and SQ2, and the current token counts of SQ0, SQ1, SQ2, and SQ3 are updated, the current token counts of SQ0, SQ1, SQ2, and SQ3 are 4, 1, 2, and 3, respectively. If based on... Figure 4A The described method determines the target scheduling queues as SQ3 and SQ1, and the adjustment values ​​of SQ3 and SQ1 are 1 and 0 respectively. Then, the current token counts of SQ0, SQ1, SQ2 and SQ3 are adjusted to 4, 0, 2 and 1 respectively. Based on 4, 0, 2 and 1, the priorities of SQ0, SQ1, SQ2 and SQ3 are determined.

[0153] For example, the scheduling queue with more adjusted current tokens has a higher priority. Therefore, in the example above, since the adjusted token counts of SQ0, SQ1, SQ2, and SQ3 are 4, 0, 2, and 1 respectively, the order of priority from highest to lowest is: SQ0, SQ2, SQ3, and SQ1.

[0154] In some embodiments of this disclosure, step S30, which distributes multiple second instructions to at least some of the multiple scheduling queues according to the second distribution configuration information, includes: distributing multiple second instructions sequentially to the scheduling queues in the second scheduling queue combination whose current token count is not 0, according to the priority of each scheduling queue in the second scheduling queue combination.

[0155] For example, in some embodiments of this disclosure, the priority of each scheduling queue in the second independent scheduling combination can be according to... Figures 3-5 The current token count is determined by the method shown after adjustment. For example, if the current token counts of SQ0, SQ1, SQ2, and SQ3 are adjusted to 4, 0, 2, and 1 respectively, then the priority order from highest to lowest is: SQ0, SQ2, SQ3, and SQ1. Since the current token count of SQ1 is 0, no second instruction is issued to SQ1. Therefore, in this example, step S30 is to issue multiple second instructions to SQ0, SQ2, and SQ3 sequentially.

[0156] Figure 6A schematic diagram of an instruction dispatch unit 600 provided in at least one embodiment of the present disclosure is shown. The following is in conjunction with... Figure 6 The provided instruction distribution unit 600 further illustrates the instruction distribution method provided in at least one embodiment of this disclosure.

[0157] like Figure 6 As shown, the instruction distribution unit 600 may include an AGU instruction allocation information preprocessing subunit 601, an AGU instruction issuing subunit 602, an ALU instruction allocation information preprocessing unit 603, and an ALU instruction issuing subunit 604. The ALU instruction issuing subunit 604 is configured to distribute ALU instructions executed by the ALU to SQs containing ALU channels (e.g., scheduling queues SQ0 to SQ3 in the second scheduling queue combination described above), and the AGU instruction issuing subunit 602 is configured to issue AGU instructions executed by the AGU to SQs containing AGU channels (e.g., scheduling queues SQ0 to SQ2 in the first scheduling queue combination described above).

[0158] like Figure 6 As shown, the instruction dispatch unit 600 may also include scheduling queue information for multiple scheduling queues. This scheduling queue information may include, for example, the number of AGU instructions and ALU instructions for each of SQ0 to SQ3, as well as their respective current token counts. This scheduling queue information can be updated based on the tokens generated and completed instructions from the execution units (ALU20 to ALU23 and AGU10 to AGU12).

[0159] Figure 6 The described instruction dispatch method assumes that both static and dynamic checks meet the requirement that the current token count in the scheduling queues corresponding to different instruction types is greater than or equal to the token count required for P microinstructions, thus ensuring that the current round of microinstruction dispatch will succeed. After fixed-dispatch type instructions are issued to their corresponding scheduling queues and the token counts in each scheduling queue are updated, Figure 6 The described instruction dispatch method can optimize the allocation strategy for ordinary type instructions (instructions other than fixed dispatch type instructions, such as ALU instructions, AGU instructions, etc.), thereby improving the balance of the number of instructions of the same type among the corresponding hybrid scheduling queues, and improving the balance of the total number of micro-instructions cached among the scheduling queues.

[0160] After fixed-dispatch type instructions are issued to the corresponding scheduling queues and the token counts of each scheduling queue are updated, the process of dynamically allocating AGU instructions and ALU instructions among the hybrid scheduling queues is as follows: Figure 6As shown. Since the scheduling queues (SQ0~SQ2) that AGU type micro-instructions can go to are a subset of the scheduling queues (SQ0~SQ3) that ALU type micro-instructions can go to, AGU type micro-instructions are distributed first, the tokens of the corresponding SQ are consumed and updated, and then ALU type micro-instructions are distributed.

[0161] The AGU instruction allocation information preprocessing unit 601 includes a scheduling queue distribution priority dynamic adjustment subunit 611.

[0162] First, the AGU instruction allocation information preprocessing unit 601 sorts the number of AGU instructions in the scheduling queues containing AGU channels (i.e., the first scheduling queue combinations SQ0 to SQ2) in ascending order. If two SQs have the same number of AGU instructions, their order remains unchanged. The scheduling queue distribution priority dynamic adjustment subunit 611 assigns priorities from high to low according to the sorted SQ order. For example, 0 represents the highest priority, 1 the next highest, and so on. The fewer the number of AGU instructions in an SQ containing AGU channels, the higher the priority of that SQ, thus dynamically changing the priorities of SQ0 to SQ2 during AGU instruction distribution.

[0163] The AGU instruction issuing unit 602 prioritizes distributing AGU instructions to scheduling queues with available tokens and higher priority (fewer AGU instruction allocations), thus achieving a balanced distribution of AGU instructions among scheduling queues SQ0 to SQ2. Each scheduling queue allocates one instruction at a time, preventing the instantaneous allocation of too many instructions to the same queue and creating new imbalances. Instructions are distributed sequentially from high to low priority to each SQ containing an AGU channel until all instructions are allocated. Each AGU instruction distribution consumes a token from the corresponding SQ. If any AGU instructions remain unallocated after one round of allocation, instructions are re-allocated to the corresponding SQs one by one according to the same priority until all AGU instructions are allocated. If, during the allocation process, an SQ runs out of available tokens, that scheduling queue no longer participates in AGU instruction allocation. This AGU distribution strategy ensures a balanced distribution of AGU instructions among SQ0, SQ1, and SQ2.

[0164] In some embodiments of this disclosure, SQ is a hybrid scheduling queue. Instruction distribution considers not only the balance of different instruction types but also the overall load balance of the scheduling queues. This is because if only the balance of ALU instructions among scheduling queues is considered during ALU instruction distribution, a situation might arise where a hybrid scheduling queue has a high load and few tokens, but few ALU instructions. If ALU instructions are still allocated to this scheduling queue, even if it helps balance the number of ALUs in each SQ, it will exacerbate the overall load imbalance of the scheduling queues, ultimately affecting chip performance. Therefore, while the balanced distribution of AGU instructions among scheduling queues is considered during AGU distribution, the load balance of scheduling queues needs to be considered during ALU distribution. Thus, during ALU instruction distribution, the scheduling queue priority is adjusted according to the number of tokens in each hybrid scheduling queue; scheduling queues with more tokens are allocated higher priority. Simultaneously, to maintain the balance of ALU instruction numbers among SQs, when a scheduling queue has too many ALU instructions, the number of tokens in that scheduling queue is appropriately reduced, thereby appropriately lowering its distribution priority.

[0165] The ALU instruction allocation information preprocessing unit 603 includes a scheduling queue token number dynamic adjustment subunit 613 and a scheduling queue distribution priority dynamic adjustment subunit 623.

[0166] The specific flowchart for the dynamic adjustment of the scheduling queue token count performed by the scheduling queue token count dynamic adjustment subunit 613 is as follows: Figure 5 As shown.

[0167] The ALU scheduling queue distribution priority dynamic adjustment unit 623 sorts the adjusted token counts of the scheduling queues (SQ0 to SQ3) containing the ALU Pipe in descending order. If two SQs have the same adjusted token count, their order remains unchanged. The dynamic adjustment unit assigns priorities from highest to lowest according to the sorted SQ order, with 0 representing the highest priority and 1 the next highest. The larger the adjusted token count, the higher the scheduling queue distribution priority. This allows for dynamic modification of the priorities of SQ0 to SQ3 during ALU instruction distribution.

[0168] The ALU instruction distribution unit 604 prioritizes allocation to scheduling queues with available tokens and higher priority, thereby achieving load balancing among scheduling queues SQ0 to SQ3. Each scheduling queue distributes one instruction at a time, preventing the instantaneous distribution of too many instructions to the same queue and creating new load imbalances. Instructions are distributed sequentially from high to low priority to each SQ containing an ALU channel until all instructions are allocated. Each distributed ALU instruction consumes a token from the corresponding SQ. If any ALU instructions remain unallocated after one round of distribution, instructions are re-allocated to the corresponding SQs one by one according to the same priority until all ALU instructions are allocated. If an SQ runs out of tokens during the allocation process, that scheduling queue no longer participates in ALU instruction distribution. This ALU distribution strategy ensures load balance among SQ0 to SQ3 and, to a certain extent, also guarantees a balance in the number of ALU type instructions among SQ0 to SQ3.

[0169] Figure 7 A schematic block diagram of an instruction distribution apparatus 700 provided in at least one embodiment of the present disclosure is shown.

[0170] For example, such as Figure 7 As shown, the instruction distribution device 700 includes an acquisition unit 710, a first adjustment distribution unit 720, and a second adjustment distribution unit 730.

[0171] The acquisition unit 710 is configured to acquire the scheduling queue information of each of a plurality of scheduling queues, wherein at least one of the plurality of scheduling queues is configured to store both a first type of first instruction and a second type of second instruction, and the scheduling queue information includes at least the current token count corresponding to each scheduling queue and the instruction count of the first instruction at the current time, wherein the current token count indicates the maximum number of instructions that the scheduling queue can receive at the current time.

[0172] For example, unit 710 can be executed. Figure 2 Step S10 is described.

[0173] The first adjustment and distribution unit 720 is configured to dynamically adjust the first distribution configuration information for the first instruction based on the number of instructions of the first instruction, and distribute multiple first instructions to at least some of the multiple scheduling queues according to the first distribution configuration information.

[0174] The first adjustment and distribution unit 720 can, for example, execute... Figure 2 Step S20 is described.

[0175] The second adjustment and distribution unit 730 is configured to update the current token count after distributing the plurality of first instructions, dynamically adjust the second distribution configuration information for the second instructions based on the updated current token count, and distribute the plurality of second instructions to at least a portion of the plurality of scheduling queues according to the second distribution configuration information.

[0176] The second adjustment and distribution unit 730 can, for example, perform... Figure 2 Step S30 is described.

[0177] For example, the acquisition unit 710, the first adjustment and distribution unit 720, and the second adjustment and distribution unit 730 can be hardware, software, firmware, or any feasible combination thereof. For example, the acquisition unit 710, the first adjustment and distribution unit 720, and the second adjustment and distribution unit 730 can be dedicated or general-purpose circuits, chips, or devices, or they can be a combination of a processor and memory. The embodiments of this disclosure do not limit the specific implementation of the above-mentioned units.

[0178] It should be noted that in the embodiments of this disclosure, each unit of the instruction distribution device 700 corresponds to each step of the aforementioned instruction distribution method. For the specific functions of the instruction distribution device 700, please refer to the relevant description of the instruction distribution method, which will not be repeated here. Figure 7 The components and structure of the instruction distribution device 700 shown are merely exemplary and not limiting. The instruction distribution device 700 may also include other components and structures as needed.

[0179] At least one embodiment of this disclosure also provides an electronic device including a processor and a memory, the memory including one or more computer program modules. The one or more computer program modules are stored in the memory and configured to be executed by the processor, and the one or more computer program modules include instructions for implementing the instruction dispatch method described above. This electronic device can improve the balance of the number of instructions of the same type among corresponding hybrid scheduling queues, and improve the balance of the total number of microinstructions cached among the various scheduling queues.

[0180] Figure 8A This is a schematic block diagram of an electronic device provided for some embodiments of this disclosure. For example... Figure 8A As shown, the electronic device 800 includes a processor 810 and a memory 820. The memory 820 stores non-transitory computer-readable instructions (e.g., one or more computer program modules). The processor 810 executes the non-transitory computer-readable instructions, which, when executed by the processor 810, can perform one or more steps in the instruction dispatch method described above. The memory 820 and the processor 810 can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0181] For example, processor 810 may be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For example, the central processing unit (CPU) may be an x86 or ARM architecture. Processor 810 may be a general-purpose processor or a special-purpose processor, capable of controlling other components in electronic device 800 to perform desired functions.

[0182] For example, memory 820 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and processor 810 may run one or more computer program modules to implement various functions of electronic device 800. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0183] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic device 800 can be referred to the description of the instruction distribution method above, and will not be repeated here.

[0184] Figure 8B This is a schematic block diagram of another electronic device provided in some embodiments of the present disclosure. The electronic device 900 is, for example, suitable for implementing the instruction distribution method provided in the embodiments of the present disclosure. The electronic device 900 may be a terminal device, etc. It should be noted that... Figure 8B The illustrated electronic device 900 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0185] like Figure 8BAs shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 910, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 920 or a program loaded from a storage device 980 into a random access memory (RAM) 930. The RAM 930 also stores various programs and data required for the operation of the electronic device 900. The processing unit 910, the ROM 920, and the RAM 930 are interconnected via a bus 940. An input / output (I / O) interface 950 is also connected to the bus 940.

[0186] Typically, the following devices can be connected to I / O interface 950: input devices 960 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 970 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 980 including, for example, magnetic tapes, hard disks, etc.; and communication devices 990. Communication device 990 allows electronic device 900 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 8B An electronic device 900 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 900 may alternatively implement or have more or fewer devices.

[0187] For example, according to embodiments of this disclosure, the instruction distribution method described above can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the instruction distribution method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 990, or installed from a storage device 980, or installed from a ROM 920. When the computer program is executed by the processing device 910, the functions defined in the instruction distribution method provided by embodiments of this disclosure can be implemented.

[0188] At least one embodiment of this disclosure also provides a computer-readable storage medium for storing non-transitory computer-readable instructions, which, when executed by a computer, enable the instruction dispatch method described above. Using this computer-readable storage medium, the balance of the number of instructions of the same type among corresponding hybrid scheduling queues can be improved, as can the balance of the total number of microinstructions cached among the various scheduling queues.

[0189] Figure 9 This is a schematic diagram of a storage medium provided for some embodiments of this disclosure. For example... Figure 9As shown, the storage medium 1000 is used to store non-transitory computer-readable instructions 1010. For example, when the non-transitory computer-readable instructions 1010 are executed by a computer, one or more steps in the instruction dispatch method described above can be performed.

[0190] For example, the storage medium 1000 can be used in the aforementioned electronic device 800. For example, the storage medium 1000 can be... Figure 8A The memory 820 in the illustrated electronic device 800. For example, a description of the storage medium 1000 can be found here. Figure 8A The corresponding description of the memory 820 in the illustrated electronic device 800 will not be repeated here.

[0191] The following points need to be explained:

[0192] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.

[0193] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0194] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. The scope of protection of this disclosure should be determined by the scope of protection of the claims.

Claims

1. An instruction dispatch method, comprising: Obtain scheduling queue information for each of multiple scheduling queues, wherein at least one of the multiple scheduling queues is configured to store both a first type of first instruction and a second type of second instruction, and the scheduling queue information includes at least the current token count and the number of first instructions at the current moment corresponding to each scheduling queue, wherein the current token count indicates the maximum number of instructions that the scheduling queue can receive at the current moment; Based on the number of instructions in the first instruction, the first distribution configuration information for the first instruction is dynamically adjusted, and according to the first distribution configuration information, multiple first instructions are distributed to at least a portion of the multiple scheduling queues. The first scheduling queue combination in the multiple scheduling queues is configured to store instructions of the first type, and the first distribution configuration information includes the priority of each scheduling queue in the first scheduling queue combination; and After distributing the plurality of first instructions, the current token count is updated. Based on the updated current token count, the second distribution configuration information for the second instructions is dynamically adjusted. Then, according to the second distribution configuration information, the plurality of second instructions are distributed to at least a portion of the plurality of scheduling queues. The second scheduling queue combination among the plurality of scheduling queues is configured to store instructions of the second type. The second distribution configuration information includes the priority of each scheduling queue in the second scheduling queue combination. The step of dynamically adjusting the first distribution configuration information for the first instruction based on the number of instructions in the first instruction includes: The priority of the scheduling queue with the larger number of first instructions in the first scheduling queue combination is adjusted to be lower than the priority of the scheduling queue with the smaller number of first instructions in the first scheduling queue combination. Specifically, after distributing the plurality of first instructions, the current token count is updated, and based on the updated current token count, the second distribution configuration information used for the second instruction is dynamically adjusted, including: For the scheduling queues in the second scheduling queue combination, after distributing the plurality of first instructions, the current token count is updated, and the scheduling queue with the larger updated current token count has a higher priority than the scheduling queue with the smaller updated current token count.

2. The method according to claim 1, wherein, The first scheduling queue combination is a subset of the second scheduling queue combination.

3. The method according to claim 1, wherein, Based on the first distribution configuration information, a plurality of first instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the first distribution configuration information, multiple first instructions are distributed to the first scheduling queue combination; Based on the second distribution configuration information, a plurality of second instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the second distribution configuration information, multiple second instructions are distributed to the second scheduling queue combination.

4. The method according to claim 1, wherein, Based on the first distribution configuration information, a plurality of first instructions are distributed to at least a portion of the plurality of scheduling queues, including: According to the priority of each scheduling queue in the first scheduling queue combination, the plurality of first instructions are sequentially distributed to the scheduling queues in the first scheduling queue combination whose current token count is not 0.

5. The method according to claim 1, wherein, Based on the second distribution configuration information, a plurality of second instructions are distributed to at least a portion of the plurality of scheduling queues, including: According to the priority of each scheduling queue in the second scheduling queue combination, the plurality of second instructions are sequentially distributed to the scheduling queues in the second scheduling queue combination whose current token count is not 0.

6. The method according to claim 1 or 5, wherein, The scheduling queue information also includes the number of instructions corresponding to the second instruction in each scheduling queue in the second scheduling queue combination at the current time. After distributing the plurality of first instructions, the current token count is updated, and based on the updated current token count, the second distribution configuration information used for the second instruction is dynamically adjusted, including: Based on the number of instructions of the second instruction in each scheduling queue in the second scheduling queue combination, select the target scheduling queue from the second scheduling queue combination that needs to adjust the updated current token count; Based on the number of instructions of the second type, the updated current token count of the target scheduling queue is adjusted to an adjustment value; Based on the adjustment value, the priority of each scheduling queue in the second scheduling queue combination is determined.

7. The method according to claim 6, wherein, Based on the number of instructions of the second type, select the target scheduling queue from the second scheduling queue combination that needs to adjust the updated current token count, including: Obtain a token to adjust the threshold; Obtain the average number of instructions of the second type in the second scheduling queue combination; Obtain the difference between each of the second scheduling queue combinations and the mean; and For each of the second scheduling queue combinations, the scheduling queue with the difference greater than the token adjustment threshold is designated as the target scheduling queue.

8. The method according to claim 7, wherein, Based on the number of instructions of the second type, the updated current token count of the target scheduling queue is adjusted to an adjustment value, including: Compare the difference in the target scheduling queue with the updated current token count in the target scheduling queue; In response to the updated current token count of the target scheduling queue being greater than the difference, the adjustment value is the updated current token count of the target scheduling queue minus the difference; and The adjustment value is 0 in response to the updated current token count of the target scheduling queue being less than or equal to the difference.

9. The method according to claim 1, wherein, Based on the first distribution configuration information, a plurality of first instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the first distribution configuration information, one of the plurality of first instructions is distributed sequentially to each of the first scheduling queue combination; Based on the second distribution configuration information, a plurality of second instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the second distribution configuration information, one of the plurality of second instructions is distributed sequentially to each of the second scheduling queue combination.

10. The method according to claim 1, wherein, Based on the first distribution configuration information, a plurality of first instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the first distribution configuration information, N of the plurality of first instructions are distributed sequentially to each of the first scheduling queue combination; Based on the second distribution configuration information, a plurality of second instructions are distributed to at least a portion of the plurality of scheduling queues, including: Based on the second distribution configuration information, M of the plurality of second instructions are sequentially distributed to each of the second scheduling queue combinations. Where N and M are integers greater than or equal to 1.

11. The method according to claim 1, wherein, The multiple scheduling queues are either arithmetic and logic operation unit scheduling queues or virtual address calculation unit scheduling queues.

12. An instruction distribution device, comprising: The acquisition unit is configured to acquire the scheduling queue information of each of a plurality of scheduling queues, wherein at least one of the plurality of scheduling queues is configured to store both a first type of first instruction and a second type of second instruction, and the scheduling queue information includes at least the current token count corresponding to each scheduling queue and the instruction count of the first instruction at the current time, wherein the current token count indicates the maximum number of instructions that the scheduling queue can receive at the current time; A first adjustment and distribution unit is configured to dynamically adjust first distribution configuration information for the first instruction based on the number of instructions in the first instruction, and to distribute multiple first instructions to at least a portion of the multiple scheduling queues according to the first distribution configuration information. The first scheduling queue combination in the multiple scheduling queues is configured to store instructions of the first type, and the first distribution configuration information includes the priority of each scheduling queue in the first scheduling queue combination. The second adjustment and distribution unit is configured to update the current token count after distributing the plurality of first instructions, dynamically adjust the second distribution configuration information for the second instructions based on the updated current token count, and distribute the plurality of second instructions to at least a portion of the plurality of scheduling queues according to the second distribution configuration information. The second scheduling queue combination in the plurality of scheduling queues is configured to store instructions of the second type, and the second distribution configuration information includes the priority of each scheduling queue in the second scheduling queue combination. The step of dynamically adjusting the first distribution configuration information for the first instruction based on the number of instructions in the first instruction includes: The priority of the scheduling queue with the larger number of first instructions in the first scheduling queue combination is adjusted to be lower than the priority of the scheduling queue with the smaller number of first instructions in the first scheduling queue combination. Specifically, after distributing the plurality of first instructions, the current token count is updated, and based on the updated current token count, the second distribution configuration information used for the second instruction is dynamically adjusted, including: For the scheduling queues in the second scheduling queue combination, after distributing the plurality of first instructions, the current token count is updated, and the scheduling queue with the larger updated current token count has a higher priority than the scheduling queue with the smaller updated current token count.

13. An electronic device, comprising: processor; Memory, which includes one or more computer program instructions; The one or more computer program instructions are stored in the memory and, when executed by the processor, implement the instructions of the instruction dispatch method according to any one of claims 1-11.

14. A computer-readable storage medium that non-transitoryly stores computer-readable instructions, wherein, The instruction dispatch method according to any one of claims 1-11 is implemented when the computer-readable instructions are executed by a processor.

Citation Information

Patent Citations

  • Method, device and system for realizing addition of traffic shaping token

    CN101599905A

  • Context-aware dynamic command scheduling for a data storage system

    CN108958907A