Instruction scheduling device, processor, chip and electronic equipment

By introducing instruction shunt and buffer multiplexing mechanisms into the instruction scheduling device, the resource waste caused by instruction cache and buffer decoupling is solved, and hardware resource conservation and system efficiency are improved.

CN120353498AActive Publication Date: 2025-07-22MOORE THREADS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510417292.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In the prior art, in processors with single-instruction multi-threading or single-instruction multi-data architecture, the complete decoupling of the instruction cache and the instruction buffer leads to wasted resources and increases the implementation area.

Method used

The instruction scheduling device is adopted, including an instruction shunt module, a control buffer, an instruction buffer, an instruction buffer and a scheduler. The instruction control information is sent to the control buffer through the shunt processing, and the instruction execution information is sent to the instruction buffer. The scheduler reads the execution information from the buffer according to the instruction control information, and uses the usage characteristics of the instructions at different stages to multiplex the buffer resources.

Benefits of technology

It effectively reduces the implementation area of the instruction buffer, saves hardware resources, and is versatile for different instruction types and processor types, improving the operating efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353498A_ABST
    Figure CN120353498A_ABST
Patent Text Reader

Abstract

The invention relates to an instruction scheduling device, a processor, a chip and electronic equipment, and the instruction scheduling device comprises an instruction shunting module, a control buffer, an instruction buffer, an instruction buffer and a scheduler, instruction control information carried by a to-be-processed instruction is sent to the control buffer, instruction execution information carried by the to-be-processed instruction is sent to the instruction buffer, the instruction buffer reads the instruction control information from the control buffer, and the scheduler sends the instruction execution information to the control buffer according to the instruction control information from the instruction buffer. And reading instruction execution information corresponding to the instruction control information from the instruction buffer. According to the instruction scheduling device, the implementation area of the instruction buffer can be effectively reduced, and implementation resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of chip technology, and in particular, to an instruction scheduling device, a processor, a chip, and an electronic device. Background Art

[0002] For a processor with a Single Instruction Multiple Threads (SIMT) or Single Instruction Multiple Data (SIMD) architecture, which is suitable for processing a large amount of parallel data, the instruction set it supports generally retrieves instruction data from multiple-level caches and places it into the lowest-level instruction cache, and then requests to read the instruction data and send it to the instruction buffer. The instructions in the instruction buffer are selected by the scheduler and issued to the downstream execution unit to perform related operations.

[0003] In related technical solutions, the instruction cache and the instruction buffer are completely decoupled, and the instruction buffer stores the complete instruction information in the instruction cache. Placing the entire instruction into the instruction buffer causes a large amount of resource waste and increases the implementation area. Summary of the Invention

[0004] In view of this, the present disclosure proposes an instruction scheduling scheme.

[0005] According to one aspect of the present disclosure, there is provided an instruction scheduling device, which includes an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction splitting module is connected to the control buffer and the instruction cache. The control buffer is connected to the instruction buffer, and the instruction buffer and the instruction cache are connected to the scheduler. Wherein, the instruction splitting module performs splitting processing on the acquired instruction to be processed, sends the instruction control information carried by the instruction to be processed to the control buffer, and sends the instruction execution information carried by the instruction to be processed to the instruction cache. The instruction buffer reads the instruction control information from the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer.

[0006] In a possible implementation manner, the instruction to be processed is an instruction requested by a request message. The instruction scheduling device further includes a detection module, and the detection module is connected to the control buffer. The detection module is configured to receive the request message, detect the request message, and obtain a detection result, and the detection result is used to characterize the existence state of the instruction control information of the instruction to be processed requested by the request message in the control buffer.

[0007] In a possible implementation, when the detection result indicates that the instruction control information of the to-be-processed instruction requested by the request information exists in the control buffer, the instruction buffer reads the instruction control information from the control buffer; or, when the detection result indicates that the instruction control information of the to-be-processed instruction requested by the request information does not exist in the control buffer, the instruction shunt module is used to request the to-be-processed instruction from the upper-level cache according to the request information, perform shunt processing on the requested to-be-processed instruction, send the instruction control information carried by the to-be-processed instruction into the control buffer, and send the instruction execution information carried by the to-be-processed instruction into the instruction cache; the instruction buffer reads the instruction control information from the control buffer.

[0008] In a possible implementation, the instruction buffer is further used to record the position information of the instruction control information read from the control buffer in the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer, including: the scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information and the position information from the instruction buffer.

[0009] In a possible implementation, the instruction control information and the instruction execution information are used in different stages.

[0010] In a possible implementation, the instruction control information is used in the scheduling stage to distinguish the execution order between tasks, and the instruction execution information is used in the execution stage to provide operation information for executing tasks to the execution units downstream of the scheduler.

[0011] In a possible implementation, the scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer, including: the scheduler selects target instruction control information and the target position information of the target instruction control information from the instruction control information in the instruction buffer according to a preset selection policy; the scheduler reads the instruction execution information corresponding to the target instruction control information from the instruction cache according to the target position information.

[0012] According to another aspect of the present disclosure, a processor is provided, and the processor includes the instruction scheduling device as described above.

[0013] According to another aspect of the present disclosure, a chip is provided, and the chip includes the processor as described above.

[0014] According to another aspect of the present disclosure, there is provided an electronic device, which includes the chip as described above.

[0015] The instruction scheduling device of the embodiments of the present disclosure may include an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. Among them, the instruction splitting module performs splitting processing on the acquired instructions to be processed, sends the instruction control information carried by the instructions to be processed to the control buffer, and sends the instruction execution information carried by the instructions to be processed to the instruction cache; the instruction buffer reads the instruction control information from the control buffer; the scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer. In this way, the instruction scheduling device of the embodiments of the present disclosure can effectively reduce the implementation area of the instruction buffer and save implementation resources.

[0016] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings included in the specification and constituting a part of the specification show exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.

[0018] Figure 1 A schematic diagram showing an instruction scheduling structure in the related art.

[0019] Figure 2 A schematic diagram showing an instruction scheduling device according to an embodiment of the present disclosure.

[0020] Figure 3 A schematic diagram showing another instruction scheduling device according to an embodiment of the present disclosure.

[0021] Figure 4 A block diagram showing an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following will describe various exemplary embodiments, features, and aspects of the present disclosure in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0023] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the existence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.

[0024] When an element is referred to as being "connected", "coupled", "responsive" or variations thereof to another element, it can be directly connected, coupled or responsive to the other element, or intervening elements may be present.

[0025] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0026] The term "exemplary" used herein means "serving as an example, instance, or illustration". Any embodiment illustrated herein as "exemplary" should not necessarily be construed as superior or better than other embodiments.

[0027] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0028] To facilitate the understanding of the technical solution of the present disclosure, professional names or terms in the related art are first explained herein.

[0029] A cache is a high-speed memory used to store data that is frequently accessed. The cache can be located between the processor and the main memory, acting as a bridge between the two. Since the cache is much faster than the main memory, it can significantly improve the data access speed, thereby enhancing the overall system performance.

[0030] The working principle of the cache is to copy data from the main memory to a smaller and faster cache. For a multi-level cache structure, when the processor needs to access data, it first checks the cache at the nearest level. If the data is in this level of cache, it is called a cache hit, and the processor can directly read the data. If the data is not in this level of cache, it is called a cache miss, and the processor needs to read the data from the upper-level cache or the main memory and then load it into the cache for future use.

[0031] A buffer is a storage space built with registers, which is faster or the fastest in access speed than the cache. For example, a buffer can read data in the current cycle and can also read out data in the current cycle.

[0032] It can be seen that the role of cache is to speed up the access to data. For example, when the processor completes a very complex computing task, the result of the computing task can be stored in the cache, so it does not need to be calculated again next time, which speeds up the access to data.

[0033] Figure 1 FIG. 1 is a schematic diagram showing an instruction scheduling structure in the related art. Figure 1 As shown, the detection module 11 can perform a hit-miss check on the acquired request information. If the detection result is a hit, it indicates that the currently requested instruction exists in the current instruction cache 13. If the detection result is a miss, it indicates that the instruction requested by the current request information does not exist in the current instruction cache 13, and it is necessary to request instruction information from the upper level instruction cache, such as the first level cache 12 in the figure. After the first level cache 12 returns the instruction information, it stores the instruction in the instruction cache 13. At this time, the detection result of the re-initiated request will become a hit. After the hit, the instruction data can be read from the instruction cache 13 and stored in the instruction buffer 14. The scheduler 15 performs scheduling based on the instruction information stored in the instruction buffer 14 according to a certain selection strategy. The selected instruction is read out from the instruction buffer 14 and sent to the downstream execution unit for execution.

[0034] like Figure 1 As shown, in the related art, the instruction cache 13 and the instruction buffer 14 are completely decoupled, and the instruction buffer 14 stores the complete instruction information in the instruction cache 13. As a single instruction supports more and more functions, the length of the instruction is getting longer and longer. For a complex instruction set, the length of an instruction is relatively long, for example, 128 bits, but the instruction information actually used by the scheduler 15 only accounts for a very small part of the entire instruction length, and the remaining information is only required by the downstream execution unit. Putting the entire instruction into the instruction buffer 14 causes a large waste of resources and increases the implementation area of the chip.

[0035] In order to reduce the implementation area of the instruction buffer 14, save hardware resources, and effectively utilize the usage characteristics of instructions at different stages, the instruction cache and instruction buffer are reused to achieve the same scheduling and transmission functions.

[0036] Figure 2 FIG. 1 is a schematic diagram showing an instruction scheduling device according to an embodiment of the present disclosure. Figure 2As shown in the figure, the instruction scheduling device includes an instruction shunt module 21, a control buffer 22, an instruction buffer 23, an instruction cache 24, and a scheduler 25. The instruction shunt module 21 is connected to the control buffer 22 and the instruction cache 24. The control buffer 22 is connected to the instruction buffer 23. The instruction buffer 23 and the instruction cache 24 are connected to the scheduler 25.

[0037] The instruction shunt module 21 performs shunt processing on the acquired instruction to be processed, sends the instruction control information carried by the instruction to be processed to the control buffer 22, and sends the instruction execution information carried by the instruction to be processed to the instruction cache 24. The instruction buffer 23 reads the instruction control information from the control buffer 22. The scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction cache 24 according to the instruction control information from the instruction buffer 23.

[0038] Among them, the instruction to be processed may include read / write instructions, logical operation instructions, interrupt instructions, program control instructions, etc., single-thread instructions, multi-thread instructions, image processing instructions, etc. The embodiments of the present disclosure do not limit the type of the instruction to be processed.

[0039] Optionally, the instruction shunt module 21 is used to perform shunt processing on the instruction to be processed, send the instruction control information carried by the instruction to be processed to the control buffer 22, and send the instruction execution information carried by the instruction to be processed to the instruction cache 24. For example, assume that the length of the instruction to be processed is L bits, and the ratio of the instruction control information and the instruction execution information carried by the instruction to be processed is K1:K2. The instruction shunt module 21 can select the instruction control information with a length of L×K1 / (K1+K2) bits from the instruction to be processed with a length of L bits, and send the instruction control information with a length of L×K1 / (K1+K2) bits to the control buffer 22; and select the instruction execution information with a length of L×K2 / (K1+K2) bits from the instruction to be processed with a length of L bits, and send the instruction execution information with a length of L×K2 / (K1+K2) bits to the instruction cache 24. Among them, L, K1, and K2 are positive integers, and the embodiments of the present disclosure do not limit the specific sizes of L, K1, and K2.

[0040] Among them, the instruction shunt module 21 can be implemented by a hardware circuit composed of general digital circuit components. The present disclosure does not limit the specific implementation manner of the instruction shunt module 21.

[0041] Optionally, the control buffer 22 and the instruction buffer 23 can be implemented using a buffer. Among them, the control buffer 22 and the instruction buffer 23 have the same bit width, and the depth of the control buffer 22 can be greater than the depth of the instruction buffer 23. Herein, the bit width represents the number of bits (bit) occupied by storing each piece of information, and the depth represents the size of the total number of pieces of information stored.

[0042] Optionally, the instruction cache 24 can be implemented using a cache (Cache). The depth of the instruction cache 24 is the same as the depth of the control buffer 22, and the sum of the bit widths of the instruction cache 24 and the instruction buffer 23 is the same as the bit width of a complete instruction to be processed (i.e., the number of bits of the binary code in the instruction to be processed).

[0043] Optionally, the scheduler 25 triggers the execution unit to execute the corresponding task according to the instruction control information, at the time or under the conditions indicated by the control information. The scheduler 25 can determine which execution unit should obtain the execution right at what time according to a preset scheduling algorithm (such as first come first served, shortest job first, etc.). Among them, the scheduler 25 can be implemented by a hardware circuit composed of general digital circuit components. The present disclosure does not limit the specific implementation manner of the scheduler 25.

[0044] In this way, after retrieving the instruction to be processed from the upper-level instruction cache (such as Figure 2 the middle-level cache 12), based on the usage characteristics of the instruction to be processed in different stages, that is, using the instruction control information of the instruction to be processed in the scheduling stage and the instruction execution information of the instruction to be processed by the downstream execution unit in the execution stage, the instruction splitter module 21 splits the instruction to be processed, puts the instruction control information used in the scheduling stage into the control buffer 22, and puts the instruction execution information required by the downstream execution unit into the instruction cache 24. When the instruction scheduling device receives the request information, it can read the instruction control information from the control buffer 22 and send it to the instruction buffer 23. These instruction control information will participate in the scheduling process of the scheduler 25. For the selected instruction to be processed, the instruction execution data can be read from the instruction buffer 23 to the downstream execution unit. In this way, while achieving the same function, the instruction buffer 23 only stores a small amount of instruction control information, and most of the information (such as instruction execution information) is still stored in the instruction cache 24, which is equivalent to reusing the resources of the instruction cache 24, thereby saving the hardware resources of the instruction cache 24.

[0045] For example, assume that the instruction to be processed has a length of 128 bits. The ratio of the instruction control information to the instruction execution information carried by the instruction to be processed is 1:3. That is, the instruction control information carried by the instruction to be processed occupies 32 bits, and the instruction execution information carried by the instruction to be processed occupies 96 bits. It should be understood that the embodiments of the present disclosure do not limit the bit width of the instruction to be processed, nor the specific ratio of the instruction control information to the instruction execution information carried by the instruction to be processed, which can be set according to the actual application scenario.

[0046] In the related art, as Figure 1 shown, the space size of the instruction buffer 13 is 1024×128, and the instruction buffer 13 can store 1024 instructions to be processed with a length of 128 bits; the space size of the instruction buffer 14 is 16×128, and the instruction buffer 14 can cache 16 instructions to be processed with a length of 128 bits. The instruction buffer 14 stores the complete instruction information in the instruction buffer 13, all of which are 128 bits.

[0047] On the contrary, in the embodiments of the present disclosure, as Figure 2 shown, the space size of the control buffer 22 is 1024×32, and the control buffer 22 can cache the instruction control information of 1024 instructions to be processed, and each instruction control information occupies 32 bits. The space size of the instruction buffer 23 is 16×32, and the instruction buffer 23 can cache the instruction control information of 16 instructions to be processed, and each instruction control information occupies 32 bits. The space size of the instruction buffer 24 is 1024×96, and the instruction buffer 24 can store the instruction execution information of 1024 instructions to be processed.

[0048] Comparing Figure 1 and Figure 2 it can be seen that Figure 2 the total hardware storage resources occupied by the control buffer 22 and the instruction buffer 24 in Figure 1 (for example, 1024×32 + 1024×96) are similar to the hardware storage resources occupied by the instruction buffer 13 in Figure 1 (for example, 1024×128). However, Figure 2 the hardware storage resources occupied by the instruction buffer 23 in Figure 1 (for example, 16×32) are much smaller than the hardware storage resources of the instruction buffer 14 in Figure 1 (for example, 16×128). The instruction buffer 23 stores 25% of the information of the original instruction to be processed, that is, the instruction buffer 23 can complete the same function with 25% of the overhead of the original hardware storage resources. It can be seen that the embodiments of the present disclosure reduce the implementation area of the instruction buffer 23 and save the implementation resources. In addition, the instruction scheduling device of the embodiments of the present disclosure is not only general for processing different instruction types, but also general for processing instructions of different architectures and different processor types.

[0049] In a possible implementation, the instruction control information and the instruction execution information are used in different stages. The instruction control information is used in the scheduling stage to distinguish the execution order between tasks, and the instruction execution information is used in the execution stage to provide the execution unit downstream of the scheduler 25 with the operation information for executing tasks. The operation information includes operation data for executing tasks, specific behaviors of executing tasks, and operations such as writing back results. For example, the operation information may include an operation code and an address code; the operation code is used to indicate the type or nature of the operation behavior to be completed by the instruction execution information, such as read operation, write operation, logical operation, arithmetic operation, etc. The address code is used to indicate the content of the operation object or the address of the storage unit where it is located.

[0050] Among them, the instruction buffer 23 is a thread group resource, and each thread group can store the instruction control information of multiple instructions to be processed. The instructions to be processed in each thread group are executed sequentially. In the scheduling stage, the scheduler 25 can select an instruction to be processed from a thread group of the instruction buffer 23 every cycle, and in the execution stage, read the corresponding instruction execution information from the instruction cache 24 and send it to the execution unit for execution.

[0051] In this way, by utilizing the usage characteristics of instructions in different stages, the instruction buffer 23 does not need to store the instruction execution information, reducing the implementation area of the instruction buffer 23 and saving implementation resources.

[0052] In a possible implementation, the instruction buffer 23 is further used to record the position information of the instruction control information read from the control buffer 22 in the control buffer 22. The scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction cache 24 according to the instruction control information from the instruction buffer 23, including: the scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction cache 24 according to the instruction control information and the position information from the instruction buffer 23.

[0053] Optionally, the storage locations of the instruction control information and the instruction execution information are in one-to-one correspondence. For example Figure 2As shown in the figure, assume that there are 1024 instructions to be processed. The instruction control information of the first instruction to be processed is stored in the first row position of the control buffer 22, and the instruction execution information of the first instruction to be processed will be stored in the first row position of the instruction cache 24; the instruction control information of the second instruction to be processed is stored in the second row position of the control buffer 22, and the instruction execution information of the second instruction to be processed will be stored in the second row position of the instruction cache 24; and so on. The instruction control information of the 1024th instruction to be processed is stored in the 1024th row position of the control buffer 22, and the instruction execution information of the 1024th instruction to be processed will be stored in the 1024th row position of the instruction cache 24.

[0054] In this way, the instruction buffer 23 can read the instruction control information and the position information of the instruction control information (such as the position number cacheline id) from the control buffer 22, so that the scheduler 25 can select one or more pieces of instruction control information to be currently executed based on the selection policy, and based on the position information of the selected instruction control information, read the instruction execution information from the instruction cache 24 to the downstream execution unit.

[0055] By setting the position information, the instruction execution information corresponding to each instruction control information can be quickly and accurately found from the instruction cache 24. In this way, in addition to storing the instruction control information, the instruction buffer 23 does not need to store the instruction execution information, and can use the position information to quickly and accurately find the instruction execution information corresponding to each instruction control information in the instruction buffer 23, reducing the implementation area of the instruction buffer 23 and saving implementation resources.

[0056] In a possible implementation manner, the scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction cache 24 according to the instruction control information from the instruction buffer 23, including: the scheduler 25 selects target instruction control information and the target position information of the target instruction control information from the instruction control information of the instruction buffer 23 according to a preset selection policy; the scheduler 25 reads the instruction execution information corresponding to the target instruction control information from the instruction cache 24 according to the target position information.

[0057] Among them, the scheduler 25 may include a Completely Fair Scheduler (CFS), a Real-time Scheduler, a Multiqueue Scheduler, etc. Correspondingly, the preset selection strategy may include a strategy for fairly selecting all instruction control information, a selection strategy for setting the response time as the highest priority, a selection strategy for achieving load balancing of execution units among multiple instruction control information, etc., which can be set according to the actual application scenario, and the embodiments of the present disclosure do not limit this.

[0058] For example, assume that the scheduler 25 receives N pieces of instruction control information from the instruction buffer 23. It can select one or more target instruction control information and the target position information of each target instruction control information from the N pieces of instruction control information according to the preset selection strategy. Then, the scheduler 25 can read the instruction execution information corresponding to each target instruction control information from the instruction cache 24 according to the target position information of each target instruction control information.

[0059] Among them, the instruction buffer 23 can be a thread group resource, which may include multiple thread groups, and each thread group can store the instruction control information of multiple instructions to be processed. The instructions to be processed in each thread group can be executed sequentially. In the scheduling stage, the scheduler 25 can select the instruction control information of one instruction to be processed as the target instruction control information from each thread group of the instruction buffer 23 according to the preset selection strategy every cycle, and record the target position information (such as position number or address) of each target instruction control information. Then, the scheduler 25 can read the instruction execution information corresponding to each target instruction control information from the instruction cache 24 according to the target position information of each target instruction control information.

[0060] In this way, the scheduler 25 can quickly identify and select the target instruction control information and its target position information from the instruction buffer 23 through the preset selection strategy. This process reduces the delay of instruction processing, enabling the instructions to be processed to enter the execution stage faster. At the same time, accurately reading the instruction execution information from the instruction cache 24 according to the target position information further accelerates the execution process of the instructions to be processed, improving the overall system operation efficiency.

[0061] In addition, the preset selection strategy can be adjusted and optimized according to different application scenarios and requirements. This enables the scheduler 25 to flexibly adapt to various complex computing tasks, improving the adaptability and scalability of the scheduler 25. The scheduler 25 can further integrate new scheduling algorithms and strategies to meet the needs of future higher-performance and more complex computing tasks.

[0062] Figure 3 Schematic diagram showing another instruction scheduling device according to an embodiment of the present disclosure. As Figure 3 shown, in addition to including an instruction splitting module 31, a control buffer 32, an instruction buffer 33, an instruction cache 34, and a scheduler 35, the instruction scheduling device further includes a detection module 36. The detection module 36 is connected to the control buffer 32 and the upper-level cache (such as the level-1 cache 30). The level-1 cache 30 is connected to the instruction splitting module 31. The instruction splitting module 31 is also connected to the control buffer 32 and the instruction cache 34. The control buffer 32 is connected to the instruction buffer 33. The instruction buffer 33 and the instruction cache 34 are connected to the scheduler 35.

[0063] It should be understood that Figure 3 the instruction splitting module 31, the control buffer 32, the instruction buffer 33, the instruction cache 34, and the scheduler 35 in Figure 2 can reuse the instruction splitting module 21, the control buffer 22, the instruction buffer 23, the instruction cache 24, and the scheduler 25 in

[0064] As Figure 3 shown, the instruction to be processed is the instruction requested by the request information. The detection module 36 is configured to receive the request information and perform a detection on the request information to obtain a detection result. Among them, the detection performed by the detection module 36 includes, for example, a hit-miss check. The detection result is used to characterize the existence status of the instruction control information of the instruction to be processed requested by the request information in the control buffer 32.

[0065] In response to the instruction scheduling device receiving the request information, the detection can be performed by the detection module 36. When the detection result characterizes that the instruction control information of the instruction to be processed requested by the request information exists in the control buffer 32, the instruction buffer 33 reads the instruction control information from the control buffer 32. For example, in a loop or iteration operation scenario, the same instruction will be repeatedly executed multiple times. In this case, if the instruction to be processed requested by the request information is an instruction that has been recently processed, then the instruction control information of this instruction may already exist in the control buffer 32.

[0066] Alternatively, when the detection result indicates that the instruction control information of the to-be-processed instruction requested by the request information does not exist in the control buffer 32, the instruction splitting module 31 is configured to request the to-be-processed instruction from the upper-level cache (such as the level-1 cache 30) according to the request information, and perform splitting processing on the requested to-be-processed instruction, send the instruction control information carried by the to-be-processed instruction into the control buffer 32, and send the instruction execution information carried by the to-be-processed instruction into the instruction cache 34; the instruction buffer 33 reads the instruction control information from the control buffer 32.

[0067] Then, the scheduler 35 can read the instruction execution information corresponding to the instruction control information from the instruction cache 34 according to the instruction control information from the instruction buffer 33.

[0068] By setting the detection module 36, it is possible to quickly determine whether the instruction control information of the required to-be-processed instruction exists in the control buffer 32, thereby avoiding unnecessary access or remote data retrieval (such as reading the to-be-processed instruction from the upper-level cache), and improving the data access speed.

[0069] In summary, in the instruction scheduling device according to the embodiment of the present disclosure, in response to the instruction splitting module obtaining a to-be-processed instruction, the instruction splitting module performs splitting processing on the to-be-processed instruction, sends the instruction control information carried by the to-be-processed instruction into the control buffer, and sends the instruction execution information carried by the to-be-processed instruction into the instruction cache; the instruction buffer reads the instruction control information from the control buffer; the scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer. In this way, by using the usage characteristics of the to-be-processed instruction in different stages, that is, using the instruction control information of the to-be-processed instruction in the scheduling stage, and using the instruction execution information of the to-be-processed instruction by the execution unit downstream of the scheduler in the execution stage, there is no need to store the instruction execution information in the instruction buffer, reducing the implementation area of the instruction buffer and saving implementation resources. Moreover, the instruction scheduling device according to the embodiment of the present disclosure is not only general for processing different instruction types, but also general for processing instructions of different architectures and different processor types.

[0070] It should be noted that although the instruction scheduling device is introduced above by taking Figure 2 and Figure 3 as examples, those skilled in the art can understand that the present disclosure should not be limited thereto. In fact, users can flexibly set according to personal preferences and / or actual application scenarios.

[0071] In a possible implementation, an embodiment of the present disclosure provides a processor, which includes the instruction scheduling device as described above. Among them, the type of the processor may include but is not limited to: Central Processing Unit (CPU), Graphic Processing Unit (GPU), General-Purpose Computing on Graphics Processing Units (GPGPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Tensor Processing Unit (TPU), Field Programmable Gate Array (FPGA), or other programmable logic devices. It may also include a microprocessor, etc. The embodiments of the present disclosure do not limit this.

[0072] In a possible implementation, an embodiment of the present disclosure provides a chip, which includes the processor as described above.

[0073] In a possible implementation, an embodiment of the present disclosure provides an electronic device, which includes the chip as described above. The electronic device may be a User Equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a Personal Digital Assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The embodiments of the present disclosure do not limit this.

[0074] Figure 4 is a block diagram of an electronic device 1900 shown according to an exemplary embodiment. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 4 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors (for example, including the instruction scheduling device described above), and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0075] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows ServerTM), the graphical user interface-based operating system launched by Apple Inc. (MacOS XTM), the multi-user and multi-process computer operating system (UnixTM), the free and open-source Unix-like operating system (LinuxTM), the open-source Unix-like operating system (FreeBSDTM), or the like.

[0076] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, the segment of a program, or the part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0078] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary technical personnel in the technical field to understand the embodiments disclosed herein.

Claims

1. An instruction scheduling device, characterized in that, The instruction scheduling device includes an instruction shunt module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction shunt module is connected to the control buffer and the instruction cache. The control buffer is connected to the instruction buffer. The instruction buffer and the instruction cache are connected to the scheduler; Among them, the instruction shunt module performs shunt processing on the acquired instructions to be processed, sends the instruction control information carried by the instructions to be processed to the control buffer, and sends the instruction execution information carried by the instructions to be processed to the instruction cache; The instruction buffer reads the instruction control information from the control buffer; The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer.

2. The instruction scheduling device according to claim 1, wherein The instruction to be processed is the instruction requested by the request information. The instruction scheduling device further includes a detection module. The detection module is connected to the control buffer. The detection module is used to receive the request information and detect the request information to obtain a detection result. The detection result is used to characterize the existence state of the instruction control information of the instruction to be processed requested by the request information in the control buffer.

3. The instruction scheduling device according to claim 2, wherein When the detection result characterizes that the instruction control information of the instruction to be processed requested by the request information exists in the control buffer, the instruction buffer reads the instruction control information from the control buffer; Alternatively, when the detection result characterizes that the instruction control information of the instruction to be processed requested by the request information does not exist in the control buffer, the instruction shunt module is used to request the instruction to be processed from the upper-level cache according to the request information, and perform shunt processing on the requested instruction to be processed, send the instruction control information carried by the instruction to be processed to the control buffer, and send the instruction execution information carried by the instruction to be processed to the instruction cache; the instruction buffer reads the instruction control information from the control buffer.

4. The instruction scheduling device according to any one of claims 1 to 3, characterized in that The instruction buffer is further used to record the position information of the instruction control information read from the control buffer in the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer, including: The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information and the position information from the instruction buffer.

5. The instruction scheduling device according to any one of claims 1 to 3, characterized in that, The instruction control information and the instruction execution information are used in different stages.

6. The instruction scheduling device according to claim 5, wherein The instruction control information is used in the scheduling stage to distinguish the execution order between tasks. The instruction execution information is used in the execution stage to provide operation information for the execution unit downstream of the scheduler to execute tasks.

7. The instruction scheduling device according to claim 4, wherein The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache according to the instruction control information from the instruction buffer, including: The scheduler selects target instruction control information and target position information of the target instruction control information from the instruction control information in the instruction buffer according to a preset selection strategy; The scheduler reads instruction execution information corresponding to the target instruction control information from the instruction cache according to the target position information.

8. A processor, characterized in that, The processor includes the instruction scheduling device according to any one of claims 1 to 7.

9. A chip, characterized in that, The chip includes the processor according to claim 8.

10. An electronic device, characterized in that, The electronic device includes the chip according to claim 9.

Citation Information

Patent Citations

  • Method and processorfor prefetching instruction lines

    CN101013360A

  • Decoupled processor instruction window and operand buffer

    CN107810476A

  • Instruction fetch circuit

    JP1992361329A

  • Processor provided with buffer for reading instruction

    JP1997319657A

  • Method and apparatus for constructing a pre-scheduled instruction cache

    SG89191A1