Instruction scheduling device, processor, chip and electronic equipment

By introducing instruction splitting modules and buffer cache reuse into the instruction scheduling device, the problem of wasted instruction cache and buffer resources is solved, thereby saving hardware resources and improving system efficiency.

CN120353498BActive Publication Date: 2025-10-28MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510417292.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-10-28
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

In existing technologies, processors with single-instruction multithreading or single-instruction multiple data architectures completely decouple instruction caches and instruction buffers, leading to resource waste and increased implementation area.

Method used

An instruction scheduling device is adopted, including an instruction splitting module, a control buffer, an instruction buffer, an instruction register, and a scheduler. Through splitting processing, instruction control information is sent to the control buffer, and instruction execution information is sent to the instruction register. The scheduler reads the execution information from the register according to the instruction control information, thereby saving resources of the instruction buffer.

Benefits of technology

It effectively reduces the implementation area of ​​the instruction buffer, saves hardware resources, and is universal for different instruction types and processor types, thus improving the system's operating efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353498B_ABST
    Figure CN120353498B_ABST
Patent Text Reader

Abstract

This disclosure relates to an instruction scheduling device, processor, chip, and electronic device. The instruction scheduling device includes an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction splitting module splits the acquired instructions to be processed, sending the instruction control information carried by the instructions to be processed into the control buffer and the instruction execution information carried by the instructions to be processed into the instruction cache. The instruction buffer reads the instruction control information from the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache based on the instruction control information from the instruction buffer. The instruction scheduling device of this disclosure can effectively reduce the implementation area of ​​the instruction buffer and save implementation resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of chip technology, and in particular to an instruction scheduling device, processor, chip, and electronic device. Background Technology

[0002] For processors with Single Instruction Multiple Threads (SIMT) or Single Instruction Multiple Data (SIMD) architectures, they are suitable for processing large amounts of parallel data. The instruction set they support generally retrieves instruction data from multiple levels of cache and puts it into the lowest level instruction cache. Then, it requests to read the instruction data and send it to the instruction buffer. The instructions in the instruction buffer are selected by the scheduler and issued to the downstream execution unit to perform the relevant operations.

[0003] In related technical solutions, the instruction cache and instruction buffer are completely decoupled, with the instruction buffer storing complete instruction information from the instruction cache. Placing the entire instruction in the instruction buffer results in significant resource waste and increases the implementation area. Summary of the Invention

[0004] In view of this, this disclosure proposes an instruction scheduling scheme.

[0005] According to one aspect of this disclosure, an instruction scheduling device is provided, comprising an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction splitting module is connected to the control buffer and the instruction cache, the control buffer is connected to the instruction buffer, and the instruction buffer, the instruction cache, and the scheduler are connected. The instruction splitting module performs splitting processing on acquired instructions to be processed, sending instruction control information carried by the instructions to be processed into the control buffer and sending instruction execution information carried by the instructions to be processed into the instruction cache. The instruction buffer reads the instruction control information from the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache based on the instruction control information from the instruction buffer.

[0006] In one possible implementation, the instruction to be processed is an instruction requesting information. The instruction scheduling device further includes a detection module connected to the control buffer. The detection module is used to receive the request information and detect the request information to obtain a detection result. The detection result is used to characterize the existence status of the instruction control information of the instruction requesting information to be processed in the control buffer.

[0007] In one possible implementation, if the detection result indicates that the instruction control information of the pending instruction requested by the request information exists in the control buffer, the instruction buffer reads the instruction control information from the control buffer; or, if the detection result indicates that the instruction control information of the pending instruction requested by the request information does not exist in the control buffer, the instruction splitting module is used to request the pending instruction from the upper-level cache according to the request information, and to split the requested pending instruction, sending the instruction control information carried by the pending instruction into the control buffer, and sending the instruction execution information carried by the pending instruction into the instruction buffer; the instruction buffer reads the instruction control information from the control buffer.

[0008] In one possible implementation, the instruction buffer is further configured to record the position information of the instruction control information read from the control buffer within the control buffer. The scheduler reads instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information from the instruction buffer, including: the scheduler reads instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information from the instruction buffer and the position information.

[0009] In one possible implementation, the instruction control information and the instruction execution information are used at different stages.

[0010] In one possible implementation, the instruction control information is used in the scheduling phase to distinguish the execution order between tasks, and the instruction execution information is used in the execution phase to provide the execution units downstream of the scheduler with operation information for executing tasks.

[0011] In one possible implementation, the scheduler reads instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information from the instruction buffer, including: the scheduler selecting target instruction control information and target location information of the target instruction control information from the instruction control information in the instruction buffer according to a preset selection strategy; and the scheduler reading instruction execution information corresponding to the target instruction control information from the instruction buffer based on the target location information.

[0012] According to another aspect of this disclosure, a processor is provided, the processor including the instruction scheduling means described above.

[0013] According to another aspect of this disclosure, a chip is provided that includes the processor described above.

[0014] According to another aspect of this disclosure, an electronic device is provided, the electronic device including the chip described above.

[0015] The instruction scheduling device of this disclosure embodiment may include an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction splitting module splits the acquired instructions to be processed, sending the instruction control information carried by the instructions to be processed into the control buffer and the instruction execution information carried by the instructions to be processed into the instruction cache. The instruction buffer reads the instruction control information from the control buffer. The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache based on the instruction control information from the instruction buffer. In this way, the instruction scheduling device of this disclosure embodiment can effectively reduce the implementation area of ​​the instruction buffer and save implementation resources.

[0016] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0017] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.

[0018] Figure 1 A schematic diagram of the instruction scheduling structure in related technologies is shown.

[0019] Figure 2 A schematic diagram of an instruction scheduling apparatus according to an embodiment of the present disclosure is shown.

[0020] Figure 3 A schematic diagram of another instruction scheduling apparatus according to an embodiment of the present disclosure is shown.

[0021] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation

[0022] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0023] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.

[0024] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.

[0025] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.

[0026] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0027] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.

[0028] To facilitate understanding of the technical solutions disclosed herein, this document first explains the technical names or terms used in the relevant technologies.

[0029] A cache is a high-speed memory used to store frequently accessed data. It can reside between the processor and main memory, acting as a bridge between the two. Because a cache is much faster than main memory, it can significantly improve data access speed, thereby enhancing overall system performance.

[0030] Caching works by copying data from main memory to a smaller, faster cache. In a multi-level cache architecture, when the processor needs to access data, it first checks the nearest cache level. If the data is in that cache level, it's called a cache hit, and the processor can read the data directly. If the data is not in that cache level, it's called a cache miss, and the processor needs to read the data from the next higher cache level or main memory before loading it into the cache for later use.

[0031] A buffer is a storage space built using registers, which is faster, or the fastest, to access than a cache. For example, a buffer can read data in the current cycle and also read data out in the current cycle.

[0032] As we can see, the purpose of caching is to speed up data retrieval. For example, when a processor completes a complex computational task, and the result of that task is needed for the next computation, it can be stored in the cache so that it doesn't need to be computed again, thus speeding up data retrieval.

[0033] Figure 1 A schematic diagram of the instruction scheduling structure in related technologies is shown. For example... Figure 1 As shown, the detection module 11 performs a hit-miss check on the acquired request information. If the check result is a hit, it indicates that the requested instruction exists in the current instruction cache 13. If the check result is a miss, it indicates that the instruction requested by the current request information does not exist in the current instruction cache 13 and needs to be requested from the next higher level instruction cache, such as the first-level cache 12 in the figure. After the first-level cache 12 returns the instruction information, it stores the instruction in the instruction cache 13. At this time, the detection result of the re-initiated request will become a hit. After a hit, the instruction data can be read from the instruction cache 13 and stored in the instruction buffer 14. The scheduler 15 schedules instructions based on the instruction information stored in the instruction buffer 14 according to a certain selection strategy. The selected instruction is read from the instruction buffer 14 and sent to the downstream execution unit for execution.

[0034] like Figure 1 As shown, in related technologies, instruction buffer 13 and instruction cache 14 are completely decoupled. Instruction cache 14 stores the complete instruction information from instruction buffer 13. As a single instruction supports more and more functions, the instruction length also increases. For complex instruction sets, an instruction can be quite long, for example, 128 bits. However, the instruction information actually used by scheduler 15 only accounts for a small portion of the total instruction length; the remaining information is only needed by downstream execution units. Placing the entire instruction in instruction cache 14 results in significant resource waste and increases the chip's implementation area.

[0035] To reduce the implementation area of ​​instruction buffer 14, save hardware resources, and effectively utilize the usage characteristics of instructions at different stages, instruction registers and instruction buffers are reused to achieve the same scheduling and issue functions.

[0036] Figure 2 A schematic diagram of an instruction scheduling apparatus according to an embodiment of the present disclosure is shown. Figure 2As shown, the instruction scheduling device includes an instruction splitting module 21, a control buffer 22, an instruction buffer 23, an instruction cache 24, and a scheduler 25. The instruction splitting module 21 is connected to the control buffer 22 and the instruction cache 24. The control buffer 22 is connected to the instruction buffer 23. The instruction buffer 23 and the instruction cache 24 are connected to the scheduler 25.

[0037] The instruction splitting module 21 splits the acquired instructions to be processed, sending the instruction control information carried by the instructions to be processed into the control buffer 22 and sending the instruction execution information carried by the instructions to be processed into the instruction buffer 24; the instruction buffer 23 reads the instruction control information from the control buffer 22; the scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction buffer 24 according to the instruction control information from the instruction buffer 23.

[0038] The instructions to be processed may include read / write instructions, logical operation instructions, interrupt instructions, program control instructions, single-threaded instructions, multi-threaded instructions, image processing instructions, etc. The embodiments of this disclosure do not limit the type of instructions to be processed.

[0039] Optionally, the instruction splitting module 21 is used to split the instructions to be processed, sending the instruction control information carried by the instructions to be processed into the control buffer 22, and sending the instruction execution information carried by the instructions to be processed into the instruction buffer 24. For example, assuming the length of the instructions to be processed is L bits, and the ratio of instruction control information to instruction execution information carried by the instructions to be processed is K1:K2, the instruction splitting module 21 can select L×K1 / (K1+K2) bits of instruction control information from the L-bit instructions to be processed, and send the L×K1 / (K1+K2) bits of instruction control information into the control buffer 22; and select L×K2 / (K1+K2) bits of instruction execution information from the L-bit instructions to be processed, and send the L×K2 / (K1+K2) bits of instruction execution information into the instruction buffer 24. Here, L, K1, and K2 are positive integers, and the specific sizes of L, K1, and K2 are not limited in the embodiments of this disclosure.

[0040] The instruction splitting module 21 can be implemented using hardware circuits composed of general digital circuit components. This disclosure does not limit the specific implementation method of the instruction splitting module 21.

[0041] Optionally, the control buffer 22 and the instruction buffer 23 can be implemented using buffers. The control buffer 22 and the instruction buffer 23 have the same bit width, and the depth of the control buffer 22 can be greater than the depth of the instruction buffer 23. Here, the bit width represents the number of bits occupied by each piece of information, and the depth represents the total number of pieces of information to be stored.

[0042] Optionally, the instruction buffer 24 can be implemented using a cache. The depth of the instruction buffer 24 is the same as the depth of the control buffer 22. The sum of the bit width of the instruction buffer 24 and the bit width of the instruction buffer 23 is the same as the bit width of a complete instruction to be processed (i.e., the number of bits of binary code in the instruction to be processed).

[0043] Optionally, the scheduler 25 triggers the execution unit to execute the corresponding task according to the instruction control information and the time or conditions indicated by the control information. The scheduler 25 can determine which execution unit should obtain the execution right and when according to a preset scheduling algorithm (such as first-come, first-served, shortest job first, etc.). The scheduler 25 can be implemented by hardware circuit composed of general digital circuit components, and this disclosure does not limit the specific implementation method of the scheduler 25.

[0044] Thus, from the upper-level instruction cache (e.g.) Figure 2 After retrieving the instruction to be processed from the first-level cache 12), the instruction can be distributed based on the usage characteristics of the instruction at different stages. Specifically, the instruction control information used during the scheduling stage and the instruction execution information used by downstream execution units during the execution stage are distributed through the instruction distribution module 21. The instruction control information used during the scheduling stage is placed in the control buffer 22, and the instruction execution information needed by downstream execution units is placed in the instruction buffer 24. When the instruction scheduling device receives a request, it can read the instruction control information from the control buffer 22 and send it to the instruction buffer 23. This instruction control information participates in the scheduling process of the scheduler 25. For the selected instruction to be processed, the instruction execution data can be read from the instruction buffer 23 and sent to the downstream execution unit. Thus, while performing the same function, the instruction buffer 23 stores only a small amount of instruction control information, while most of the information (such as instruction execution information) is still stored in the instruction buffer 24, effectively reusing the resources of the instruction buffer 24 and saving its hardware resources.

[0045] For example, assuming a 128-bit instruction to be processed, the ratio of instruction control information to instruction execution information carried by the instruction is 1:3, that is, the instruction control information occupies 32 bits, and the instruction execution information occupies 96 bits. It should be understood that the embodiments of this disclosure do not limit the bit width of the instruction to be processed, nor the specific ratio of instruction control information to instruction execution information carried by the instruction; these can be set according to the actual application scenario.

[0046] In related technologies, such as Figure 1 As shown, the instruction buffer 13 has a space size of 1024×128, and can store 1024 instructions of 128 bits each to be processed; the instruction buffer 14 has a space size of 16×128, and can buffer 16 instructions of 128 bits each to be processed. The instruction buffer 14 stores the complete instruction information from the instruction buffer 13, all of which are 128 bits.

[0047] Conversely, in embodiments of this disclosure, such as Figure 2 As shown, the control buffer 22 has a space size of 1024×32, and can cache the instruction control information of 1024 pending instructions, with each instruction control information occupying 32 bits. The instruction buffer 23 has a space size of 16×32, and can cache the instruction control information of 16 pending instructions, with each instruction control information occupying 32 bits. The instruction buffer 24 has a space size of 1024×96, and can store the instruction execution information of 1024 pending instructions.

[0048] contrast Figure 1 and Figure 2 It can be seen that, Figure 2 The hardware storage resources shared by the control buffer 22 and the instruction cache 24 (e.g., 1024×32+1024×96) are... Figure 1 The hardware storage resources (e.g., 1024×128) occupied by the instruction register 13 are roughly the same. However, Figure 2 The hardware storage resources occupied by the instruction buffer 23 (e.g., 16×32) are much smaller than those occupied by the instruction buffer 23. Figure 1 The instruction buffer 14 has hardware storage resources (e.g., 16×128). The instruction buffer 23 stores 25% of the original instructions to be processed, meaning that the instruction buffer 23 can perform the same function using only 25% of the original hardware storage resources. Therefore, the embodiments of this disclosure reduce the implementation area of ​​the instruction buffer 23 and save implementation resources. Furthermore, the instruction scheduling device of the embodiments of this disclosure is not only universal for processing different instruction types, but also universal for processing instructions from different architectures and processor types.

[0049] In one possible implementation, the instruction control information and the instruction execution information are used at different stages. The instruction control information is used in the scheduling stage to distinguish the execution order between tasks, while the instruction execution information is used in the execution stage to provide operation information for the execution units downstream of the scheduler 25. The operation information includes operation data for the executed task, the specific behavior of the executed task, and operations such as writing back the result. For example, the operation information may include an opcode and an address code; the opcode is used to indicate the type or nature of the operation behavior to be performed by the instruction execution information, such as a read operation, a write operation, a logical operation, an arithmetic operation, etc. The address code is used to indicate the content of the operation object or the address of its storage unit.

[0050] The instruction buffer 23 is a thread group resource, and each thread group can store instruction control information for multiple pending instructions. The pending instructions in each thread group are executed sequentially. During the scheduling phase, the scheduler 25 can select one pending instruction from one thread group in the instruction buffer 23 each cycle, and read its corresponding instruction execution information from the instruction cache 24 during the execution phase and send it to the execution unit for execution.

[0051] In this way, by taking advantage of the usage characteristics of instructions at different stages, instruction buffer 23 does not need to store instruction execution information, which reduces the implementation area of ​​instruction buffer 23 and saves implementation resources.

[0052] In one possible implementation, the instruction buffer 23 is further used to record the position information of the instruction control information read from the control buffer 22 in the control buffer 22, and the scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction buffer 24 according to the instruction control information from the instruction buffer 23, including: the scheduler 25 reads the instruction execution information corresponding to the instruction control information from the instruction buffer 24 according to the instruction control information from the instruction buffer 23 and the position information.

[0053] Optionally, the storage locations for instruction control information and instruction execution information are in a one-to-one correspondence. For example... Figure 2As shown, assuming there are 1024 instructions to be processed, the instruction control information of the first instruction to be processed is stored in the first row of the control buffer 22, and the instruction execution information of the first instruction to be processed is stored in the first row of the instruction buffer 24; the instruction control information of the second instruction to be processed is stored in the second row of the control buffer 22, and the instruction execution information of the second instruction to be processed is stored in the second row of the instruction buffer 24; and so on, the instruction control information of the 1024th instruction to be processed is stored in the 1024th row of the control buffer 22, and the instruction execution information of the 1024th instruction to be processed is stored in the 1024th row of the instruction buffer 24.

[0054] In this way, the instruction buffer 23 can read instruction control information and the location information of the instruction control information (e.g., cacheline ID) from the control buffer 22, so that the scheduler 25 can select one or more instruction control information to be executed based on the selection strategy, and read instruction execution information from the instruction cache 24 to the downstream execution unit based on the location information of the selected instruction control information.

[0055] By setting the location information, the instruction execution information corresponding to each instruction control information can be quickly and accurately located from the instruction buffer 24. In this way, the instruction buffer 23 only stores instruction control information and does not need to store instruction execution information. The location information can be used to quickly and accurately locate the instruction execution information corresponding to each instruction control information in the instruction buffer 23, reducing the implementation area of ​​the instruction buffer 23 and saving implementation resources.

[0056] In one possible implementation, the scheduler 25 reads instruction execution information corresponding to the instruction control information from the instruction buffer 24 based on the instruction control information from the instruction buffer 23, including: the scheduler 25 selecting target instruction control information and target location information of the target instruction control information from the instruction control information in the instruction buffer 23 according to a preset selection strategy; and the scheduler 25 reading instruction execution information corresponding to the target instruction control information from the instruction buffer 24 based on the target location information.

[0057] The scheduler 25 may include a Completely Fair Scheduler (CFS), a Real-time Scheduler, a Multiqueue Scheduler, etc. Correspondingly, the preset selection strategies may include strategies for fair selection of all instruction control information, selection strategies that set response time as the highest priority, and selection strategies that achieve load balancing of execution units among multiple instruction control information, etc., which can be set according to the actual application scenario, and the embodiments of this disclosure do not impose limitations on them.

[0058] For example, suppose the scheduler 25 receives N instruction control messages from the instruction buffer 23. It can select one or more target instruction control messages from the N instruction control messages according to a preset selection strategy, as well as the target location information of each target instruction control message. Then, the scheduler 25 can read the instruction execution information corresponding to the target instruction control message of each target instruction control message from the instruction buffer 24 according to the target location information of each target instruction control message.

[0059] The instruction buffer 23 can be a thread group resource, which may include multiple thread groups. Each thread group can store instruction control information for multiple pending instructions. The pending instructions in each thread group can be executed sequentially. During the scheduling phase, the scheduler 25 can select the instruction control information of a pending instruction from each thread group in the instruction buffer 23 as the target instruction control information according to a preset selection strategy each cycle, and record the target location information (e.g., location number or address) of each target instruction control information. Then, the scheduler 25 can read the instruction execution information corresponding to the target instruction control information of each target instruction control information from the instruction buffer 24 according to the target location information of each target instruction control information.

[0060] In this way, the scheduler 25 can quickly identify and select the target instruction control information and its target location information from the instruction buffer 23 using a preset selection strategy. This process reduces instruction processing latency, allowing the instructions to be processed to enter the execution phase more quickly. At the same time, accurately reading instruction execution information from the instruction cache 24 based on the target location information further accelerates the execution process of the instructions to be processed and improves the overall system operating efficiency.

[0061] Furthermore, the preset selection strategy can be adjusted and optimized according to different application scenarios and needs. This allows scheduler 25 to flexibly adapt to various complex computing tasks, improving its adaptability and scalability. Scheduler 25 can further integrate new scheduling algorithms and strategies to meet the needs of future higher-performance and more complex computing tasks.

[0062] Figure 3 A schematic diagram of another instruction scheduling apparatus according to an embodiment of the present disclosure is shown. Figure 3 As shown, the instruction scheduling device includes an instruction splitting module 31, a control buffer 32, an instruction buffer 33, an instruction cache 34, and a scheduler 35. The instruction scheduling device also includes a detection module 36. The detection module 36 is connected to the control buffer 32 and the next-level cache (e.g., the first-level cache 30). The first-level cache 30 is connected to the instruction splitting module 31. The instruction splitting module 31 is also connected to the control buffer 32 and the instruction cache 34. The control buffer 32 is connected to the instruction buffer 33. The instruction buffer 33 and the instruction cache 34 are connected to the scheduler 35.

[0063] It should be understood that Figure 3 The instruction distribution module 31, control buffer 32, instruction buffer 33, instruction register 34, and scheduler 35 can be reused. Figure 2 The instruction distribution module 21, control buffer 22, instruction buffer 23, instruction cache 24, and scheduler 25 are described in detail above and will not be repeated here.

[0064] like Figure 3 As shown, the instruction to be processed is an instruction requesting information. The detection module 36 is used to receive the request information and detect it to obtain a detection result. The detection performed by the detection module 36 includes, for example, a hit-miss check. The detection result is used to characterize the presence status of the instruction control information of the instruction requesting information in the control buffer 32.

[0065] In response to the instruction scheduling device receiving a request message, the detection module 36 can perform detection. If the detection result indicates that the instruction control information of the requested instruction exists in the control buffer 32, the instruction buffer 33 reads the instruction control information from the control buffer 32. For example, in a loop or iterative operation scenario, the same instruction may be executed multiple times. In this case, if the requested instruction is a recently processed instruction, then the instruction control information of that instruction may already exist in the control buffer 32.

[0066] Alternatively, if the instruction control information of the pending instruction requested by the detection result does not exist in the control buffer 32, the instruction splitting module 31 is used to request the pending instruction from the upper-level cache (e.g., level 1 cache 30) according to the request information, and to split the requested pending instruction, sending the instruction control information carried by the pending instruction into the control buffer 32, and sending the instruction execution information carried by the pending instruction into the instruction buffer 34; the instruction buffer 33 reads the instruction control information from the control buffer 32.

[0067] Then, the scheduler 35 can read the instruction execution information corresponding to the instruction control information from the instruction buffer 34 based on the instruction control information from the instruction buffer 33.

[0068] By setting up the detection module 36, it is possible to quickly determine whether the instruction control information of the instruction to be processed exists in the control buffer 32, thereby avoiding unnecessary access or remote data retrieval (such as reading the instruction to be processed from the previous level cache) and improving data access speed.

[0069] In summary, in the instruction scheduling apparatus of this embodiment, in response to the instruction splitting module acquiring a pending instruction, the instruction splitting module performs splitting processing on the pending instruction, sending the instruction control information carried by the pending instruction to a control buffer and sending the instruction execution information carried by the pending instruction to the instruction cache; the instruction buffer reads the instruction control information from the control buffer; and the scheduler reads the instruction execution information corresponding to the instruction control information from the instruction cache based on the instruction control information from the instruction buffer. Thus, by utilizing the usage characteristics of the pending instruction at different stages—that is, using the instruction control information of the pending instruction during the scheduling stage and using the instruction execution information of the pending instruction in the execution stage downstream of the scheduler—there is no need to store instruction execution information in the instruction buffer, reducing the implementation area of ​​the instruction buffer and saving implementation resources. Furthermore, the instruction scheduling apparatus of this embodiment is not only universal for processing different instruction types but also universal for processing instructions from different architectures and processor types.

[0070] It should be noted that, although... Figure 2 and Figure 3 The instruction scheduling device described above is an example, but those skilled in the art will understand that this disclosure is not limited thereto. In fact, users can flexibly configure it according to their personal preferences and / or actual application scenarios.

[0071] In one possible implementation, embodiments of this disclosure provide a processor including the instruction scheduling device described above. The processor may be of a type including, but is not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a General-Purpose Computing on Graphics Processing Units (GPGPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Tensor Processing Unit (TPU), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, and may also include microprocessors, etc. Embodiments of this disclosure do not limit this.

[0072] In one possible implementation, embodiments of this disclosure provide a chip that includes the processor described above.

[0073] In one possible implementation, embodiments of this disclosure provide an electronic device including the chip described above. The electronic device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc., and embodiments of this disclosure do not limit this.

[0074] Figure 4 This is a block diagram illustrating an electronic device 1900 according to an exemplary embodiment. For example, the electronic device 1900 may be provided as a server or a terminal device. (Refer to...) Figure 4 The electronic device 1900 includes a processing component 1922, which further includes one or more processors (e.g., including the instruction scheduling means described above), and memory resources represented by memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0075] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). Electronic device 1900 can operate on an operating system stored in memory 1932, such as Microsoft Windows Server™, Apple's graphical user interface-based operating system (MacOS X™), a multi-user, multi-process computer operating system (Unix™), a free and open-source Unix-like operating system (Linux™), an open-source Unix-like operating system (FreeBSD™), or similar.

[0076] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0078] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A command scheduling device, characterized in that, The instruction scheduling device includes an instruction splitting module, a control buffer, an instruction buffer, an instruction cache, and a scheduler. The instruction splitting module is connected to the control buffer and the instruction cache. The control buffer is connected to the instruction buffer. The instruction buffer and the instruction cache are connected to the scheduler. The instruction splitting module splits the acquired instructions to be processed, sends the instruction control information carried by the instructions to be processed into the control buffer, and sends the instruction execution information carried by the instructions to be processed into the instruction buffer. The instruction buffer reads instruction control information from the control buffer; The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information from the instruction cache.

2. The instruction scheduling device according to claim 1, characterized in that, The instruction to be processed is an instruction requesting information. The instruction scheduling device further includes a detection module, which is connected to the control buffer. The detection module is used to receive the request information and detect the request information to obtain a detection result. The detection result is used to characterize the existence status of the instruction control information of the instruction to be processed by the requesting information in the control buffer.

3. The instruction scheduling device according to claim 2, characterized in that, If the instruction control information representing the pending instruction of the request information is present in the control buffer, the instruction buffer reads the instruction control information from the control buffer. Alternatively, if the detection result indicates that the instruction control information of the pending instruction requested by the request information does not exist in the control buffer, the instruction splitting module is used to request the pending instruction from the upper-level cache according to the request information, and to split the requested pending instruction by sending the instruction control information carried by the pending instruction into the control buffer and the instruction execution information carried by the pending instruction into the instruction buffer; the instruction buffer reads the instruction control information from the control buffer.

4. The instruction scheduling device according to any one of claims 1 to 3, characterized in that, The instruction buffer is further used to record the position information of the instruction control information read from the control buffer within the control buffer. The scheduler, based on the instruction control information from the instruction buffer, reads instruction execution information corresponding to the instruction control information from the instruction buffer, including: The scheduler reads the instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information and the position information.

5. The instruction scheduling device according to any one of claims 1 to 3, characterized in that, The instruction control information and the instruction execution information are used at different stages.

6. The instruction scheduling device according to claim 5, characterized in that, The instruction control information is used in the scheduling phase to distinguish the execution order between tasks, and the instruction execution information is used in the execution phase to provide the execution units downstream of the scheduler with operation information for executing tasks.

7. The instruction scheduling device according to claim 4, characterized in that, The scheduler reads instruction execution information corresponding to the instruction control information from the instruction buffer based on the instruction control information from the instruction cache, including: The scheduler selects target instruction control information and target location information from the instruction control information in the instruction buffer according to a preset selection strategy. The scheduler reads the instruction execution information corresponding to the target instruction control information from the instruction buffer based on the target location information.

8. A processor, characterized in that, The processor includes the instruction scheduling device as described in any one of claims 1 to 7.

9. A chip, characterized in that, The chip includes the processor as described in claim 8.

10. An electronic device, characterized in that, The electronic device includes the chip as described in claim 9.

Citation Information

Patent Citations

  • Decoupled processor instruction window and operand buffer

    CN107810476A

  • Method and apparatus for constructing a pre-scheduled instruction cache

    SG89191A1