Instruction execution module for use in processor, chip, device, and method

By converging multiple memory fetch instructions into merge instructions in the dispatch queue and executing them in the pipeline, the problem of micro instruction bandwidth limitation is solved, efficient execution of processor memory fetch instructions is achieved, and processor performance is improved.

WO2025140132A1PCT designated stage expired Publication Date: 2025-07-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141553
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-12-23
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, due to the expansion of micro instruction bandwidth, the execution efficiency of the pipeline to access instructions is difficult to improve, resulting in limited processor performance.

Method used

Add new execution logic to the dispatch queue, integrate multiple memory fetch instructions into a merge instruction, and execute the merge instruction through the pipeline to realize the synchronous execution of multiple memory fetch instructions, expanding the scope and efficiency of instruction merging.

Benefits of technology

By combining instructions, the pipeline can execute multiple memory access instructions in one clock cycle, which improves the execution efficiency of the instruction execution module to perform memory access instructions and improves the performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141553_03072025_PF_FP_ABST
    Figure CN2024141553_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of chips, and discloses an instruction execution module for use in a processor, a chip, a device, and a method. The instruction execution module comprises: a dispatch queue and a pipeline. The dispatch queue is used for: storing at least one memory access instruction to be executed; selecting n memory access instructions from among the at least one memory access instruction to be executed, n being a positive integer greater than 1; and fusing the n memory access instructions to obtain a first merged instruction, functions of the first merged instruction being equivalent to functions of the n memory access instructions. The pipeline is used for executing the first merged instruction. The memory access instructions stored in the dispatch queue are merged to obtain a merged instruction, which is beneficial for increasing the number of memory access instructions comprised in an individual instruction, and is beneficial for improving the efficiency of the instruction execution module executing the memory access instructions when the transmission bandwidth of the memory access instructions is limited, thereby improving the performance of the instruction execution module.
Need to check novelty before this filing date? Find Prior Art

Description

Instruction execution module, chip, device and method applied to processor

[0001] This application claims priority to the Chinese patent application filed on December 25, 2023, with application number 202311809091.9 and invention name “Instruction execution module, chip, device and method for processor”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of chip technology, and in particular to an instruction execution module, chip, device, and method applied to a processor. Background Art

[0003] Memory access instructions are used to write or retrieve instruction data from memory. A large number of memory access instructions appear during the processor's program execution process. The speed at which the processor processes these instructions directly affects the processor's performance.

[0004] In related art, for a microinstruction bandwidth, a pipeline can fetch one memory access instruction within one clock cycle. Specifically, a dispatch queue dispatches the memory access instruction to an issue queue, which then issues the memory access instruction to the pipeline. The pipeline receives and executes one memory access instruction within one clock cycle.

[0005] However, in related technologies, due to the limitation of the expansion of microinstruction bandwidth, it is difficult to improve the execution efficiency of the pipeline for memory access instructions. Summary of the Invention

[0006] The present application provides an instruction execution module, chip, device, and method for a processor. The technical solutions provided in the embodiments of the present application are as follows.

[0007] According to one aspect of an embodiment of the present application, there is provided an instruction execution module applied to a processor, the instruction execution module comprising: a dispatch queue and a pipeline;

[0008] The dispatch queue is configured to store at least one memory access instruction to be executed; n memory access instructions are selected from the at least one memory access instruction to be executed, where n is an integer greater than 1; and the n memory access instructions are merged into one instruction to obtain a first merged instruction, wherein an effect of the first merged instruction is equivalent to an effect of the n memory access instructions;

[0009] The pipeline is used to execute the first merge instruction.

[0010] According to one aspect of an embodiment of the present application, a chip is provided, comprising a processor and the instruction execution module as described above.

[0011] According to one aspect of an embodiment of the present application, an electronic device is provided. The electronic device includes a chip, and the chip includes the instruction execution module as described above.

[0012] According to one aspect of an embodiment of the present application, there is provided an instruction execution method applied to an instruction execution module, wherein the instruction execution module includes: an allocation queue and a pipeline, and the method includes:

[0013] The dispatch queue stores at least one memory access instruction to be executed;

[0014] The dispatch queue selects n memory access instructions from the at least one memory access instruction to be executed, where n is an integer greater than 1;

[0015] The dispatch queue merges the n memory access instructions into one instruction to obtain a first merged instruction, where the effect of the first merged instruction is equivalent to the effect of the n memory access instructions;

[0016] The pipeline executes the first merge instruction.

[0017] The embodiment of the present application provides an instruction execution module that can achieve the following beneficial effects: by adding new execution logic to the dispatch queue, the effect of fusing multiple memory access instructions stored in the dispatch queue into a merged instruction is achieved. The merged instruction is executed by the pipeline to achieve the same effect as executing the above-mentioned multiple memory access instructions. Relying on the relatively wide range of instruction merging provided in the dispatch queue, it helps to merge a large number of memory access instructions into one merged instruction, and only transmits one merged instruction to achieve the effect of transmitting the above-mentioned large number of memory access instructions, so that the pipeline obtains multiple memory access instructions in one clock cycle, which helps to realize the execution of multiple memory access instructions by the instruction execution module in one clock cycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG1 is a schematic diagram of the architecture of an instruction execution module provided by a related art;

[0019] FIG2 is a schematic diagram of a binary disassembly file provided by an exemplary embodiment of the present application;

[0020] FIG3 is a schematic diagram of an inventive concept provided by an exemplary embodiment of the present application;

[0021] FIG4 is a schematic diagram of an instruction execution module applied to a processor provided by an exemplary embodiment of the present application;

[0022] FIG5 is a schematic diagram of a binary disassembly file provided by another exemplary embodiment of the present application;

[0023] FIG6 is a schematic diagram of an instruction merging process provided by an exemplary embodiment of the present application;

[0024] FIG7 is a schematic diagram of an instruction re-issuing process provided by an exemplary embodiment of the present application;

[0025] FIG8 is a schematic diagram of an instruction re-issuing process provided by another exemplary embodiment of the present application;

[0026] FIG9 is a schematic diagram of an instruction execution method applied to an instruction execution module provided by an exemplary embodiment of the present application;

[0027] FIG10 is a schematic diagram of an instruction execution method applied to an instruction execution module provided by another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0029] First, the terms involved in the embodiments of this application are explained.

[0030] Memory access instructions are instructions used to store or access data. They are also called load / store instructions. A memory access instruction consists of a register identifier and a memory address. The register identifier identifies the register used during the execution of the memory access instruction, and the memory address indicates the storage address in the memory space where the data being operated on by the memory access instruction is stored.

[0031] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0032] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0033] Artificial intelligence technology often involves a large amount of calculations. The instruction execution module provided in this application can be used to perform operations involved in artificial intelligence technology and promote the development and progress of artificial intelligence technology.

[0034] FIG1 is a schematic diagram of the architecture of an instruction execution module provided by a related art.

[0035] As shown in Figure 1 , the instruction execution module primarily includes a decoder, a map and renamer, a DQ (Dispatch Queue), a scheduler, and pipelines (including the different types of ALUs (Arithmetic and Logic Units), load pipelines, and store pipelines shown in Figure 1 ). This architecture diagram includes multiple types of DQs, each used to store different types of instructions. For example, Load / Store instructions enter the Load Store Dispatch Queue (LSDQ) 100. A DQ includes at least one entry, each used to store one instruction. As shown in Figure 1 , the LSDQ 100 has 10 entries, meaning it can store up to 10 memory access instructions.

[0036] In the architecture shown in Figure 1, units such as the decoder perform instruction fetch and decoding to generate instructions. The generated instructions are then converted from logical registers to physical registers through a map and rename process. Subsequently, the instructions are transferred to the corresponding DQ based on their type. The DQ then dispatches the instructions to the corresponding scheduler, which then sends them to the corresponding pipeline for execution.

[0037] Optionally, the decoder in FIG1 can decode six microinstructions in one clock cycle, i.e., the decoder is a six-way decoder. An instruction fusion module is provided next to the decoder. The instruction fusion module is used to merge the instructions decoded by the decoder in one clock cycle.

[0038] Since the decoder decodes 6 instructions in one clock cycle, the instruction fusion module can fuse up to 6 instructions. Moreover, due to the limitation of instruction type, the instruction fusion module can actually only merge two adjacent instructions. For example, the lui instruction (used to set the high 16-bit value of a constant in a register) and the load instruction. It can be seen that the architecture of the instruction execution module needs to be improved to improve the effect of the instruction merging process, thereby improving the efficiency of the pipeline's execution of instructions.

[0039] Figure 2 is a schematic diagram of a binary disassembly file provided by an exemplary embodiment of the present application. As shown in Figure 2, during operation, the processor needs to process a large number of sequentially executed memory access instructions 210. Limited by the microinstruction bandwidth, the pipeline can only receive a limited number of memory access instructions in a single clock cycle, which affects pipeline performance. Compression and packaging consecutive memory access instructions into a single merged instruction can reduce the pressure on the pipeline to transmit memory access instructions, thereby improving the pipeline's efficiency in executing memory access instructions.

[0040] FIG3 is a schematic diagram of the inventive concept of the present application.

[0041] In the instruction execution module provided in the embodiment of the present application, a new execution logic is added to the dispatch queue, and the newly added execution logic is used to merge multiple memory access instructions stored in the dispatch queue to obtain a merged instruction. The merged instruction can be transmitted through the transmission path of the ordinary memory access instruction, so that after the pipeline receives the merged instruction, it can determine the instruction content of the above-mentioned multiple memory access instructions based on the merged instruction. That is, based on the one merged instruction, the effect of synchronously executing the above-mentioned multiple memory access instructions in one clock cycle is achieved, which saves the bandwidth consumed by transmitting multiple memory access instructions, improves the speed of the pipeline in obtaining memory access instructions, reduces the clock cycle that the pipeline needs to wait for to obtain memory access instructions, and helps to improve the execution efficiency of memory access instructions.

[0042] Since the pipeline may encounter instruction execution exceptions during the instruction execution process, and redirection and other operations are required between instructions with timing violations, the decoder generates multiple instructions in one clock cycle and cannot be immediately executed in the subsequent clock cycle, resulting in a pile-up of memory access instructions in the dispatch queue. Therefore, the number of entries in the dispatch queue used to store memory access instructions needs to be greater than or equal to the number of decoders generated by the decoder in one clock cycle. Instruction merging in the dispatch queue provides a wide instruction selection window for the instruction merging process, increases the range of memory access instructions that can be selected during the instruction merging process, helps to merge more memory access instructions into one merged instruction, and helps the pipeline obtain more memory access instructions in one clock cycle, thereby further improving the pipeline's processing efficiency for memory access instructions.

[0043] The instruction execution module provided in the embodiments of this application can be a complete high-performance processor, or a portion of a high-performance processor encapsulated in a chip. The chip can be an AI chip for model training, a chip for image processing, a chip for video processing, etc. This application does not limit the type of chip used by the instruction execution module.

[0044] The chip corresponding to the instruction execution module provided in the embodiment of the present application can also be applied to ITS (Intelligent Traffic System). Intelligent transportation system, also known as intelligent transportation system (Intelligent Transportation System), is the effective and comprehensive application of advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) to transportation, service control and vehicle manufacturing, strengthening the connection between vehicles, roads and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment and saves energy. For example, the instruction execution device provided in the embodiment of the present application can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, assisted driving, etc., to improve the speed at which the processor executes memory access instructions in the intelligent transportation system, thereby prompting the data processing performance of the processor.

[0045] From the perspective of CPU (Central Processing Unit) design, the most critical part affecting CPU performance is the processing of memory access instructions, that is, the design of the memory access unit. The efficiency of the execution of memory access instructions directly determines the performance of the CPU. In the CPU design process, for the design scheme using Dispatch Queue, further compression function of memory access instructions can be implemented based on Dispatch Queue. Compared with the design scheme of performing memory access instruction fusion in the decoder, the instruction execution module provided in the embodiment of the present application can be used as a supplement to further improve CPU performance.

[0046] FIG4 is a schematic diagram of an instruction execution module applied to a processor according to an exemplary embodiment of the present application. The instruction execution module 400 includes: a dispatch queue 410 and a pipeline 430;

[0047] The dispatch queue 410 is used to store at least one memory access instruction to be executed; select n memory access instructions from the at least one memory access instruction to be executed, where n is an integer greater than 1; merge the n memory access instructions into one instruction to obtain a first merged instruction, and the effect of the first merged instruction is equivalent to the effect of the n memory access instructions.

[0048] In some embodiments, a memory access instruction refers to an instruction for implementing data loading or data storage. Optionally, the types of memory access instructions include load instructions and store instructions. The load instruction is used to obtain data stored in the memory access address, and the store instruction is used to store data in the storage space corresponding to the memory access address.

[0049] The instructions that can be fused in the embodiments of the present application are conventional memory access instructions, such as LB (Load Byte), LW (Load Word), LH (Load Half-word), and LD (Load Double-word) types in load instructions. Memory access instructions used to implement mechanisms such as lock grabbing are not suitable for instruction merging.

[0050] Exemplarily, the memory access instruction is generated by decoding the instruction data stored in the instruction buffer by the decoder in the processor and performing renaming processing through the renaming module. After completing the renaming of the instruction, the renaming module sends the instruction to the corresponding dispatch queue according to the type of the instruction.

[0051] Optionally, the instruction content of the memory access instruction includes: an instruction identifier, a register number, and a memory access address, wherein the instruction identifier is used to indicate the type of the memory access instruction, and the register number is used to represent the register used when the pipeline executes the memory access instruction.

[0052] Exemplarily, the register number is the number of a physical register. A physical register refers to a register that actually exists in the processor. During the renaming process, the renaming module maps the logical registers in the memory access instruction to physical registers. Logical registers refer to the callable registers provided to software in the instruction set. A mapping relationship exists between logical registers and physical registers, and the renaming module stores this mapping relationship.

[0053] Optionally, the memory access instruction belongs to a uop (Micro Operation).

[0054] In some embodiments, the dispatch queue 410 is used to store at least one memory access instruction to be executed. The dispatch queue is also called a scheduling queue, an allocation queue, etc. Optionally, the dispatch queue 410 is used to store at least two memory access instructions to be executed.

[0055] Optionally, the dispatch queue only stores memory access instructions and does not store other types of instructions. Exemplarily, after completing the instruction renaming process, the renaming module sends the instruction to the corresponding dispatch queue based on the instruction type. For example, memory access instructions are sent to dispatch queue 410, while arithmetic instructions are allocated to other dispatch queues of the processor. In this case, dispatch queue 410 can be an LSDQ.

[0056] Optionally, both memory access instructions and operation instructions are stored in the dispatch queue 410. The instruction type stored in the dispatch queue 410 is set according to actual needs and is not limited in this application.

[0057] In some embodiments, a pending memory access instruction refers to a memory access instruction that has not been executed by the pipeline. The pending memory access instruction is waiting to be sent to the pipeline for execution. Optionally, the dispatch queue 410 determines the storage entry of the pending memory access instruction in the dispatch queue 410 according to the execution timing of at least one pending memory access instruction. The execution timing of the memory access instruction refers to the order of the memory access instruction in the instruction stream, and the instruction stream refers to the serial arrangement order of each instruction when executing each instruction sequentially. Instructions earlier in the instruction stream are executed before instructions later in the instruction stream.

[0058] Furthermore, by adding new execution logic to the dispatch queue 410 , n memory access instructions can be selected from at least one memory access instruction to be executed stored in the dispatch queue, and the n memory access instructions can be merged into one instruction to obtain a merged instruction.

[0059] Optionally, the dispatch queue 410 includes s entries for storing memory access instructions, where s is a positive integer. For example, s is equal to 16 or 32. s is also called the capacity of the dispatch queue. The larger the value of s, on the one hand, the more memory access instructions to be executed can be accommodated in the dispatch queue 410, the wider the merging range when performing instruction merging, and the more memory access instructions can be selected; on the other hand, the larger the value of s, the longer it takes to select n memory access instructions from the dispatch queue 410, and the larger the volume occupied by the dispatch queue 410 in the hardware. The value of s is determined by compromise based on actual performance requirements, and the specific value of s is not limited here.

[0060] In order to avoid pipeline blockage that causes the upstream decoder to be unable to continue decoding and generate new instructions, the capacity of the dispatch queue is usually greater than the number of instructions generated by the decoder in one clock cycle. That is, the memory access instructions to be executed stored in the dispatch queue are unlocked by the decoder in different clock cycles. Therefore, compared to performing instruction merging in the decoder, the embodiment of the present application performs instruction merging in the dispatch queue, which improves the scope of instruction merging, so that the instruction merging process has more optional memory access instructions, helps to expand the field of view of the instruction merging process, and helps to represent more memory access instructions through merging instructions.

[0061] Optionally, n is a positive integer greater than or equal to 1. For example, n is equal to 6, that is, the first merge instruction includes 6 memory access instructions. Exemplarily, the value of n changes dynamically. In any clock cycle during the operation of the instruction execution module, at least one of the following situations will occur: the memory access instruction to be executed leaves the dispatch queue 410, and a new memory access instruction to be executed is added to the dispatch queue 410. It can be seen that the at least one memory access instruction to be executed stored in the dispatch queue 410 changes dynamically. Therefore, in different clock cycles, the number of memory access instructions selected from at least one memory access instruction to be executed that is suitable for generating a merge instruction will also change.

[0062] In one example, n has a maximum value and a minimum value. That is, the number of memory access instructions selected from the at least one memory access instruction is no greater than the maximum value of n and no less than the minimum value of n. For example, the maximum value is 6 and the minimum value is 2.

[0063] In some cases, the maximum value of n is affected by the instruction merging method. To ensure that dispatch queue 410 can use the transmission path of one memory access instruction to dispatch the first merged instruction, the first merged instruction does not exceed the bandwidth used to transmit the memory access instruction. In other words, the maximum value of n is used to ensure that the generated merged instruction does not exceed the microinstruction bandwidth used to transmit one memory access instruction. In addition, the maximum value of n is affected by the pipeline's ability to process memory access instructions. The maximum value of n is less than or equal to the upper limit of the memory access instructions that the pipeline can execute in one clock cycle.

[0064] Optionally, the n memory access instructions are of the same instruction type. For example, the n memory access instructions are all load instructions. For another example, the n memory access instructions are all store instructions.

[0065] Optionally, the n memory access instructions are n memory access instructions whose execution sequence is continuous among at least one memory access instruction to be executed. In one example, the execution sequence of the memory access instructions can be represented by instruction numbers, where a smaller number indicates an earlier execution sequence of the memory access instruction and a larger number indicates a later execution sequence of the memory access instruction. The n memory access instructions are numbered [0, n-1].

[0066] In some embodiments, the first merge instruction is used to represent the instruction content of n memory access instructions. In other words, the instruction content of the n memory access instructions can be determined based on the first merge instruction. Issuing the first merge instruction to pipeline 430 is equivalent to issuing n memory access instructions to pipeline 430. The difference between the two is that issuing the first merge instruction to pipeline 430 only requires 1 microinstruction bandwidth within one clock cycle, while issuing n memory access instructions separately to pipeline 430 requires a maximum of n clock cycles. Merging n memory access instructions improves the efficiency of the pipeline in obtaining memory access instructions.

[0067] Instruction fusion is equivalent to compressing and packaging n memory access instructions so that the first merged instruction obtained can represent the instruction content of the n memory access instructions, so as to achieve the effect of the pipeline executing n memory access instructions in one clock cycle.

[0068] Optionally, in the process of fusing n memory access instructions into one instruction, the dispatch queue 410 generates the instruction content of a first merged instruction based on the instruction content of the n memory access instructions, thereby obtaining the first merged instruction. Subsequently, the first merged instruction is emitted to the pipeline 430 instead of emitting n memory access instructions to the pipeline 430 separately. Exemplarily, the data path used when emitting the first merged instruction is the same as the data path used when emitting a single memory access instruction. Therefore, this architecture has little impact on other parts of the existing instruction execution architecture and can directly improve the architecture of the existing instruction execution module when making major changes.

[0069] In some embodiments, the instruction execution module further includes a reservation station (RS). The reservation station is used to transfer pending instructions between the dispatch queue and the pipeline. For example, the reservation station receives a pending instruction dispatched by the dispatch queue, determines the pipeline 430 to execute the pending instruction, and determines the clock cycle at which to transmit the pending instruction to pipeline 430. The reservation station transmits the pending instruction to pipeline 430, which then executes the pending instruction.

[0070] Optionally, the type of instruction to be executed includes at least one of the following: a merge instruction, or other memory access instructions other than the n memory access instructions in the at least one memory access instruction to be executed. A merge instruction is obtained by fusing multiple memory access instructions. For example, the first merge instruction in the above embodiment is a merge instruction.

[0071] The reservation station is also called a scheduler, etc. The reservation station in the embodiment of the present application is used to complete the function of issuing memory access instructions to the pipeline. The reservation station is represented by different names in different processor types, and this application does not explain these names one by one.

[0072] It should be noted that if n memory access instructions are merged to generate a first merge instruction, the dispatch queue 410 only needs to send the first merge instruction to the reservation station, and will not send n memory access instructions to the reservation station separately, so as to avoid any of the n memory access instructions being executed repeatedly, resulting in errors in the program running results.

[0073] Because n of the n memory access instructions has a maximum limit, and the n memory access instructions selected from at least one pending memory access instruction must meet the instruction merging condition, at least one pending memory access instruction may contain memory access instructions that cannot be merged. These memory access instructions are referred to as "other memory access instructions." Dispatch queue 410 must dispatch these other memory access instructions separately. That is, these other memory access instructions are transmitted to the pipeline using a single microinstruction bandwidth.

[0074] Optionally, dispatch queue 410 is coupled to a reservation station, which is coupled to pipeline 430. Exemplarily, the port width between dispatch queue 410 and the reservation station is a microinstructions, indicating that dispatch queue 410 can send a maximum of a microinstructions to the reservation station in one clock cycle. a is a positive integer, such as 4.

[0075] As described above, memory access instructions are microinstructions, and the size of the merged instruction generated by instruction merging does not exceed one microinstruction. Therefore, within a clock cycle, dispatch queue 410 dispatches at most a merged instructions to the reservation station and at least one other memory access instruction, or no pending instructions. For example, if a is 4, dispatch queue 410 dispatches two merged instructions and one memory access instruction to the reservation station.

[0076] In some embodiments, the instruction execution module includes at least one reservation station, and different reservation stations are used to issue different types of memory access instructions. Optionally, the instruction execution module includes at least a load instruction reservation station (Load RS) and a store instruction reservation station (Store RS).

[0077] The load reservation station is used to issue load instructions, and the store reservation station is used to issue store instructions. For example, if the first merge instruction includes n load instructions, the dispatch queue 410 dispatches the first merge instruction to the load reservation station; if the first merge instruction includes n store instructions, the dispatch queue 410 dispatches the first merge instruction to the store reservation station.

[0078] Exemplarily, the port width between the reservation station and pipeline 430 is b microinstructions, indicating that the reservation station can send a maximum of b microinstructions to pipeline 430 in one clock cycle. b is a positive integer, such as 4. For example, the reservation station can send one merge instruction and two memory access instructions to the pipeline in one clock cycle.

[0079] Pipeline 430 is used to execute the first merge instruction.

[0080] In some embodiments, the pipeline 430 is used to execute the first merge instruction to achieve synchronous execution of n memory access instructions. The first pipeline 430 can also be used to execute other memory access instructions.

[0081] The synchronous execution of n memory access instructions can be understood as follows: within a clock cycle, the pipeline obtains the instruction content of n memory access instructions according to the first merge instruction, and executes the n access instructions sequentially according to the execution timing of the n memory access instructions. Compared to the related art, in which a microinstruction path can only issue one memory access instruction to the same pipeline 430 per clock cycle, the embodiment of the present application issues the first merge instruction to pipeline 430, allowing pipeline 430 to obtain n memory access instructions in the same clock cycle, thereby achieving the synchronous execution of n memory access instructions in the same clock cycle.

[0082] In some embodiments, the instruction execution module 400 includes multiple pipelines, wherein the multiple pipelines include at least two types of pipelines, and the two types of pipelines are respectively used to execute different types of memory access instructions. Optionally, the multiple pipelines include a load instruction pipeline (Load Pipeline) for executing load instructions and a store instruction pipeline (Store Pipeline) for executing store instructions.

[0083] The instruction execution module 400 may also include a calculation pipeline, etc., which is not limited in this application.

[0084] When the n memory access instructions are load instructions, pipeline 430 is a load pipeline, and the load reservation station sends the first merge instruction to the load pipeline. When the n memory access instructions are store instructions, pipeline 430 is a store pipeline, and the store reservation station sends the first merge instruction to the store pipeline.

[0085] FIG5 is a schematic diagram of a binary disassembly file provided by an exemplary embodiment of the present application.

[0086] From the binary disassembly file returned by the performance testing tool (Dhrystone Benchmark), it can be seen that during the processor's function call process, a large number of stack push and pop operations occur. This means that there are a large number of consecutive memory access instructions (such as the consecutive store instructions 510 in Figure 5). Statistics show that approximately 36% of the memory access instructions in some existing CPUs are consecutive. This shows that the execution efficiency of memory access instructions has a crucial impact on the processor's task performance.

[0087] The instruction execution module provided by the embodiments of the present application has a wider range of optional memory access instructions during the instruction merging process. Without affecting the working cycle of the pipeline, multiple memory access instructions can be merged into a single merged instruction as much as possible, reducing the limitation of the microinstruction port width on the transmission of multiple memory access instructions to the pipeline. By transmitting a single merged instruction to the pipeline, the pipeline can obtain multiple instruction contents in a single clock cycle, which helps improve the efficiency of the pipeline in obtaining and executing memory access instructions.

[0088] In summary, the embodiment of the present application provides an instruction execution module, which achieves the effect of fusing multiple storage instructions stored in the dispatch queue into a merged instruction by adding new execution logic to the dispatch queue. The merged instruction is executed by the pipeline to achieve the same effect as executing the above-mentioned multiple storage instructions. Relying on the relatively wide range of instruction merging provided in the dispatch queue, it helps to merge a large number of memory access instructions into one merged instruction, and only transmits one merged instruction to achieve the effect of transmitting the above-mentioned large number of memory access instructions, so that the pipeline obtains multiple memory access instructions in one clock cycle, which helps to realize the execution of multiple memory access instructions by the pipeline in one clock cycle, thereby improving the execution efficiency of the instruction execution module for memory access instructions.

[0089] The functional components of the dispatch queue are introduced and explained through several embodiments below.

[0090] In some embodiments, a dispatch queue includes: a first buffer and an instruction fusion unit; the first buffer is used to store at least one memory access instruction to be executed; the instruction fusion unit is used to select n memory access instructions with continuous execution sequence from at least one memory access instruction to be executed based on an instruction merging condition, as n memory access instructions, wherein the instruction merging condition includes that the instruction types of the n memory access instructions are the same; the instruction fusion unit is further used to fuse the n memory access instructions into one instruction according to the instruction contents of the n memory access instructions to obtain a first merged instruction.

[0091] In some embodiments, the instruction type includes at least one of the following: a load instruction, a store instruction.

[0092] In some embodiments, at least one pending memory access instruction is stored in the first buffer. Optionally, the first buffer stores the at least one pending memory access instruction in a queue data structure. That is, for pending memory access instructions obtained earlier in the dispatch queue, these pending memory access instructions are stored at the head of the first buffer, and for pending memory access instructions obtained later in the dispatch queue, these pending memory access instructions are stored at the tail of the first buffer. Optionally, the dispatch queue prioritizes dispatching the pending memory access instructions at the head of the queue.

[0093] Optionally, the first buffer is an existing functional unit in the dispatch queue. The newly added execution logic in the dispatch queue is embodied in the instruction fusion unit. That is, by adding the instruction fusion unit to the dispatch queue in the related art, n merge instructions from at least one memory access instruction stored in the dispatch queue can be merged to generate a first merge instruction.

[0094] The instruction fusion unit is configured to select n memory access instructions and generate a first merged instruction. Optionally, the instruction fusion unit is coupled to a first buffer. Exemplarily, the instruction fusion unit determines, through the first buffer, the instruction content of at least one pending memory access instruction; based on the instruction content, determines whether the at least one pending memory access instruction complies with an instruction merge condition, thereby selecting n memory access instructions from the at least one pending memory access instruction and executing a process to generate the first merged instruction based on the n memory access instructions.

[0095] To ensure successful execution of the first merged instruction and reduce the number of exceptions that occur during instruction execution, an instruction merge condition must be followed when determining n instructions from at least one pending memory access instruction. The instruction merge condition is a basis for selecting n instructions from the at least one memory access instruction stored in the first buffer.

[0096] In some embodiments, the instruction merging condition includes at least one condition where the n memory access instructions are of the same instruction type. For example, the n instructions are n load instructions that are executed sequentially. For another example, the n instructions are n store instructions that are executed sequentially. In other words, the n memory access instructions are n memory access instructions of the same type that are executed sequentially.

[0097] As can be seen from the above embodiments, the instruction execution module may include a load pipeline and a store pipeline, and the load instructions and store instructions must be executed in different pipelines. When selecting multiple memory access instructions based on the instruction merging condition, it is ensured that the selected multiple memory access instructions are of the same instruction type. This allows the merged instruction obtained by fusing the multiple memory access instructions to be successfully executed by the same pipeline, helping to reduce the number of exceptions that occur during the execution of the merged instructions.

[0098] In some embodiments, the instruction merging condition further includes at least one of the following: the labels of the physical registers of the n memory access instructions are continuous; among the n memory access instructions, there is a stable offset between the memory access addresses of any two memory access instructions with adjacent execution sequences.

[0099] Optionally, the existence of a stable offset between the memory addresses of any two memory access instructions executed sequentially adjacently means that the memory addresses of the n memory access instructions are continuous. Optionally, the stable offset means that the offsets between the memory addresses of any two memory access instructions executed sequentially adjacently are equal.

[0100] For example, for memory access instruction 1, memory access instruction 2 and memory access instruction 3 among n memory access instructions, the memory access address of memory access instruction 1 is 200 (sp), the memory access address of memory access instruction 2 is 192 (sp), and the memory access address of memory access instruction 3 is 184 (sp); then the offset of the memory access address of memory access instruction 2 relative to the memory access address of memory access instruction 1 is 8, and the offset of the memory access address of memory access instruction 3 relative to the memory access address of memory access instruction 2 is 8, that is, there is a stable offset between the memory access addresses of memory access instruction 1, memory access instruction 2 and memory access instruction 3.

[0101] In other embodiments, among the n memory access instructions, there is a first memory access instruction and a second memory access instruction; the execution timing of the first memory access instruction and the second memory access instruction is continuous, and the offset between the memory access address of the first memory access instruction and the memory access address of the second memory access instruction is not equal to the stable offset. In this case, an exception may occur during the pipeline execution of the first merge instruction. The instruction execution module also includes a mechanism for handling such an exception. For details on this process, please refer to the embodiments below.

[0102] Optionally, the label of the physical register is used to represent the physical register of the memory access instruction. The physical register of the memory access instruction refers to the physical register used in the process of executing the memory access instruction. For example, when the memory access instruction is a load instruction, the physical register of the memory access instruction is used to store data read from the memory access address. For another example, when the memory access instruction is a store instruction, the physical register of the store instruction is used to store data that needs to be stored in the memory access address, that is, the data in the physical register of the store instruction is stored in the storage area corresponding to the memory access address.

[0103] It should be noted that the physical register number of the memory access instruction is determined by the reservation station. For the specific method of determining the physical register of the memory access instruction, please refer to the relevant technology and it is not limited here.

[0104] In some embodiments, the instruction fusion unit includes: at least one level of comparator and an instruction generator; for the first level comparator in the at least one level of comparator, the kth comparator in the first level comparator is used to determine the comparison result of the kth comparator according to the memory access instructions stored in the 2kth entry and the 2k+1th entry in the first buffer, wherein the entries of the first buffer are used to store memory access instructions, k is a positive integer less than or equal to 0.5*s, s is the total number of entries included in the first buffer, and the comparison result of the kth comparator is used to characterize whether the memory access instructions stored in the 2kth entry and the 2k+1th entry meet the instruction fusion condition; if there is an i-th level comparator in the at least one level of comparator, i is a positive integer greater than 1, then the j-th comparator in the i-th level comparator is used to determine the comparison result of the j-th comparator according to the comparison result of the 2j-1th comparator in the i-1th level comparator and the comparison result of the 2jth comparator, and the comparison result of the j-th comparator is used to characterize the (j-1)*2 i-1 +1 entry to j*2 i-1 The memory access instructions stored in the entries are respectively compared to see whether the instruction merging condition is met; an instruction generator is used to select n memory access instructions from at least one memory access instruction according to the comparison results of at least one level of comparator; and the n memory access instructions are merged according to the instruction contents of the n memory access instructions to obtain a first merged instruction.

[0105] In some embodiments, the comparator is a combinational logic for determining whether a memory access instruction complies with an instruction merge condition. Optionally, a comparator is used to compare a limited number of memory access instructions. For example, one comparator is used to compare whether two adjacent memory access instructions comply with the instruction merge condition.

[0106] Optionally, the comparator has two inputs and one output. In an embodiment of the present application, the two inputs of the comparator include instruction types of two memory access instructions, and the output of the comparator is whether the two memory access instructions meet the instruction merging condition.

[0107] The conditions for instruction merging are: met and not met. If two memory access instructions meet the instruction merging conditions, then the two memory access instructions cannot participate in the same instruction merging process, i.e., the two memory access instructions cannot be merged into a merged instruction.

[0108] Exemplarily, if the instruction types of two adjacent memory access instructions are the same, the comparator determines that the two adjacent memory access instructions meet the instruction merging condition; if the instruction types of the memory access instructions of two adjacent memory access instructions are different, the comparator determines that the two adjacent memory access instructions do not meet the instruction merging condition. The two adjacent memory access instructions refer to the memory access instructions stored in two adjacent entries in the first buffer. That is, when the first buffer stores the memory access instructions according to the execution sequence, the execution sequence of the memory access instructions stored in the two adjacent entries is adjacent. At this time, the comparator only needs to compare whether the types of the two adjacent memory access instructions are consistent.

[0109] Optionally, when the instruction merge condition also includes the presence of a stable offset between the memory access addresses of any two memory access instructions whose execution sequence is adjacent, the two inputs of the comparator include the instruction types of the two memory access instructions and the memory access addresses of the two memory access instructions, and the output of the comparator is whether the two memory access instructions meet the instruction merge condition. If the instruction types of the two adjacent memory access instructions are the same and the memory access address offset of the two memory access instructions is a stable offset, the comparator determines that the two adjacent memory access instructions meet the instruction merge condition.

[0110] In some embodiments, a plurality of comparators are provided in the instruction fusion unit, and the plurality of comparators form a hierarchical structure. For example, the plurality of comparators constitute at least one comparator hierarchy, and at least one comparator hierarchy is in a tree structure. Exemplarily, the comparators in the same comparator hierarchy execute comparison logic in parallel, that is, the comparators included in the same comparator hierarchy execute comparison logic in the same time period. Exemplarily, the inputs of the comparators in the same comparator hierarchy are different.

[0111] Optionally, the first-level comparator refers to a comparator at the lowest level in at least one comparator hierarchy, and for any comparator belonging to the first-level comparator, the comparator is used to determine whether the memory access instructions stored in two entries in the first buffer meet the instruction merging condition. In other words, any comparator in the first-level comparator corresponds to two entries in the first buffer, and the comparator is used to determine whether the memory access instructions stored in the corresponding two entries respectively meet the instruction merging condition.

[0112] Exemplarily, any comparator in the first-stage comparator corresponds to two adjacent entries in the first buffer. For example, the first buffer includes 16 entries, namely e1, e2, e3, e4, e5, e6, e7, e8, e9, e10, e11, e12, e13, e14, e15, and e16. The first-stage comparator includes 8 comparators, where comparator 1 corresponds to entries e1 and e2; comparator 2 corresponds to entries e3 and e4; comparator 3 corresponds to entries e5 and e6, and so on.

[0113] Exemplarily, the number of comparators included in the first-stage comparator is proportional to the total number s of entries in the first buffer. The greater the total number of entries in the first buffer, the greater the number of comparators included in the first-stage comparator; and the fewer the total number of entries in the first buffer, the fewer the number of comparators included in the first-stage comparator. For example, the number of comparators included in the first-stage comparator is equal to half the total number of entries in the first buffer.

[0114] Optionally, when the instruction fusion unit includes multiple levels of comparators, for two adjacent levels of comparators, the two inputs of the comparator in the higher level are the outputs of the comparator in the lower level. For example, for any comparator in the i-th level, the input of the comparator is the output of the two comparators in the i-1th level, where i is a positive integer greater than 1.

[0115] The output of the comparator is called a comparison result. Exemplarily, the comparison result includes result information. The comparison result also includes at least one of the following: instruction type information and an offset. The result information indicates whether the instruction merge condition is met or not met. If the result information indicates that the instruction merge condition is met, the comparison result also includes instruction type information, which is used to characterize the instruction type of the memory access instruction corresponding to the comparator. Exemplarily, the instruction type information includes a load instruction and a store instruction.

[0116] When the result information meets the instruction merging condition, the comparison result also includes an offset. The offset is used to represent the offset value between the memory access addresses of two memory access instructions with adjacent execution sequences among the multiple memory access instructions corresponding to the comparator.

[0117] In some embodiments, if the instruction fusion unit includes at least one level of comparators and there is an i-th level comparator, i is a positive integer greater than 1, then the j-th comparator in the i-th level comparator is used to determine the comparison result of the j-th comparator according to the comparison result of the 2j-1-th comparator in the i-1-th level comparator and the comparison result of the 2j-th comparator, and the comparison result of the j-th comparator is used to represent the (j-1)*2 i-1 +1 entry to j*2 i-1 Each entry stores the memory access instructions and the compliance with the instruction merging conditions.

[0118] Optionally, the jth comparator obtains two instruction type information from the comparison result of the 2j-1th comparator and the comparison result of the 2jth comparator. If the two instruction type information are of the same type, the result information included in the comparison result of the jth comparator is in compliance with the instruction merging condition. Optionally, when the result information in the comparison result of the jth comparator is in compliance with the instruction merging condition, the comparison result of the jth comparator also includes the (j-1)*2th instruction type information in the first buffer. i-1 +1 entry to j*2 i-1 Each entry stores the instruction type of the memory access instruction.

[0119] In some embodiments, the instruction generator is used to select n memory access instructions from at least one memory access instruction based on the comparison results of at least one level of comparator, and merge the n memory access instructions according to the instruction contents of the n memory access instructions to obtain a first merged instruction.

[0120] Optionally, the instruction generator selects a comparator with the highest level from at least one level of comparators, and determines that the result information included in the comparison result meets the instruction merging conditions as the target comparator. The instruction generator uses the memory access instructions stored in the multiple entries corresponding to the target comparator in the first buffer as n memory access instructions.

[0121] Optionally, fusing n memory access instructions into one instruction means compressing and packaging the instruction contents of the n memory access instructions to obtain a first merged instruction. For the instruction contents of the first merged instruction, please refer to the following embodiments.

[0122] By setting multiple comparators in the instruction fusion unit and arranging the multiple comparators in a hierarchical manner, it helps to improve the speed of determining n memory access instructions from at least one memory access instruction and shorten the time spent on the instruction fusion process.

[0123] The following describes the content of the first merge instruction through several embodiments.

[0124] In some embodiments, the instruction content of the first merged instruction includes: basic operation information, used to indicate the basic memory access instruction with the earliest execution sequence among n memory access instructions; register information, used to represent the physical registers of the n memory access instructions, and the physical registers of the memory access instructions are used to record the memory access data corresponding to the memory access instructions during the execution of the memory access instructions; memory access address information, used to represent the memory access addresses of the n memory access instructions, and the memory access addresses of the memory access instructions are used to store the memory access data indicated by the memory access instruction, or to provide the memory access data for the memory access instruction to read.

[0125] In some embodiments, the instruction content of the memory access instruction includes: an instruction type identifier, a register number, and an access address. The first merge instruction can represent the instruction content of n memory access instructions. The register number is used to uniquely identify a physical register. The numbers of different physical registers are not repeated.

[0126] In some embodiments, the time taken to execute the first merge instruction in the pipeline is less than or equal to one clock cycle. Executing the first merge instruction in the pipeline means determining the instruction contents of n memory access instructions based on the first merge instruction and executing the n memory access instructions in the pipeline.

[0127] Exemplarily, the pipeline determines the instruction content of n memory access instructions one by one. For example, after the i-th memory access instruction among the n memory access instructions is executed, the pipeline determines the instruction content of the i+1-th memory access instruction and executes the i+1-th memory access instruction, where i is a positive integer less than or equal to n.

[0128] Exemplarily, after receiving the first merge instruction, the pipeline first determines the instruction content of n memory access instructions and executes the n memory access instructions in parallel in the current clock cycle. Executing n memory access instructions in parallel can be understood as: the pipeline executes and completes n memory access instructions one by one according to the execution sequence within one clock cycle. For example, n is equal to 6, and according to the execution sequence from first to last, the n memory access instructions are: memory access instruction 1, memory access instruction 2, memory access instruction 3, memory access instruction 4, memory access instruction 5, and memory access instruction 6. Then, the pipeline determines the instruction content of the above 6 memory access instructions within one clock cycle, and executes the above 6 memory access instructions starting from memory access instruction 1.

[0129] In some embodiments, the first merged instruction includes register information and memory access address information. Optionally, the register information is used to directly indicate the physical registers of the n memory access instructions; and the memory access address information is used to directly indicate the memory access addresses of the n memory access instructions. For example, the register information includes the label of the physical register of each of the n memory access instructions, and the memory access address information includes the memory access instruction of each of the n memory access instructions.

[0130] In other embodiments, the first merged instruction includes basic operation information, register information, and memory access address information. Optionally, the basic operation information is used to represent a basic memory access instruction among n memory access instructions, where the basic memory access instruction refers to the memory access instruction with the earliest execution sequence among the n memory access instructions. The register information is used to represent the label of the physical register of the selected memory access instruction based on the physical register of the basic memory access instruction. The memory access address information is used to represent the memory access address of the selected memory access instruction based on the memory address of the basic memory access instruction. The selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions.

[0131] Optionally, the execution timing of the basic memory access instruction takes precedence over the execution timing of the selected memory access instruction.

[0132] Exemplarily, after receiving the first merge instruction, the pipeline determines the instruction content of the basic memory access instruction based on the basic operation information in the first merge instruction, and then executes the basic memory access instruction based on the instruction content of the basic memory access instruction. Furthermore, the pipeline can also determine the instruction content of the selected memory access instruction based on the instruction content of the basic memory access instruction, thereby executing n memory access instructions in parallel within a single clock cycle.

[0133] Of course, the basic memory access instruction refers to any one of the n memory access instructions. For example, the basic memory access instruction refers to the memory access instruction with the latest execution sequence among the n memory access instructions, which is not limited in this application.

[0134] Optionally, the register information is used to represent the physical register of the selected memory access instruction, and the offset label relative to the physical register of the basic memory access instruction; the access address information is used to represent the memory access address of the selected memory access instruction, and the address offset relative to the memory access address of the basic memory access instruction.

[0135] In this embodiment, register information and memory access address information are encoded based on basic memory access instructions, which can reduce the number of bits used for register information and memory access address information. This encoding method helps control the data volume of the instruction content of the merge instruction, allowing the merge instruction to be represented with fewer bits. This facilitates the use of the microinstruction bandwidth used to transmit memory access instructions in the instruction execution module to dispatch and issue the merge instruction, reducing the need to modify the instruction transmission port of the instruction execution module.

[0136] The following describes several encoding formats of register information supported by the instruction execution module in the embodiments of the present application.

[0137] Example 1: The occupancy status of one physical register is represented by one bit.

[0138] In some embodiments, register information is represented by m bits, different bits in the m bits correspond to different physical registers, one bit in the m bits corresponding to the physical register of the selected memory access instruction represents an occupied state, and the other bits in the m bits represent an idle state. The selected memory access instruction refers to a memory access instruction other than the basic memory access instruction among the n memory access instructions.

[0139] Register information may be represented as Phy_reg_off.

[0140] Optionally, m is a positive integer, and the value of m is positively correlated with the total number of entries included in the first buffer. The smaller the total number of entries included in the first buffer, the smaller m; the larger the total number of entries included in the first buffer, the larger m. Exemplarily, m is greater than or equal to the total number of entries included in the first buffer. For example, if the total number of entries included in the first buffer is 16, m=16.

[0141] Optionally, for any one of the m bits, the bit corresponds to the index of a physical register, that is, the bit corresponds to a physical register. Each of the m bits has a one-to-one correspondence with a physical register. For example, the physical register of the x-th bit among the m bits has the index X, where x is a positive integer less than or equal to m. A mapping relationship exists between X and x, such as X=x.

[0142] Exemplarily, the label of the physical register corresponding to the m bits is related to the label of the physical register of the basic memory access instruction.

[0143] For example, the index of the physical register corresponding to the first bit of the m bits (i.e., the least significant bit of the m bits) is offset by 1 from the index of the physical register corresponding to the basic memory access instruction. Assuming that the index of the physical register corresponding to the basic memory access instruction is P80, the index of the physical register corresponding to the first bit of the m bits is P81, the index of the physical register corresponding to the second bit of the m bits is P82, and so on.

[0144] For another example, the first bit among the m bits corresponds to the physical register number of the basic memory access instruction. That is, the first bit corresponds to the physical register of the basic memory access instruction. Assuming that the physical register number corresponding to the basic memory access instruction is P80, the first bit among the m bits corresponds to the physical register number P80, the second bit among the m bits corresponds to the physical register number P81, and so on.

[0145] Optionally, any one of the m bits may represent an idle state or an occupied state. The specific value of the bit may be "0" or "1," wherein one of "0" and "1" represents an idle state and the other represents an occupied state.

[0146] For example, "1" indicates an occupied state, and "0" indicates an idle state. For example, if the value of the first bit among m bits is 1, the physical register corresponding to that bit is occupied. In other words, among the n memory access instructions, there is one memory access instruction 1, and this memory access instruction 1 uses the physical register corresponding to that bit.

[0147] Optionally, n-1 bits among the m bits represent an occupied state, and m-n+1 bits among the m bits excluding n-1 bits represent an idle state. For example, in the above example where "the index of the physical register corresponding to the first bit among the m bits (i.e., the lowest bit among the m bits) is offset by 1 relative to the index of the physical register corresponding to the basic memory access instruction," n-1 bits among the m bits represent an occupied state, and m-n+1 bits among the m bits excluding n-1 bits represent an idle state.

[0148] Optionally, n bits among the m bits represent an occupied state, and mn bits represent an idle state. For example, in the example above where "the physical register number corresponding to the first bit among the m bits is the physical register number of the basic memory access instruction," n bits among the m bits represent an occupied state, and mn bits represent an idle state.

[0149] In some embodiments, m bits are represented using a bit field, where a bit field refers to a data structure composed of binary bits. That is, register information is represented using a bit field. Optionally, the bit field includes m binary bits, where m binary bits are m bits. Each binary bit in the bit field corresponds to the index of a physical register. That is, based on the specific value of the binary bit, the index of the physical register used by each of the n memory access instructions can be determined.

[0150] Table 1 Correspondence between binary bits in the bit field and the labels of physical registers

[0151] As shown in Table 1, taking m = 16 as an example, the bit field includes 16 binary bits, represented from the lowest order to the highest order as Bit 0, Bit 1, ..., Bit 15. Each binary bit corresponds to a physical register number. Optionally, the value of the physical register number corresponding to the lowest order binary bit is smaller than the value of the physical register number corresponding to the highest order binary bit.

[0152] In the example in Table 1, the physical register for the basic memory access instruction is labeled P80, the physical register corresponding to bit 0 is labeled P81, the physical register corresponding to bit 2 is labeled P82, and so on. This arrangement facilitates calculation of the correspondence between m bits and physical register labels, helping to simplify the pipeline's processing logic when determining the physical registers for n memory access instructions based on register information.

[0153] Optionally, the specific value of the binary bit indicates whether the corresponding physical register is in an occupied state. For example, if the specific value of the binary bit is "1", it indicates that the physical register corresponding to the binary bit is in an occupied state; if the specific value of the binary bit is "0", it indicates that the physical register corresponding to the binary bit is in an idle state, that is, the physical registers corresponding to the binary bit are not included in the physical registers of the n memory access instructions.

[0154] The register information generated using the above encoding format can represent both n memory access instructions with consecutive physical register numbers and n memory access instructions with discontinuous physical register numbers. The following describes two examples of register information encoding formats for n memory access instructions with consecutive physical register numbers and n memory access instructions with discontinuous physical register numbers.

[0155] Table 2 shows a bit field encoding format 1 of register information provided by an exemplary embodiment of the present application.

[0156] Table 2 Bit field encoding format of register information 1

[0157] It should be noted that the information within “()” in Table 2 and Tables 3 and 4 below is displayed for ease of understanding. In fact, the register information consists of only a few bits, that is, the register information does not actually include the content displayed within “()” in the second row of the table.

[0158] In this example, the instruction merging conditions include: the n memory access instructions have the same instruction type, and the physical registers for the n memory access instructions have consecutive numbers; m = 16, n = 6, the physical register for the base memory access instruction is numbered P80, and in the bit field corresponding to the m bits: the physical register corresponding to binary bit 0 is numbered P81, the physical register corresponding to binary bit 1 is numbered P82, the physical register corresponding to binary bit 2 is numbered P83, and so on. Assume that the physical registers for the selected memory access instructions among the n memory access instructions are numbered P81, P82, P83, P84, and P85, respectively. This means that the specific values ​​of binary bits Bit 0, Bit 1, Bit 2, Bit 3, and Bit 4 in the bit field are "1," and the specific values ​​of the remaining 11 high-order bits in the bit field are all 1. The m bits can be represented as h0003, where h represents 16 binary bits.

[0159] Table 3 shows a bit field encoding format 2 of a register provided by an exemplary embodiment of the present application.

[0160] Table 3 Register information bit field encoding format 2

[0161] In this example, the instruction merging conditions include: n memory access instructions have the same instruction type; m = 16, n = 6; the physical register for the base memory access instruction is labeled P80; and in the bit field corresponding to the m bits, the physical register corresponding to bit 0 is labeled P81, bit 1 is labeled P82, bit 2 is labeled P83, and so on. Assume that among the n memory access instructions, the physical registers for the remaining five selected memory access instructions, in descending order of execution, are labeled P83, P87, P88, P89, and P92. This means that bits 2, 6, 7, 8, and 11 in the bit field have values ​​of "1"; the remaining 11 high-order bits in the bit field have values ​​of 0. For example, in Table 3, Bit[1:0] represents bits 0 and 1, and the values ​​of bits 0 and 1 are 0.

[0162] This register information encoding method can represent register information with fewer bits, which helps prevent the merged instructions from exceeding the transmission bandwidth of a microinstruction.

[0163] Example 2: The occupancy status of a physical register is indicated by k bits.

[0164] In some embodiments, register information is represented by (n-1)*k bits, k=ceil(log2s), ceil() represents rounding up, s is the total number of entries in the first buffer included in the allocation queue for storing memory access instructions, and the k bits corresponding to the selected memory access instruction in the register information are used to represent the label of the physical register of the memory access instruction. The selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions.

[0165] ceil() means rounding up. For example, when s is 9, ceil(log2s)=4, and when s is 8, ceil(log2s)=3.

[0166] Optionally, for any selected memory access instruction, the corresponding k bits are used to represent the index of the physical register of the selected memory access instruction. For example, if the k bits corresponding to a selected memory access instruction are 0111, it means that the index of the physical register of the selected memory access instruction is 7.

[0167] Optionally, for any selected memory access instruction, the corresponding k bits are used to represent the offset number of the physical register of the selected memory access instruction relative to the physical register of the base memory access instruction. For example, if the physical register of the base memory access instruction is numbered P80, and the k bits corresponding to a selected memory access instruction are 0011, then the physical register of the selected memory access instruction is numbered P83, (83=80+3).

[0168] In some embodiments, the n-1 selected memory access instructions in the register information correspond to k bits respectively and are arranged according to the execution sequence of the selected memory access instructions. For example, the least significant k bits in the register information are used to represent the physical register of the selected memory access instruction with the earliest execution sequence.

[0169] In one example, n is equal to 6, and m is equal to 16. Then, in the encoding format of the register information provided in Example 2, the register information includes (6-1)*ceil(log216)=20 bits.

[0170] Table 4 shows a bit field encoding format 3 of register information provided by an exemplary embodiment of the present application.

[0171] Table 4 Register information bit field encoding format 3

[0172] In this example, the instruction merging conditions include: the n memory access instructions are of the same instruction type and are executed sequentially; m = 16, n = 6, the physical register of the base memory access instruction is labeled P80, and the five selected memory access instructions, in descending order of execution, are instruction 1, instruction 2, instruction 3, instruction 4, and instruction 5. As shown in Table 4, the k bits corresponding to instruction 1 are "0001," the k bits corresponding to instruction 2 are "0010," the k bits corresponding to instruction 3 are "0011," the k bits corresponding to instruction 4 are "0100," and the k bits corresponding to instruction 5 are "0101." The binary representation of the 20 bits of register information is "01010100001100100001."

[0173] This encoding format helps reduce the number of bits used to represent register information when the maximum value of n is small.

[0174] The identification format of the access address information is introduced and explained through several embodiments below.

[0175] In some embodiments, the memory access address information includes offset information, and the offset information is used to represent the relative offset between the memory access addresses of n memory access instructions.

[0176] A memory access address is a physical address for accessing memory. It corresponds to a memory location in memory. In other words, the memory access address allows data to be stored in or written to that memory location. Optionally, the memory access address indicates the offset of the selected memory access instruction's memory address relative to the base memory access address.

[0177] Optionally, the offset information includes: the offset direction of the memory address of any selected memory access instruction relative to the memory address of the basic memory access instruction, and the offset of the memory address of any selected memory access instruction relative to the memory address of the basic memory access instruction. In other words, the offset information includes n-1 offset sub-information, and for any one of the n-1 offset sub-information, the offset sub-information is used to represent the offset direction of the memory address of a selected memory access instruction relative to the memory address of the basic memory access instruction, and the offset of the memory address of the selected memory access instruction relative to the access address of the basic memory access instruction.

[0178] By encoding the offset information in this way, the representation of each selected memory access instruction is relatively flexible, which reduces the restrictions when selecting memory access instructions, helps to increase the number of memory access instructions selected from the dispatch queue, and makes the effect of executing the merge instruction equivalent to more memory access instructions, which helps to improve the execution efficiency of the pipeline on memory access instructions.

[0179] Exemplarily, the offset information includes: an offset direction of the memory address of the third memory access instruction relative to the memory address of the basic memory access instruction, and an offset of the memory address of the third memory access instruction relative to the memory address of the basic memory access instruction. The third memory access instruction refers to an access instruction that is adjacent to the execution sequence of the basic memory access instruction among the n-1 other selected memory access instructions. In this case, the offset directions of the memory addresses of the n-1 other selected memory access instructions relative to the access address of the basic memory access instruction are the same, and, among the n memory access instructions, the offsets between the access addresses of two memory access instructions that are adjacent in execution sequence are equal.

[0180] Representing the instruction addresses of n-1 selected memory access instructions by using offset information helps reduce the number of bits occupied by the memory access address information and also helps quickly obtain the instruction addresses of the n selected memory access instructions based on the offset information.

[0181] In some embodiments, the offset information includes direction information and an offset value; the direction information is used to represent the offset direction of the memory access address of the selected memory access instruction relative to the memory access address of the basic memory access instruction, and the offset direction includes at least one of the following: left offset, right offset; the offset value is used to represent the offset between the memory access addresses of any two memory access instructions with adjacent execution sequences among n memory access instructions.

[0182] The offset information can be expressed as Offset_bit.

[0183] Optionally, the offset direction is used to characterize the positional relationship between the memory address of the selected memory access instruction and the memory address of the base memory access instruction. A leftward offset means that the memory address of the selected memory access instruction is located to the left of the memory address of the base memory access instruction, i.e., the value of the memory address of the base memory access instruction is smaller than the value of the memory address of the base memory access instruction; and when arranged from earliest to latest according to execution time sequence, the memory addresses of the n memory access instructions decrease in sequence.

[0184] Among them, the right offset means that the memory access address of the selected memory access instruction is located to the right of the memory access address of the basic memory access instruction, that is, the value of the memory access address of the basic memory access instruction is greater than the value of the memory access address of the basic memory access instruction; when arranged from early to late according to the execution sequence, the memory access addresses of n memory access instructions increase sequentially.

[0185] Exemplarily, the offset direction is represented by one bit. For example, a specific value of "0" for this bit indicates an offset to the left, and a specific value of "1" for this bit indicates an offset to the right. For another example, a specific value of "1" for this bit indicates an offset to the left, and a specific value of "0" for this bit indicates an offset to the right.

[0186] Optionally, the offset value is a positive integer. Exemplarily, the offset value is equal to the absolute value of the difference between the memory access address of the third memory access instruction and the memory access address of the basic memory access instruction.

[0187] Exemplarily, the offset value is represented using ceil(log2log2(0+1)) bits. Where O represents the offset value. Exemplarily, if O is equal to 8, the offset value is represented using 2 bits. In this embodiment, the offset information is represented using at least ceil(log2log2(0+1))+1 bits.

[0188] Table 5 Coding format of offset information

[0189] As shown in Table 5, in one example, the basic memory access instruction is: ld p80,200(sp), where ld indicates that the instruction type is a load instruction, p80 indicates the label of the physical register, and 200(sp) indicates the memory access address of the basic memory access instruction.

[0190] The third memory access instruction is ld p80,192(sp), and among n memory access instructions, the offset values ​​of the memory access addresses of adjacent memory access instructions are equal. The offset value is 8, and this offset value is represented by 2 bits. The offset direction is left, and the offset direction is represented by 1 bit. Therefore, the offset information is represented by 3 bits, specifically: 0b111, where 0b represents binary encoding, bit 3 represents the offset direction, and bits [2:1] are 0b11, indicating an offset value of 8.

[0191] The following describes the generation process of the first merge instruction through an example.

[0192] FIG6 is a schematic diagram of an instruction merging process provided by an exemplary embodiment of the present application.

[0193] As shown in Figure 6, in the decode stage, six instructions are retrieved from the previous stage for decoding. In the rename stage, the registers of the six decoded instructions are renamed to obtain the memory access instructions to be executed. In the dispatch stage, the memory access instructions to be executed are written to the LSDQ. Subsequently, in the dispatch stage, the memory access instructions stored in the dispatch queue, in descending order of execution time, include: 0: Ld p80, 200 (sp); 1: Ld p81, 192 (sp); 2: Ld p82, 184 (sp); 3: Ld p83, 176 (sp); 4: Ld p84, 168 (sp); 5: Ld p85, 160 (sp); 6: Ld p20, 8 (p80); 7: Ld p3, 16 (p92); 8: st p12, 0 (p1).

[0194] Optionally, the instruction fusion unit is at the Dispatch2RS level. The instruction fusion unit selects the first six of the eight memory access instructions to be executed as n memory access instructions, and generates a first merge instruction MergeUop0 based on the six memory access instructions. The first merge instruction includes basic operation information: Ld p80,200 (sp), register information Phy_reg_off: 16'h0003, and memory access address information: Offset_addr: 3'b11.

[0195] As shown in Figure 6, assuming the port width between Dispatch2RS and Load RS is 4 uops, since the instruction fusion resulting from MergeUop0 corresponds to the first six Load instructions, this MergeUop occupies 1 uop in the port. The remaining 3 uops of the 4 uops are free and can be used to transmit the load instruction Ld p20, (p80) and the load instruction Ld p3, 16 (p92). Therefore, 6+2 instructions can be dispatched to Load RS in one clock cycle. Since instruction 8 in LSDQ is a store instruction, instruction 8 cannot be dispatched to Load RS in the same clock cycle and must be dispatched to Store RS in the next clock cycle.

[0196] The following describes the contents related to the reservation station through several embodiments.

[0197] In some embodiments, the instruction execution module also includes: a reservation station; the reservation station is used to use a microinstruction bandwidth to transmit a first merge instruction to the pipeline; the reservation station is also used to determine an abnormal memory access instruction based on an abnormal execution signal sent from the reorder cache, the abnormal execution signal is used to indicate an abnormal memory access instruction in the first merge instruction, and the abnormal memory access instruction is not successfully executed in the pipeline; re-transmitting the memory access instructions that have not started to be executed among the n memory access instructions to the pipeline, and the execution timing of the memory access instructions that have not started to be executed lags behind the execution timing of the abnormal memory access instruction.

[0198] In some embodiments, the instruction execution module further includes a reorder cache, which is used to determine the actual execution sequence of memory access instructions in the pipeline. If an exception occurs in a memory access instruction executed in the pipeline, the reorder cache is further used to broadcast an exception execution signal and re-determine the actual execution sequence for the unexecuted memory access instructions in the pipeline.

[0199] In the instruction execution module provided in the embodiment of the present application, new logic is added to the reservation station to deal with exceptions that occur during the execution of the merge instruction. From the above content, it can be seen that after the pipeline receives the merge instruction, it can determine the instruction content of the multiple selected memory access instructions based on the merge instruction and execute n memory access instructions in one clock cycle. When a memory access instruction among the n memory access instructions fails to execute in the pipeline, the reservation station needs to re-transmit the memory access instruction that has not been executed in the n memory access instructions.

[0200] Optionally, the abnormal execution signal refers to a flash signal emitted by the reorder buffer. When an exception occurs during pipeline execution of the first merge instruction, the reorder buffer broadcasts the abnormal execution signal. The abnormal execution signal is used to indicate the execution of the memory access instruction that encountered the exception. For example, the abnormal execution information is used to indicate the sequence number of the abnormal memory access instruction among n memory access instructions.

[0201] For example, a first merge instruction is generated by fusing n memory access instructions. The pipeline's first merge instruction actually executes n memory access instructions. If the tth memory access instruction among the n memory access instructions generates an exception during pipeline execution, the pipeline stops executing the remaining unexecuted memory access instructions among the n memory access instructions, and the reorder buffer broadcasts an exception execution signal related to the tth memory access instruction. In other words, the tth memory access instruction among the n instructions is an abnormal memory access instruction. Exemplarily, an abnormal memory access instruction is one with an abnormal memory access address.

[0202] For example, if the instruction merge conditions do not include a stable offset between the memory access addresses of sequentially adjacent memory access instructions, the n memory access instructions may include one with an abnormal memory address, preventing the pipeline from successfully executing the memory access instruction. In this case, the newly added execution logic in the reservation station needs to take effect.

[0203] Optionally, the reservation station records the abnormal execution signal sent by the reorder cache, and determines the abnormal memory access instruction from the n memory access instructions based on the abnormal execution information; then the reservation station re-transmits the memory access instructions that have not started to be executed among the n memory access instructions to the pipeline, and the execution timing of the memory access instructions that have not started to be executed lags behind the execution timing of the abnormal memory access instruction.

[0204] Exemplarily, a memory access instruction that has not yet begun execution refers to a memory access instruction whose execution sequence lags behind the execution sequence of the abnormal memory access instruction among n memory access instructions. For example, if the abnormal memory access instruction is the tth memory access instruction among n memory access instructions, then the memory access instructions that have not yet begun execution refer to the memory access instructions from the t+1th to the nth memory access instructions among the n memory access instructions. It should be noted that abnormal memory access instructions cannot be considered as memory access instructions that have not yet begun execution.

[0205] Optionally, the reservation station sends the unexecuted memory access instructions to the pipeline one by one. For example, the reservation station uses 1 micro-instruction bandwidth to send 1 unexecuted memory access instruction.

[0206] Optionally, the reservation station merges multiple memory access instructions that have not yet started to be executed into one instruction to obtain a second merged instruction, and the reservation station sends the second merged instruction to the pipeline.

[0207] By adding new logic to the reservation station, it is possible to handle situations where abnormal memory access instructions fail to execute in the pipeline, which helps maintain the execution order of memory access instructions in the pipeline.

[0208] In some embodiments, the reservation station includes: an exception search unit and an instruction re-fusion unit; the exception search unit is used to determine an abnormal memory access instruction from n memory access instructions based on an abnormal execution signal; the instruction re-fusion unit is used to fuse at least two memory access instructions that have not started execution among the n memory access instructions into one instruction to obtain a second merged instruction, the effect of the second merged instruction is equivalent to the effect of the at least two memory access instructions that have not started execution, and transmitting the second merged instruction to the pipeline occupies one microinstruction bandwidth.

[0209] The reservation station then issues a second merge instruction to the pipeline.

[0210] Optionally, the exception search unit is used to determine the abnormal memory access instruction from the n memory access instructions according to the instruction sequence number included in the abnormal execution signal, and determine the instruction content of at least one memory access instruction that has not started to be executed according to the first merge instruction.

[0211] Exemplarily, the exception search unit determines the sequence number of the fourth memory access instruction based on the instruction sequence number included in the abnormal execution signal, and determines the instruction content of the fourth memory access instruction based on the first merge instruction. The fourth memory access instruction is the earliest one among the multiple memory access instructions that have not yet started execution.

[0212] Optionally, the structure of the instruction re-fusion unit is similar to the above instruction fusion unit, and for details, please refer to the above embodiment, which will not be described in detail here. Exemplarily, the execution timing of at least two memory access instructions that have not yet started execution is continuous.

[0213] As can be seen from the above embodiments, the instruction content of the first merge instruction includes basic operation information, register information and memory access address. The instruction content of the fourth memory access instruction can be calculated based on the above three information and the sequence number of the fourth memory access instruction in n memory access instructions.

[0214] Subsequently, the instruction re-fusion unit uses the fourth memory access instruction as the basic memory access instruction among at least two memory access instructions that have not yet started execution, and determines the register information and memory access address information of the second merged instruction based on the register information and memory access address information of the first memory access instruction to obtain the instruction content of the second merged instruction.

[0215] FIG7 is a schematic diagram of an instruction re-issuing process provided by an exemplary embodiment of the present application.

[0216] In this embodiment, n is 6, and the n memory access instructions, in descending order of execution time, include: memory access instruction 1, memory access instruction 2, memory access instruction 3, memory access instruction 4, memory access instruction 5, and memory access instruction 6. The dispatch queue merges these six memory access instructions into one instruction, generating a first merged instruction. This first merged instruction is then transmitted to the load pipeline via the load reservation station. After receiving the first merged instruction, the load pipeline determines the instruction content of the n memory access instructions based on the first merged instruction and begins executing the base memory access instruction (memory access instruction 1) within a certain clock cycle. If memory access instruction 4 executes abnormally in the load pipeline, the load pipeline notifies the reorder cache that memory access instruction 4 has executed abnormally, and the reorder cache broadcasts an exception execution signal. Based on the exception execution signal, the load reservation station determines that memory access instruction 4 has executed abnormally in the pipeline and merges memory access instructions 5 and 6 into a second merged instruction. The load reservation station then transmits the second merged instruction to the load pipeline.

[0217] By performing secondary packaging on multiple memory access instructions that have not yet been executed, it helps to reduce the clock cycles for re-issuing the multiple memory access instructions that have not yet been executed to the pipeline, so that the secondary packaging of the multiple memory access instructions that have not yet been executed can be completed within one clock cycle of the pipeline. In the event of an abnormality in the execution of the memory access instruction value, it helps to ensure the execution efficiency of the pipeline for the memory access instruction.

[0218] FIG8 is a schematic diagram of an instruction re-issuing process provided by an exemplary embodiment of the present application.

[0219] In this embodiment, n is 6, and the n memory access instructions, in descending order of execution time, include: memory access instruction 1, memory access instruction 2, memory access instruction 3, memory access instruction 4, memory access instruction 5, and memory access instruction 6. The dispatch queue merges these six memory access instructions into a single instruction, generating a first merged instruction. This merged instruction is then transmitted to the load pipeline via the load reservation station. After receiving the first merged instruction, the load pipeline determines the instruction contents of the n memory access instructions based on the first merged instruction and begins executing the base memory access instruction (memory access instruction 1) within a certain clock cycle. If memory access instruction 4 executes abnormally in the load pipeline, the load pipeline notifies the reorder cache that memory access instruction 4 has executed abnormally, and the reorder cache broadcasts an abnormal execution signal. Based on the abnormal execution signal, the load reservation station determines that memory access instruction 4 has executed abnormally in the pipeline and determines that the memory access instructions that have not yet begun execution among the n memory access instructions include memory access instruction 5 and memory access instruction 6. The load reservation station then transmits memory access instruction 5 and memory access instruction 6, respectively, to the load pipeline.

[0220] Although memory access instruction 5 and memory access instruction 6 are not merged again in this embodiment, since the memory access instruction 1, memory access instruction 2 and memory access instruction 3 of the first merged instruction have been executed in the same clock cycle, compared with not compressing and packaging the memory access instructions and issuing n memory access instructions separately at one time, the clock cycles consumed by the pipeline to execute memory access instruction 1, memory access instruction 2 and memory access instruction 3 are shortened.

[0221] In some embodiments, the reservation station also includes a second buffer, which is used to temporarily store the first merge instruction and other memory access instructions in at least one candidate memory access instruction except n memory access instructions; each entry in the second buffer is used to store the instruction content of a merge instruction, or the instruction content of another memory access instruction; the instruction content includes: instruction identification information, register label information, immediate value information and synchronous execution information, the instruction identification is used to represent the instruction type of the instruction stored in the entry, the register information is used to represent the physical register of the instruction stored in the entry, the memory access address information is used to represent the memory access address of the instruction stored in the entry, and the synchronous execution information is used to represent the number of memory access instructions executed synchronously in the pipeline.

[0222] Because the reservation station needs to store the first merge instruction before issuing the first merge instruction, the entry in the reservation station for storing the instruction also needs to be modified accordingly.

[0223] In some embodiments, the entries in the second buffer are used to store instruction types including: single memory access instruction, first merge instruction, and second merge instruction. Table 6 is a schematic table illustrating the format of entries in a reservation station provided by an exemplary embodiment of the present application.

[0224] Table 6 Format of entries in reservation station

[0225] Since multiple memory access instructions are merged into one merge instruction (including the first merge instruction and the second merge instruction), the merge instruction only occupies one entry in the reservation station. Therefore, the entry in the reservation station needs to adapt to the instruction content of the merge instruction.

[0226] Each entry in the second buffer includes instruction identification information (Opcode), register label information, immediate value information, and synchronous execution information. Optionally, the instruction identification information is also used to represent the instruction identifier of the instruction stored in the entry, and the instruction identifier can be the sequence number of the memory access instruction in the instruction stream.

[0227] Optionally, the register label information of the merged instruction includes the label of the physical register of the basic memory access instruction corresponding to the merged instruction, and the register information of the merged instruction. Optionally, the immediate value information of the merged instruction includes the memory access address of the basic memory access instruction corresponding to the merged instruction, and the memory access address information of the merged instruction.

[0228] Among them, register label information needs to be expanded. For example, multiple bits can be added to the register information bit field (p80) to encode the physical register labels of n memory access instructions. For example, 16 bits can be added to the register information bit field. Immediate value information needs to be expanded, such as adding 3 bits to the immediate value information bit field to encode the difference between n memory access addresses. Synchronous execution information also needs to be expanded, such as adding 15 bits to the synchronous execution information bit field to encode up to 16 instructions to be executed simultaneously.

[0229] Illustratively, for an entry storing a memory access instruction, a zero-fill operation is performed on the high bits of the bit fields of each information in the entry to ensure that the memory access instruction and the merge instruction have consistent formats in the entry.

[0230] By adjusting the format of the entries in the second buffer, the adjusted entries in the second buffer can correctly record the merge instruction.

[0231] In some embodiments, a merge instruction can retrieve a cache line (64B) of data at once. This means that the bandwidth of the data processing stage in the instruction execution module accessing the data cache is increased to 64B. This allows the pipeline to concurrently execute multiple memory access instructions within the merge instruction within a single clock cycle to meet data access requirements.

[0232] In some embodiments, the instruction execution module also includes: a write-back cache, a write-back buffer, which is used to write the execution result of the instruction to be written back to the general register through a first write-back path, and the first write-back path is adapted to the pipeline to execute a single memory access instruction in one clock cycle; the write-back buffer is also used to write the execution results of different instructions to be written back to the general register through the first write-back path and the second write-back path when there is a second write-back path between the write-back buffer and the general register, and the second write-back path is adapted to the execution speed change generated by the pipeline based on the first merge instruction; wherein, the instruction to be written back refers to the memory access instruction that has been executed in the pipeline.

[0233] Optionally, after the memory access instruction is executed in the pipeline, the pipeline transfers the execution result of the instruction to be written back to the write-back buffer, and the write-back buffer writes the execution result of the instruction to be written back back to the general register. In the related art, the pipeline only executes one memory access instruction in one clock cycle, so the write-back pressure on the write-back buffer is relatively small. Optionally, there is a first write-back path between the write-back buffer and the general register. The first write-back path is adapted to the speed at which the pipeline generates an execution result of the instruction to be written back in one clock cycle. In other words, the first write-back path refers to the write-back path set between the write-back buffer and the general register before the instruction execution device proposed in this application appears.

[0234] The instruction to be written back refers to a memory access instruction that has been executed successfully in the pipeline. Optionally, the instruction to be written back includes any one of n memory access instructions that has been executed successfully by the pipeline. Because the pipeline can execute multiple memory access instructions within a single clock cycle based on the first merge instruction, the pipeline generates n execution results of instructions to be written back within a single clock cycle, which may result in a blockage when writing these execution results to the general register.

[0235] In one example, the instruction execution module writes the instructions to be written back into the general-purpose registers one by one according to the first write-back path established between the write-back buffer and the general-purpose registers. That is, in the instruction execution module provided by this application, there is no need to modify the data path between the write-back buffer and the general-purpose registers. The write-back registers write the execution results of each instruction to be written back back into the general-purpose registers in the order in which they are executed in the pipeline.

[0236] The second write-back path is a newly added write-back path between the write-back buffer and the general register to improve the efficiency of pipeline execution of memory access instructions. Optionally, the write-back processes of the first write-back path and the second write-back path do not interfere with each other.

[0237] In another example, when a second write-back path is added between the write-back buffer and the general registers, the instruction execution module can use both the first and second write-back paths to concurrently write the execution results of multiple instructions to be written back to the general registers, improving write-back efficiency. This configuration helps reduce the occurrence of blockages.

[0238] In some embodiments, the instruction execution module also includes a write-back buffer and a buffer bypass network, the write-back buffer is coupled to the buffer bypass network, and the buffer bypass network is coupled to the read port of the pipeline; the buffer bypass network is used to transmit the execution result of the instruction to be written back to the read port of the pipeline before the execution result of the instruction to be written back is written back to the general register, and the instruction to be written back refers to the memory access instruction executed in the pipeline.

[0239] For example, if an execution unit in the pipeline needs to use the execution result of an instruction to be written back, but the execution result is still queued for writing back to a general register, the execution unit obtains the execution result from the write-back buffer via the register bypass network. This configuration helps reduce the number of clock cycles the execution unit must wait to obtain data, thereby improving instruction execution efficiency.

[0240] The following is an example to illustrate the instruction execution module.

[0241] The instruction execution module includes: dispatch queue, reservation station, pipeline, and write-back register.

[0242] A dispatch queue is used to store at least one memory access instruction to be executed; n memory access instructions are selected from the at least one memory access instruction to be executed; the n memory access instructions are merged into one instruction to obtain a first merged instruction, and the effect of the first merged instruction is equivalent to the effect of the n memory access instructions.

[0243] Exemplarily, the maximum value of n is equal to 6, and the minimum value is equal to 2.

[0244] Optionally, the dispatch queue includes: a first buffer and an instruction fusion unit; the first buffer is used to store at least one memory access instruction to be executed; the instruction fusion unit is used to select n memory access instructions with continuous execution sequence from at least one memory access instruction to be executed based on an instruction fusion condition, as n memory access instructions, and the instruction type includes at least one of the following: load instruction, store instruction; the instruction fusion unit is also used to fuse the n memory access instructions into one instruction according to the instruction content of the n memory access instructions to obtain a first merged instruction.

[0245] Exemplarily, among the n memory access instructions, there is a stable offset between the memory access addresses of any two memory access instructions that are executed adjacently in timing.

[0246] In some embodiments, the instruction content of the first merge instruction includes: basic operation information, register information, and memory access address information.

[0247] The basic operation information includes the instruction content of the basic memory access instruction among the n memory access instructions. The register information is represented by (n-1)*k bits, where k=ceil(log2s), ceil() represents rounding up, and s is the total number of entries in the first buffer included in the dispatch queue for storing memory access instructions. The k bits corresponding to the selected memory access instruction in the register information are used to represent the label of the physical register of the memory access instruction. The selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions. The memory access address information includes: direction information and an offset value; the direction information is used to represent the memory access address of the selected memory access instruction, the offset direction relative to the memory access address of the basic memory access instruction, the offset direction including at least one of the following: left offset and right offset; the selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions; the offset value is used to represent the offset between the memory access addresses of any two memory access instructions with adjacent execution timings among the n memory access instructions.

[0248] a reservation station for issuing a first merge instruction to the pipeline using a microinstruction bandwidth;

[0249] The pipeline is configured to execute the first merge instruction.

[0250] The reservation station is further used to determine an abnormal memory access instruction based on an abnormal execution signal sent from the reorder cache, where the abnormal execution signal is used to indicate an abnormal memory access instruction in the first merge instruction, and the abnormal memory access instruction is not successfully executed in the pipeline; and re-transmit the memory access instruction that has not started to be executed among the n memory access instructions to the pipeline, where the execution timing of the memory access instruction that has not started to be executed lags behind the execution timing of the abnormal memory access instruction.

[0251] A write-back buffer is configured to write the execution result of the instruction to be written back to the general register via a first write-back path, the first write-back path being adapted to the pipeline executing a single memory access instruction in one clock cycle; the write-back buffer is further configured to write the execution results of different instructions to be written back to the general register via the first write-back path and the second write-back path, respectively, when a second write-back path exists between the write-back buffer and the general register, the second write-back path being adapted to the execution speed change of the pipeline caused by the first merge instruction;

[0252] In some embodiments, the write-back buffer is coupled to a buffer bypass network, which is coupled to a pipeline read port. The buffer bypass network is configured to transmit the execution result of the instruction to be written back to the pipeline read port before the execution result of the instruction to be written back is written back to a general register. The instruction to be written back refers to a memory access instruction that has been executed in the pipeline. The instruction execution module provided by this application helps improve the execution efficiency of memory access instructions, thereby improving processor performance.

[0253] It should be noted that in this embodiment, "ld" and "Ld" have the same meaning, both indicating that the instruction type of the memory access instruction is a load instruction. In the above embodiment, the load instruction is used as an example to illustrate the solution. In fact, the memory access instruction can also be a store instruction, which is not repeated here.

[0254] The following is an embodiment of the method of the present application. For details not disclosed in the embodiment of the method of the present application, please refer to the embodiment of the instruction execution module of the present application. Figure 9 is a flow chart of an instruction processing method applied to the instruction execution module provided by an exemplary embodiment of the present application. The instruction execution module includes: an allocation queue and a pipeline; the instruction processing method applied to the instruction execution module includes the following steps:

[0255] Step 910: The dispatch queue stores at least one memory access instruction to be executed.

[0256] Step 920 : The dispatch queue selects n memory access instructions from at least one memory access instruction to be executed, where n is an integer greater than 1.

[0257] In step 930 , the dispatch queue merges the n memory access instructions into one instruction to obtain a first merged instruction. The effect of the first merged instruction is equivalent to that of the n memory access instructions.

[0258] In step 940 , the pipeline executes the first merge instruction.

[0259] In some embodiments, an allocation queue includes: a first buffer and an instruction fusion unit; the first buffer stores at least one memory access instruction to be executed; step 920 includes: the instruction fusion unit selects n memory access instructions with continuous execution sequence from at least one memory access instruction to be executed based on an instruction merging condition, as n memory access instructions, wherein the instruction merging condition includes that the n memory access instructions have the same instruction type, and the instruction type includes at least one of the following: a load instruction, a store instruction; step 930 includes: the instruction fusion unit fuses the n memory access instructions according to the instruction content of the n memory access instructions to obtain a first merged instruction.

[0260] In some embodiments, the instruction merging condition further includes at least one of the following: the labels of the physical registers of the n memory access instructions are continuous; among the n memory access instructions, there is a stable offset between the memory access addresses of any two memory access instructions with adjacent execution sequences.

[0261] In some embodiments, the instruction fusion unit includes: at least one level of comparator and an instruction generator; step 910 includes, sub-step 912, for the first level comparator in the at least one level of comparator, the kth comparator in the first level comparator, according to the memory access instructions stored in the 2kth entry and the 2k+1th entry in the first buffer, respectively, determining the comparison result of the kth comparator, wherein the entry of the first buffer is used to store the memory access instruction, k is a positive integer less than or equal to 0.5*s, and s is the total number of entries included in the first buffer , the comparison result of the kth comparator is used to represent the memory access instructions stored in the 2kth entry and the 2k+1th entry respectively, and the compliance with the instruction merging condition; if there is an i-th level comparator in at least one level of comparators, i is a positive integer greater than 1, then the j-th comparator in the i-th level comparator determines the comparison result of the j-th comparator according to the comparison result of the 2j-1th comparator in the i-1th level comparator and the comparison result of the 2jth comparator, and the comparison result of the j-th comparator is used to represent the (j-1)*2th comparator in the first buffer. i-1 +1 entry to j*2 i-1 The memory access instructions stored in the entries are respectively compared to see whether the instruction merging conditions are met; in sub-step 914, the instruction generator selects n memory access instructions from at least one memory access instruction according to the comparison results of at least one level of comparator, and merges the n memory access instructions according to the instruction contents of the n memory access instructions to obtain a first merged instruction.

[0262] In some embodiments, the instruction content of the first merged instruction includes: basic operation information, used to indicate the basic memory access instruction with the earliest execution sequence among n memory access instructions; register information, used to represent the physical registers of the n memory access instructions, and the physical registers of the memory access instructions are used to record the memory access data corresponding to the memory access instructions during the execution of the memory access instructions; memory access address information, used to represent the memory access addresses of the n memory access instructions, and the memory access addresses of the memory access instructions are used to store the memory access data indicated by the memory access instruction, or to provide the memory access data for the memory access instruction to read.

[0263] In some embodiments, register information is represented by m bits, different bits in the m bits correspond to different physical registers, one bit in the m bits corresponding to the physical register of the selected memory access instruction represents an occupied state, and the other bits in the m bits represent an idle state. The selected memory access instruction refers to a memory access instruction other than the basic memory access instruction among the n memory access instructions.

[0264] In some embodiments, register information is represented by (n-1)*k bits, k=ceil(log2s), ceil() represents rounding up, s is the total number of entries in the first buffer included in the allocation queue for storing memory access instructions, and the k bits corresponding to the selected memory access instruction in the register information are used to represent the label of the physical register of the memory access instruction. The selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions.

[0265] In some embodiments, the memory access address information includes offset information, and the offset information is used to represent the relative offset between the memory access addresses of n memory access instructions.

[0266] In some embodiments, the offset information includes direction information and an offset value; the direction information is used to represent the memory access address of the selected memory access instruction, relative to the offset direction of the memory access address of the basic memory access instruction, and the offset direction includes at least one of the following: left offset, right offset, and the selected memory access instruction refers to the memory access instruction other than the basic memory access instruction among the n memory access instructions; the offset value is used to represent the offset between the memory access addresses of any two memory access instructions with adjacent execution timing among the n memory access instructions.

[0267] FIG10 is a schematic diagram of an instruction execution method applied to an instruction execution module provided by another exemplary embodiment of the present application.

[0268] In some embodiments, the instruction execution module further includes: a reservation station; the reservation station uses a microinstruction bandwidth to transmit a first merge instruction to the pipeline; the method shown in Figure 9 also includes: step 950, the reservation station determines the abnormal memory access instruction based on the abnormal execution signal sent from the reorder cache, the abnormal execution signal is used to indicate the abnormal memory access instruction in the first merge instruction, and the abnormal memory access instruction is not successfully executed in the pipeline; step 960, the reservation station re-transmits the memory access instruction that has not started execution among the n memory access instructions to the pipeline, and the execution timing of the memory access instruction that has not started execution lags behind the execution timing of the abnormal memory access instruction.

[0269] In some embodiments, the reservation station includes: an exception search unit and an instruction re-fusion unit; step 950 includes sub-step 952 (not shown in Figure 10), the exception search unit determines the abnormal memory access instruction from n memory access instructions based on the abnormal execution signal; sub-step 956 (not shown in Figure 10), the instruction re-fusion unit fuses at least two memory access instructions that have not started execution among the n memory access instructions to obtain a second merged instruction, the effect of the second merged instruction is equivalent to the effect of the at least two memory access instructions that have not started execution, and transmitting the second merged instruction to the pipeline occupies one microinstruction bandwidth.

[0270] In some embodiments, the reservation station also includes a second buffer, which temporarily stores the first merge instruction and other memory access instructions in at least one candidate memory access instruction except n memory access instructions; each entry in the second buffer is used to store the instruction content of a merge instruction, or the instruction content of another memory access instruction; the instruction content includes: instruction identification information, register label information, immediate value information and synchronous execution information, the instruction identification is used to represent the instruction type of the instruction stored in the entry, the register label information is used to represent the physical register of the instruction stored in the entry, the immediate value information is used to represent the memory access address of the instruction stored in the entry, and the synchronous execution information is used to represent the number of memory access instructions executed synchronously in the pipeline.

[0271] In some embodiments, the instruction execution module also includes: a write-back buffer; the method shown in Figure 9 also includes step 970, the write-back buffer writes the execution result of the instruction to be written back to the general register through the first write-back path, and the first write-back path is adapted to the pipeline to execute a single memory access instruction in one clock cycle; in the case where there is a second write-back path between the write-back buffer and the general register, the write-back buffer writes the execution results of different instructions to be written back to the general register through the first write-back path and the second write-back path, and the second write-back path is adapted to the execution speed change generated by the pipeline based on the first merge instruction; wherein, the instruction to be written back refers to the memory access instruction that has been executed in the pipeline.

[0272] In some embodiments, the instruction execution module also includes a write-back buffer and a buffer bypass network, the write-back buffer is coupled to the buffer bypass network, and the buffer bypass network is coupled to the read port of the pipeline; the method shown in Figure 9 also includes the following steps, the buffer bypass network transmits the execution result of the instruction to be written back to the read port of the pipeline before the execution result of the instruction to be written back is written back to the general register, and the instruction to be written back refers to the memory access instruction executed in the pipeline.

[0273] For the specific content of this embodiment, please refer to the above embodiment, which will not be described in detail here.

[0274] The present application also provides a chip including the instruction execution module described above. Optionally, the chip includes the instruction execution module and a memory. The chip is a processor chip and can be integrated into the processor to perform the instruction execution method.

[0275] The present application also provides an electronic device including the instruction execution module described above. Optionally, the electronic device includes a chip including a memory and an instruction execution module. The instruction execution module is configured to execute instructions using the solution provided in the above embodiment, completing the process of instruction fusion and executing the merged instructions.

[0276] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.

[0277] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent switches, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. An instruction execution module applied to a processor, the instruction execution module comprising: Dispatch queue and pipeline; The dispatch queue is used to store at least one memory access instruction to be executed; Select n memory access instructions from the at least one memory access instruction to be executed, where n is an integer greater than 1; fuse the n memory access instructions into one instruction to obtain a first merged instruction, and the function of the first merged instruction is equivalent to the function of the n memory access instructions; The pipeline is used to execute the first merged instruction.

2. The instruction execution module according to claim 1, wherein, The dispatch queue includes: a first buffer and an instruction fusion unit; The first buffer is used to store the at least one memory access instruction to be executed; The instruction fusion unit is used to select n memory access instructions with consecutive execution timings from the at least one memory access instruction to be executed as the n memory access instructions based on an instruction merging condition, where the instruction merging condition includes that the instruction types of the n memory access instructions are the same; The instruction fusion unit is further used to fuse the n memory access instructions into one instruction according to the instruction contents of the n memory access instructions to obtain the first merged instruction.

3. The instruction execution module according to claim 2, wherein, The instruction merging condition further includes at least one of the following: The labels of the physical registers of the n memory access instructions are consecutive; Among the n memory access instructions, there is a stable offset between the memory access addresses of any two memory access instructions with adjacent execution timings.

4. The instruction execution module according to claim 2 or 3, wherein The instruction fusion unit includes: at least one stage of comparators and an instruction generator; For the first-stage comparator in the at least one stage of comparators, the k-th comparator in the first-stage comparator is used to determine the comparison result of the k-th comparator according to the memory access instructions stored in the 2k-th entry and the 2k + 1-th entry in the first buffer, where the entries of the first buffer are used to store the memory access instructions, k is a positive integer less than or equal to 0.5*s, s is the total number of entries included in the first buffer, and the comparison result of the k-th comparator is used to represent the compliance of the memory access instructions stored in the 2k-th entry and the 2k + 1-th entry with the instruction merging condition; If there is an i-th level comparator in the at least one level of comparators, where i is a positive integer greater than 1, then the j-th comparator in the i-th level comparator is used to determine the comparison result of the j-th comparator according to the comparison results of the (2j - 1)-th comparator and the 2j-th comparator in the (i - 1)-th level comparator, and the comparison result of the j-th comparator is used to represent the compliance of the memory access instructions stored in the ((j - 1)*2 i-1 + 1)th entry to the (j*2 i-1 th entries with respect to the instruction merging condition; The instruction generator is used to select the n memory access instructions from the at least one memory access instruction according to the comparison results of the at least one stage of comparators respectively, and fuse the n memory access instructions according to the instruction contents of the n memory access instructions to obtain the first merged instruction.

5. The instruction execution module according to any one of claims 1 to 4, wherein, The instruction content of the first merged instruction includes: Basic operation information, which is used to indicate the basic memory access instruction among the n memory access instructions, and the basic memory access instruction is the memory access instruction with the earliest execution timing among the n memory access instructions; Register information, which is used to represent the physical registers of the n memory access instructions, and the physical registers of the memory access instructions are used to record the memory access data corresponding to the memory access instructions during the execution of the memory access instructions; Memory access address information, which is used to represent the memory access addresses of the n memory access instructions, and the memory access addresses of the memory access instructions are used to store the memory access data indicated by the memory access instructions or provide the memory access data for the memory access instructions to read.

6. The instruction execution module according to claim 5, wherein The register information is represented by m bits. Different bits among the m bits respectively correspond to different physical registers. One bit corresponding to the physical register of the selected memory access instruction in the m bits represents the occupied state, and the other bits in the m bits represent the idle state. The selected memory access instruction refers to the memory access instructions among the n memory access instructions other than the basic memory access instruction.

7. The instruction execution module according to claim 5, wherein, The register information is represented by (n - 1)*k bits, where k = ceil(log2s), ceil() represents rounding up, and s is the total number of entries in the first buffer included in the dispatch queue for storing the memory access instructions. The k bits corresponding to the selected memory access instruction in the register information are used to characterize the label of the physical register of the memory access instruction. The selected memory access instruction refers to the memory access instructions among the n memory access instructions other than the basic memory access instruction.

8. The instruction execution module according to any one of claims 5 to 7, wherein, The memory access address information includes offset information, and the offset information is used to characterize the relative offset situation between the memory access addresses of the n memory access instructions.

9. The instruction execution module according to claim 8, wherein, The offset information includes direction information and an offset value; The direction information is used to characterize the offset direction of the memory access address of the selected memory access instruction relative to the memory access address of the basic memory access instruction. The offset direction includes at least one of the following: offset to the left, offset to the right. The selected memory access instruction refers to the memory access instructions among the n memory access instructions other than the basic memory access instruction; The offset value is used to characterize the offset amount between the memory access addresses of any two memory access instructions with adjacent execution timings among the n memory access instructions.

10. The instruction execution module according to any one of claims 1 to 9, wherein, The instruction execution module further includes: a reservation station; The reservation station is used to issue the first merged instruction to the pipeline using one micro-instruction bandwidth; The reservation station is further used to determine an exception memory access instruction based on an exception execution signal sent from the reorder buffer. The exception execution signal is used to indicate the exception memory access instruction in the first merged instruction, and the exception memory access instruction fails to be successfully executed in the pipeline. The reservation station reissues to the pipeline the memory access instructions among the n memory access instructions that have not started execution, and the execution timings of the memory access instructions that have not started execution are behind the execution timing of the exception memory access instruction.

11. The instruction execution module according to claim 10, wherein, The reservation station includes: an exception search unit and an instruction re-fusion unit; The exception search unit is used to determine the exception memory access instruction from the n memory access instructions based on the exception execution signal; The instruction re-fusion unit is used to fuse at least two memory access instructions among the n memory access instructions that have not started execution into one instruction to obtain a second merged instruction. The function of the second merged instruction is equivalent to the functions of the at least two memory access instructions that have not started execution, and issues the second merged instruction to the pipeline, occupying one micro-instruction bandwidth.

12. The instruction execution module according to claim 10 or 11, wherein, The reservation station further includes a second buffer, and the second buffer is used to temporarily store the first merged instruction and other memory access instructions among the at least one candidate memory access instruction other than the n memory access instructions; Each entry in the second buffer is used to store the instruction content of a merge instruction or the instruction content of one of the other memory access instructions; the instruction content includes: instruction identification information, register label information, immediate value information, and synchronous execution information. The instruction identification is used to characterize the instruction type of the instruction stored in the entry, the register label information is used to characterize the physical register of the instruction stored in the entry, the immediate value information is used to characterize the memory access address of the instruction stored in the entry, and the synchronous execution information is used to characterize the number of memory access instructions synchronously executed in the pipeline.

13. The instruction execution module according to any one of claims 1 to 12, wherein, The instruction execution module further includes: a write-back buffer; The write-back buffer is used to write the execution result of the instruction to be written back to the general-purpose register through a first write-back path, and the first write-back path is adapted to execute a single memory access instruction in one clock cycle of the pipeline; The write-back buffer is further used to, when there is a second write-back path between the write-back buffer and the general-purpose register, write the execution results of different instructions to be written back to the general-purpose register through the first write-back path and the second write-back path respectively, and the second write-back path is adapted to the change in execution speed of the pipeline based on the first merge instruction; Wherein, the instruction to be written back refers to a memory access instruction that has been executed in the pipeline.

14. The instruction execution module according to any one of claims 1 to 13, wherein, The instruction execution module further includes a write-back buffer and a buffer bypass network. The write-back buffer is coupled to the buffer bypass network, and the buffer bypass network is coupled to the read port of the pipeline; The buffer bypass network is used to transfer the execution result of the instruction to be written back to the read port of the pipeline before the execution result of the instruction to be written back is written back to the general-purpose register. The instruction to be written back refers to a memory access instruction that has been executed in the pipeline.

15. A chip, the chip includes a memory and the instruction execution module according to any one of claims 1 to 14.

16. An electronic device, the electronic device includes a processor, and the processor includes the instruction execution module according to any one of claims 1 to 14.

17. An instruction execution method applied to an instruction execution module, the instruction execution module comprising: A dispatch queue and a pipeline, the method includes: The dispatch queue stores at least one memory access instruction to be executed; The dispatch queue selects n memory access instructions from the at least one memory access instruction to be executed, where n is an integer greater than 1; The dispatch queue merges the n memory access instructions into one instruction to obtain a first merge instruction, and the function of the first merge instruction is equivalent to the function of the n memory access instructions; The pipeline executes the first merge instruction.

Citation Information

Patent Citations

  • Execution method for logarithmic loading instruction

    CN108845830A

  • Fusion of microprocessor storage instructions

    CN116194885A

  • Instruction execution module, chip, equipment and method applied to processor

    CN119127312A

  • Handling and fusing load instructions in a processor

    US11249757B1

  • Method and apparatus for dynamic allocation of multiple buffers in a processor

    US5778245A