Pre-decoding method for variable-length instruction set, electronic equipment and storage medium

By sliding the window on the cache block to retrieve and decode the variable-length instruction set, the problem of difficulty in parsing the variable-length instruction set is solved, improving decoding efficiency and reducing chip area.

CN120255958AActive Publication Date: 2025-07-04METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510725580.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently parse variable-length instruction sets, which leads to difficulty in designing instruction prefetching and decoding units in chip design.

Method used

The atomic window is used to slide the instruction data of the preset length on the cache block and predecode it. By obtaining the length of the first instruction and the sliding start position, the variable-length instruction set is extracted one by one.

Benefits of technology

It improves the decoding efficiency of variable-length instruction sets, and can flexibly adapt to various variable-length instruction sets, reducing chip area footprint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255958A_ABST
    Figure CN120255958A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chip design, in particular to a pre-decoding method for a variable-length instruction set, electronic equipment and a storage medium, and the method comprises the steps: taking out N cache blocks from a memory, obtaining R cascaded atomic windows, and configuring a window length wid for each atomic window; inputting the same cache block into the R-level atomic window for pre-decoding to obtain instruction data of the first instruction in the atomic window; wherein the pre-decoding step of obtaining the instruction data of the first instruction in the atomic window through the jth sliding in the ith cache block cache i comprises the following steps: according to the initial position, the atomic window slides on the cache block to take out the instruction data with the fixed length, pre-decoding is carried out to obtain the instruction data of the first instruction and the length of the instruction data, and the initial position of the next sliding is output; according to the method, the instruction data of each variable-length instruction is extracted, so that the decoding efficiency of the variable-length instruction set is improved, and the method can be flexibly adapted to various variable-length instruction sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip design, and in particular to a pre-decoding method, electronic equipment and storage medium for a variable-length instruction set. Background Art

[0002] In modern processor architectures, variable-length instruction sets are widely used in mainstream instruction set architectures such as x86 and RISC-V due to their high code density and flexible instruction encoding characteristics. Unlike fixed-length instruction sets, the encoding length of variable-length instructions changes dynamically according to the opcode type and operand combination, which poses a severe challenge to the design of instruction prefetch and decoding units.

[0003] When the chip executes instructions, a fixed-length data needs to be obtained from the memory before issuing the instruction. However, for variable-length instruction sets, it is extremely difficult to determine how many instructions are contained in this fixed-length data. Due to the different lengths of instructions, this fixed-length data may contain instructions of various lengths, or even a partial fragment of an instruction. This results in the inability to simply implement instruction segmentation based on fixed-step shifts like fixed-length instruction sets when parsing variable-length instruction sets, making the parsing process difficult. Therefore, how to efficiently and accurately parse variable-length instruction sets has become a key issue that needs to be solved in chip design. Summary of the invention

[0004] In view of the above technical problems, the technical solution adopted by the present invention is: a pre-decoding method of a variable-length instruction set, the method comprising the following steps: S100, fetch N cache blocks from the memory, where N is greater than or equal to 1.

[0005] S200, obtaining R cascaded atomic windows, configuring a window length wid for each of the atomic windows; the starting position of the next sliding output of the atomic window of the current level is used as the input of the atomic window of the next level, where R is greater than or equal to 1.

[0006] S300, input the same cache block into the atomic window of level R for pre-decoding, and obtain the instruction data of the first instruction in the atomic window; wherein the i-th cache block cache_line i The pre-decoding steps of obtaining the instruction data of the first instruction in the atomic window by the j-th slide include: S310, get cache_line i The starting position index of the jth slide i,j .

[0007] S320: cache_line i and its index i,j Enter the Atom window.

[0008] S330, the atomic window obtains the cache_line i in which the instruction data data i,j starting from index and with a length of wid i,j .

[0009] S340, pre-decode the data i,j to obtain the length of the first instruction.

[0010] S350, intercept the data according to the length of the first instruction i,j , obtain the instruction data of the first instruction, and generate the starting position index of the (j + 1)-th sliding i,j+1 .

[0011] In addition, the present invention also provides a non-transitory computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the above method.

[0012] In addition, the present invention also provides an electronic device, including a processor and the above non-transitory computer-readable storage medium.

[0013] The present invention has at least the following beneficial effects: The present invention provides a pre-decoding method, an electronic device and a storage medium for a variable-length instruction set. By sliding an atomic window on a cache block to fetch instruction data of a preset length and performing pre-decoding to obtain the instruction data and its length of the first instruction, the atomic window intercepts the instruction data according to the instruction length of the first instruction and calculates the starting position of the next sliding. By this method, each variable-length instruction is extracted, thereby improving the decoding efficiency of the variable-length instruction set, and the length of the sliding window and the number of cascades can be flexibly configured, and can be flexibly adapted to various variable-length instruction sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0015] Figure 1 is a flowchart of a pre-decoding method for a variable-length instruction set provided by an embodiment of the present invention; Figure 2 is the cache_line provided by an embodiment of the present invention iFlowchart of the pre - decoding step for obtaining the instruction data of the first instruction in the atomic window during the j - th sliding in Figure 3 Schematic diagram of the structure of a single atomic window provided by the present invention; Figure 4 Schematic diagram of the structure of R cascaded atomic windows provided by the present invention. Detailed implementation manners

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0017] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meaning as commonly understood by those skilled in the art.

[0018] Please refer to Figure 1 , which shows a flowchart of a pre - decoding method for a variable - length instruction set. The method includes the following steps: S100, fetch N cache blocks from the memory, where N is greater than or equal to 1.

[0019] Among them, a cache line is the basic unit for data storage and transmission in the cache, enabling data to be transmitted and stored in units of blocks between the cache and the memory. A cache block includes multiple instruction data with the same or different lengths. Since the length of the cache block is a fixed length, that is, the instruction data is evenly distributed into multiple cache blocks with a fixed length. Therefore, a complete instruction may be mechanically divided into the tail of the current cache block and the head of the next cache block.

[0020] S200, obtain R cascaded atomic windows, and configure a window length wid for each of the atomic windows; the starting position of the next sliding output by the atomic window of the current level is used as the input of the atomic window of the next level, where R is greater than or equal to 1.

[0021] Among them, an atomic window is a window structure for fetching instruction data with a length of wid from a cache block and performing pre - decoding on the fetched instruction data. The wid configured for the atomic window determines the amount of instruction data fetched from the cache block each time. The pre - decoding operation of the atomic window on the fetched data is an indivisible whole and will not be interrupted or interfered by other operations. The input of each atomic window is the cache block and the starting position of the next sliding, and the output is the pre - decoded instruction and the starting position of the next set of sliding.

[0022] In one embodiment, wid is the maximum length among all instruction lengths in the current instruction set. It can be flexibly configured according to changes in the instruction set.

[0023] Among them, each atomic window may include a set of instruction data of wid fetched from cache_line i As an example, when the instructions in the instruction set include instruction A, instruction B, and instruction C, the length of instruction A is 32 bit, the length of instruction B is 64 bit, and the length of instruction C is 96 bit. Since the length of instruction C is the largest, the length of the atomic window is 96 bit. The instruction data in the cashline fetched by the atomic window may be the instruction data of 3 instruction As, or the instruction data of instruction A and instruction B, or the instruction data of instruction C, or may also be the instruction data of instruction B and part of instruction C. During the pre-decoding process, only the instruction data of the first instruction will be retained, and the instruction data of the remaining instructions will be discarded.

[0024] Among them, the starting position of the next slide refers to the starting position of the first instruction in the instruction data fetched by the next atomic window.

[0025] In one embodiment, please refer to Figure 3 , when R is equal to 1, when the atomic window slides for the first time, the starting position of the first slide is the initial position of the cache block; after inputting the initial position of the first slide and the cache block into the atomic window, the starting position of the next slide output by the atomic window is used as the input of the atomic window. Among them, using one atomic window for pre-decoding can not only achieve pre-decoding of variable-length instruction sets, but also, compared with multiple atomic windows, greatly reduce the chip area occupied.

[0026] In one embodiment, please refer to Figure 4 , in S200, when R is greater than 1, when the atomic window slides for the first time, the starting position of the first slide is the initial position of the cache block; input the initial position of the first slide into the first-level atomic window, and input the cache block into R atomic windows simultaneously, and the R-level atomic windows perform synchronous pre-decoding; among them, the starting position of the next slide output by the r-th level atomic window is used as the input of the (r + 1)-th level atomic window. Using R cascaded atomic windows can not only achieve pre-decoding of variable-length instruction sets, but also, compared with a single atomic window, the R cascaded atomic windows can complete pre-decoding within the same clock cycle, greatly improving the efficiency of instruction pre-decoding.

[0027] When the chip area is tight or the efficiency requirement is low, the value of R can be configured to a lower value. When the chip area is abundant or the efficiency requirement is high, the value of R can be configured to a higher value, and the number of cascades can be flexibly configured according to the needs of the chip.

[0028] S300, input the same cache block into the R-level atomic window for pre-decoding to obtain the instruction data of the first instruction in the atomic window.

[0029] Further, please refer to Figure 2 , which shows the pre-decoding steps for obtaining the instruction data of the first instruction in the atomic window by the j-th sliding of the i-th cache block cache_line i include: S310, obtain the starting position index i of the j-th sliding in cache_line i,j .

[0030] S320, input the cache_line i and its index i,j into the atomic window.

[0031] S330, the atomic window obtains the instruction data data i in cache_line i,j starting from index i,j with a length of wid.

[0032] Among them, when the remaining undecoded instruction data in cache_line i is incomplete, it needs to be concatenated first and then pre-decoded through the atomic window.

[0033] In one implementation, S330 further includes: S331, determine whether the instruction data starting from index i in cache_line i,j is complete. If not, wait to retrieve cache_line i+1 .

[0034] In one implementation, when the length of the instruction data starting from index i,j is less than the length configured by the length field segment carried by the current instruction, it is determined that the instruction data is incomplete.

[0035] S332, when cache_line i+1 is retrieved, concatenate cache_line i and cache_line i+1 into a whole.

[0036] S333, the atomic window obtains the instruction data data i,j starting from index i,j with a length of wid in the concatenated whole.

[0037] The problems of instruction fragments can be correctly handled through S331 - S333.

[0038] S340, for the said data i,j perform pre - decoding to obtain the length of the first instruction.

[0039] In one embodiment, the length of the first instruction is the length configured in the length field segment of the first instruction.

[0040] In one embodiment, the length of the first instruction is the fixed length of the instruction specified in the instruction set.

[0041] S350, intercept the said data according to the length of the first instruction i,j , obtain the instruction data of the first instruction, and generate the starting position index of the (j + 1)-th sliding i,j+1 .

[0042] Among them, by intercepting the corresponding instructions in the sliding window manner and performing pre - decoding, the system can more accurately parse the corresponding instructions, breaking through the confinement of only being able to parse fixed - length instructions.

[0043] It should be noted that wid is the maximum length among the lengths of all instructions in the instruction set. Therefore, no matter how many instructions are included in an atomic window, only the instruction data of the first instruction is retained according to the instruction length of the first instruction, and the instruction data of other instructions is directly discarded. In this way, the instruction data of each variable - length instruction is stripped from the cache block to complete pre - decoding, making the subsequent decoding process more efficient. It should be noted that the discarded instruction data in the atomic window does not affect the original instruction data in the cache block. The starting position of the next sliding actually starts sliding from the next address of the first instruction obtained by the current sliding, that is, the starting position of the next instruction. By analogy, all variable - length instructions are stripped one by one to complete the task of pre - decoding, so no instruction loss will occur.

[0044] In one embodiment, in S350, index i,j+1 satisfies: index i,j+1 = index i,j + size i,j , where size i,j is the instruction length of the first instruction obtained by the j - th sliding pre - decoding in cache_line i .

[0045] In one embodiment, S350 further includes: judging whether it is a branch jump instruction according to the operation code of the first instruction. If so, start the branch prediction processing operation to improve the efficiency of the pipeline. Please refer to againFigure 2 and Figure 3 In the result output by the atomic window, it also includes the judgment result of whether the current instruction obtained by pre-decoding is a branch jump instruction.

[0046] Among them, when the end position of the first instruction in the atomic window is exactly the end address of cache_line i or when the end position of the first instruction in the atomic window exceeds the end address of cache_line i the index i,j+1 of the next slide points to the next cache block.

[0047] In one implementation, S350 further includes: when the end position of the first instruction obtained by the j-th slide pre-decoding is the end address of the current cache block or crosses the end address of the current cache block, then the first instruction obtained by the j-th slide pre-decoding is the last instruction in cache_line i and the index i,j+1 is the starting position in the (i + 1)-th cache block cache_line i+1 . Please refer to Figure 2 and Figure 3 again. In the result output by the atomic window, it also includes the judgment result of whether the instruction obtained by pre-decoding is the last instruction.

[0048] In summary, the present invention provides a pre-decoding method for a variable-length instruction set. By sliding an atomic window on a cache block to fetch instruction data of a preset length and performing pre-decoding to obtain the instruction data and its length of the first instruction, the atomic window intercepts the instruction data according to the instruction length of the first instruction and calculates the starting position of the next slide. By this method, each variable-length instruction is extracted, thereby improving the decoding efficiency of the variable-length instruction set, and the length of the sliding window and the number of cascades can both be flexibly configured, and can flexibly adapt to various variable-length instruction sets.

[0049] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one segment of a program related to a method in a method embodiment. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement the method provided in the above embodiment.

[0050] An embodiment of the present invention also provides an electronic device, including a processor and the aforementioned non-transitory computer-readable storage medium.

[0051] An embodiment of the present invention also provides a computer program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps in the methods according to various exemplary embodiments of the present invention described above in this specification.

[0052] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0053] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention disclosed is defined by the appended claims.

Claims

1. A pre - decoding method for a variable - length instruction set, characterized in that, The method includes the following steps: S100, fetch N cache blocks from memory, where N is greater than or equal to 1; S200, obtain R cascaded atomic windows, and configure a window length wid for each of the atomic windows; the starting position of the next slide output by the atomic window at the current level is used as the input of the atomic window at the next level, where R is greater than or equal to 1; S300, input the same cache block into the R-level atomic windows for pre-decoding to obtain the instruction data of the first instruction in the atomic window; Wherein the i-th cache block cache_line i The pre-decoding step of obtaining the instruction data of the first instruction in the atomic window by the j-th sliding in it includes: S310, obtain the cache_line i The starting position index of the j-th sliding in i,j ; S320, input the cache_line i and its index i,j into the atomic window; S330, the atomic window obtains cache_line i in which the instruction data data i,j starting from index i,j ; S340, pre-decode the said data i,j to obtain the length of the first instruction; S350, intercept the data according to the length of the first instruction i,j , obtain the instruction data of the first instruction, and generate the starting position index of the (j + 1)-th slide i,j+1 .

2. The method according to claim 1, wherein In S350, index i,j+1 Satisfies: index i,j+1 = index i,j + size i,j where size i,j is the instruction length of the first instruction obtained by the j-th sliding pre-decoding in cache_line i in the 3. The method according to claim 1, characterized in that, In S200, when R is equal to 1, when the atomic window slides for the first time, the starting position of the first slide is the initial position of the cache block; After inputting the initial position of the first slide and the cache block into the atomic window, the starting position of the next slide output by the atomic window is used as the input of the atomic window.

4. The method according to claim 1, characterized in that, In S200, when R is greater than 1, when the atomic window slides for the first time, the starting position of the first slide is the initial position of the cache block; input the initial position of the first slide into the first-level atomic window, input the cache block into the R atomic windows simultaneously, and perform synchronous pre-decoding on the R-level atomic windows; wherein, the starting position of the next slide output by the r-level atomic window is used as the input of the (r + 1)-level atomic window.

5. The method according to claim 1, characterized in that S350 further includes: judging whether it is a branch jump instruction according to the operation code of the first instruction, and if so, starting a branch prediction processing operation.

6. The method according to claim 1, wherein S330 further includes: S331, determine whether the instruction data starting from index in cache_line is complete. If it is not complete, wait to retrieve cache_line i in i,j cache_line is complete. If not, wait to retrieve cache_line i+1 ; S332, when cache_line i+1 is retrieved, combine cache_line i and cache_line i+1 into a whole; S333, the atomic window obtains the instruction data data i,j starting from index and with a length of wid in the spliced whole i,j .

7. The method according to claim 1, characterized in that, The S350 further includes: when the end position of the first instruction obtained by the j-th sliding pre-decoding is the end address of the current cache block or crosses the end address of the current cache block, the first instruction obtained by the j-th sliding pre-decoding is the last instruction in cache_line i in, and the index i,j+1 is the starting position in the (i + 1)-th cache block cache_line i+1 in.

8. The method according to claim 1, wherein wid is the maximum length among the lengths of all instructions in the current instruction set.

9. A non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, characterized in that, The at least one instruction or the at least one program is loaded and executed by a processor to implement the method according to any one of claims 1-8.

10. An electronic device, characterized in that, It includes a processor and the non-transitory computer-readable storage medium described in claim 9.

Citation Information

Patent Citations

  • Variable length instruction booting for instruction decoding clusters

    CN116893848A

  • Instruction coding method and device, electronic equipment and medium

    CN118796270A

  • Techniques for Storing Instructions and Related Information in a Memory Hierarchy

    US20080256338A1

  • Method and Apparatus for Length Decoding and Identifying Boundaries of Variable Length Instructions

    US20090019257A1

  • Method and system configuration for simplifying the decoding system for access to an register file with overlapping windows

    US5440714A