A zero-overhead hardware loop processor, method, and storage medium

A simplified zero overhead hardware loop processor with a loop module and cache addresses complexity in RISC processors, enhancing efficiency by eliminating redundant fetching and speeding up loop execution.

CN114185601BActive Publication Date: 2025-07-15SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111509131.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-10
Publication Date
2025-07-15
Estimated Expiration
2041-12-10

AI Technical Summary

Technical Problem

The existing zero-overhead hardware loop processor implemented by the thin-instruction set is highly complex in hardware design and requires a simpler implementation method.

Method used

The zero-overhead loop module is added to the processor structure, and the cache module is configured to store the loop body. The cache module extracts loop instructions in the subsequent loop process, eliminating the repeated operations of the fetching module.

Benefits of technology

The processor structure is simplified, the extraction speed and execution efficiency of cyclic instructions are improved, the operation of the finger fetch module is reduced, and the efficiency of zero-overhead hardware cycles is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114185601B_ABST
    Figure CN114185601B_ABST
Patent Text Reader

Abstract

The present invention provides a zero-overhead hardware loop processor, method and storage medium. The structure of the processor includes: an instruction fetch module configured to obtain instructions in response to an instruction fetch request and forward them; a zero-overhead loop module configured to determine whether the obtained instruction is a loop instruction, extract the loop count in the obtained loop instruction, and load a cache to save the loop instruction as a loop body, and sequentially forward the loop instructions in the loop body, and decrement the loop count by 1 each time the loop body is forwarded; a decoding module configured to decode the received instruction; and an execution module configured to perform corresponding operations according to the decoded instruction. The zero-overhead hardware loop processor of the present invention has the characteristics of small modification and simple structure, and can extract the loop body at one time so that it is not necessary to obtain instructions through the instruction fetch module during subsequent loop instruction fetching, which can effectively improve the efficiency of zero-overhead hardware loops.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of processor design, and in particular, to a zero-overhead hardware loop processor, method, and storage medium. Background Art

[0002] Currently, most RISC (Reduced Instruction Set Computer) supports the Zero Overhead Hardware Loop instruction. The idea of the zero-overhead hardware loop is that through the direct participation of hardware, by setting certain loop count registers, the program can then automatically loop, and each time it loops, the loop count register is automatically decremented by 1. This continuous loop until the value of the loop count register becomes 0, then the loop exits.

[0003] However, the hardware design of the zero-overhead hardware loop processor implemented based on the existing reduced instruction set is relatively complex. Therefore, there is an urgent need for a simpler processor, method, and / or reduced instruction set that can implement the zero-overhead hardware loop based on the reduced instruction set. Summary of the Invention

[0004] In order to implement a simpler processor, method, and / or reduced instruction set that can implement the zero-overhead hardware loop based on the reduced instruction set. In one aspect of the present invention, a zero-overhead hardware loop processor is proposed, including: an instruction fetch module configured to obtain and forward an instruction in response to an instruction fetch request; a zero-overhead loop module configured to determine whether the obtained instruction is a loop instruction, extract the loop count in the obtained loop instruction, and load a cache to save the loop instruction as a loop body, and sequentially forward the loop instructions in the loop body, and decrement the loop count by 1 each time the loop body is forwarded; a decoding module configured to decode the received instruction; and an execution module configured to perform corresponding operations according to the decoded instruction.

[0005] In one or more embodiments, the zero-overhead loop module includes: a loop decoding module configured to determine whether the received instruction is a loop instruction, decode the loop instruction and extract the loop count therein, and forward the loop instruction and the loop count; a loop storage module configured to load a cache module and generate a loop body to store the loop instruction; a loop control module configured to manage the loop count and forward the loop instruction, and decrement the loop count by 1 each time the loop body is forwarded; and a cache module configured to temporarily save the loop body.

[0006] In one or more embodiments, the loop decoding module is further configured to, in response to the received instruction being a non-loop instruction, directly forward the non-loop instruction to the loop control module without processing the non-loop instruction.

[0007] In one or more embodiments, the loop control module is further configured to control the loop storage module to clear the loop body to release the cache when the number of loops is zero.

[0008] In one or more embodiments, the zero-overhead hardware loop processor of the present invention further includes: a memory access module configured to access a corresponding data storage module; and a write-back module configured to obtain the result of performing a corresponding operation and write it back to a corresponding register.

[0009] In a second aspect of the present invention, a zero-overhead hardware loop method is proposed. The method includes: configuring a loop body in a reduced instruction set, where the loop body includes a first reduced instruction with consecutive fetch addresses, a loop instruction, and a second reduced instruction; in response to the processor receiving a fetch request, obtaining an instruction from the reduced instruction set; in response to the obtained instruction being the first reduced instruction, performing consecutive fetching and storing the obtained loop instruction and the first reduced instruction in a loaded cache; in response to the obtained instruction being the second reduced instruction, stopping the consecutive fetching operation and storing the second reduced instruction in the loaded cache to form a loop body; and controlling the execution of a zero-overhead hardware loop based on the loop instruction in the loop body.

[0010] In one or more embodiments, the first reduced instruction includes a loop instruction start flag and the number of loops; correspondingly, the method further includes: extracting the number of loops in the first reduced instruction, and decrementing the number of loops by one each time the zero-overhead hardware loop is controlled to execute based on the loop body; and clearing the loop body to release the cache in response to the number of loops being zero.

[0011] In one or more embodiments, the method further includes: in response to the need to fetch an instruction again when the number of loops is not zero, fetching an instruction from the loop body in the cache.

[0012] In one or more embodiments, the second reduced instruction includes a loop instruction end flag; correspondingly, the decrementing of the number of loops by one each time the zero-overhead hardware loop is controlled to execute based on the loop body includes: decrementing the number of loops by one each time the second reduced instruction including the instruction end flag is fetched.

[0013] In a third aspect of the present invention, a storage medium is proposed. A computer program that can run is stored in the storage medium, and when the computer program is executed, it is used to implement the steps of the zero-overhead hardware loop method in any one of the above embodiments.

[0014] The beneficial effects of the present invention include: The zero-overhead hardware loop processor of the present invention only needs to add a zero-overhead loop module 200 between the instruction fetch module and the decoding module on the basis of the existing processor structure, so it has the characteristics of small modification and simple structure; and the zero-overhead loop module 200 of the present invention is also separately configured with a cache module 204, and the loop body can be extracted into the cache module 204 at one time, so that the loop instructions can be extracted from the cache module in the subsequent loop instruction fetch process, and there is no need to obtain instructions through the instruction fetch module anymore. Except for the first instruction fetch that needs to go through the instruction fetch module, the instruction fetch operation of the instruction fetch module is omitted in the subsequent multiple instruction loop processes, so the efficiency is higher, and because the loop body can be stored in the specified cache module local to the processor, the instruction extraction speed is also faster, which helps to further improve the efficiency of the zero-overhead hardware loop. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.

[0016] Figure 1 is a schematic structural diagram of the zero-overhead hardware loop processor of the present invention;

[0017] Figure 2 is a working flowchart of the zero-overhead hardware loop method of the present invention;

[0018] Figure 3 is a schematic structural diagram of the loop body of the present invention;

[0019] Figure 4 is a schematic structural diagram of a readable storage medium of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings.

[0021] It should be noted that all the expressions using "first" and "second" in the embodiments of the present invention are used to distinguish two entities or parameters with the same name but different, so "first" and "second" are only for the convenience of expression and should not be construed as a limitation on the embodiments of the present invention. This will not be repeated in the subsequent embodiments.

[0022] Figure 1This is a schematic structural diagram of the zero-overhead hardware loop processor of the present invention. As Figure 1 shown, the structure of the zero-overhead hardware loop processor of the present invention includes: an instruction fetch module 100, a zero-overhead loop module 200, a decoding module 300, an execution module 400, a memory access module 500, and a write-back module 600 that are connected in sequence; among them, the instruction fetch module 100 is configured to obtain an instruction in response to an instruction fetch request and forward it; the zero-overhead loop module 200 is configured to determine whether the obtained instruction is a loop instruction, extract the loop count in the obtained loop instruction, load a cache to save the loop instruction as a loop body, and forward the loop instructions in the loop body in sequence, and decrement the loop count by 1 each time the loop body is forwarded; the decoding module 300 is configured to decode the received instruction; the execution module 400 is configured to perform corresponding operations according to the decoded instruction.

[0023] In one embodiment, the zero-overhead loop module 200 of the present invention includes: a loop decoding module 201 connected to the instruction fetch module 100, a loop storage module 202 connected to the loop decoding module 201, a loop control module 203 respectively connected to the loop decoding module 201 and the loop storage module 202, and a cache module 204 connected to the loop storage module 202; among them, the loop decoding module 201 is configured to receive the instruction forwarded by the instruction fetch module 100, determine whether the received instruction is a loop instruction, decode the loop instruction and extract the loop count therein, and forward the loop instruction to the loop storage module 202 and the loop control module 203, and forward the loop count to the loop control module 203, wherein the loop control module 203 includes a register for storing and managing the loop count; the loop storage module 202 is configured to load the cache module 204 and generate a loop body to store the loop instruction; the loop control module 203 is configured to manage the loop count and forward the loop instruction to the decoding module 300, and decrement the loop count by 1 each time the loop body is forwarded, wherein the specific process of decrementing the loop count by 1 each time the loop body is forwarded is that the loop count is decremented by 1 each time all the loop instructions in the loop body are forwarded. In addition, the cache module 204 in this embodiment is only controlled by the loop storage module 202 and is not used to cache other data except loop instructions.

[0024] The structure of the existing processor includes an instruction fetch module, a decoding module, an execution module, a memory access module, and a write-back module connected in sequence. By comparison, it can be seen that the zero-overhead hardware loop processor of the present invention only needs to add a zero-overhead loop module 200 between the instruction fetch module and the decoding module on the basis of the existing processor structure. Therefore, it has the characteristics of small modification and simple structure. Moreover, the zero-overhead loop module 200 of the present invention is also separately configured with a cache module 204, and the loop body can be extracted into the cache module 204 at one time, so that in the subsequent loop instruction fetch process, loop instructions can be extracted from this cache module, and there is no need to obtain instructions through the instruction fetch module anymore. In this way, except for the first instruction fetch that needs to go through the instruction fetch module, the instruction fetch operation of the instruction fetch module is omitted in the subsequent multiple instruction loop processes. Therefore, the efficiency is higher, and since the loop body can be stored in the specified cache module local to the processor, the instruction extraction speed will also be faster, which helps to further improve the execution efficiency of the zero-overhead hardware loop.

[0025] In one embodiment, the loop decoding module 201 is further configured to, in response to the received instruction being a non-loop instruction, not process the non-loop instruction but directly forward the non-loop instruction to the loop control module 203, and then the loop control module 203 forwards it to the decoding module 300.

[0026] In one embodiment, the loop control module 203 is further configured to control the loop storage module to clear the loop body to release the cache when the number of loops is zero. In an alternative embodiment, when the size of the loop body is greater than the maximum effective storage capacity of the cache module to store the loop body, an error message is generated by the loop storage module and forwarded to the loop control module, and then the loop control module forwards it to the subsequent module, and the error message is written into the corresponding register through the memory access module. Relevant personnel can discover the error message in the register through methods such as log monitoring and adjust the size of the loop body.

[0027] In one embodiment, the memory access module 500 of the present invention is configured to access the corresponding data storage module; the write-back module 600 is configured to obtain the result of performing the corresponding operation and write it back to the corresponding register.

[0028] On the basis of the zero-overhead hardware loop processing proposed in the above embodiments, the present invention also proposes a zero-overhead hardware loop method, which aims to propose a way to implement the above zero-overhead hardware loop based on a reduced instruction set. Figure 2 This is the flowchart of the working process of the zero-overhead hardware loop method of the present invention. As Figure 2As shown, the overhead hardware loop method of the present invention includes the following steps: Step 10 configures a loop body in the reduced instruction set, and the loop body includes a first reduced instruction, a loop instruction, and a second reduced instruction with consecutive fetch addresses; Step 20, in response to the processor receiving a fetch request, obtains an instruction from the reduced instruction set; Step 30, in response to the obtained instruction being the first reduced instruction, performs consecutive fetching and stores the obtained loop instruction and the first reduced instruction in the loaded cache; Step 40, in response to the obtained instruction being the second reduced instruction, stops the consecutive fetching operation and stores the second reduced instruction in the loaded cache to form a loop body; Step 50, controls the execution of a zero-overhead hardware loop based on the loop instruction in the loop body. In this embodiment, the present invention sets a loop body structure determined by two instructions, namely the first reduced instruction and the second reduced instruction, in the reduced instruction set based on the RISC-V architecture. The fetch address of the loop instruction in the loop body structure is located between the fetch addresses of the first reduced instruction and the second reduced instruction. In response to the instruction corresponding to the fetch request being the first reduced instruction, the fetch module performs consecutive fetching starting from the fetch address of the first reduced instruction until the second reduced instruction is fetched, thereby achieving the one-time extraction of all required loop instructions. In this way, the present invention can extract loop instructions without performing complex address and address offset calculations, and can extract all required loop instructions at one time.

[0029] In one embodiment, the method of the present invention further includes: when it is necessary to fetch again in response to the loop count not being zero, fetch from the loop body in the cache. Combining with the previous embodiment, this embodiment saves the loop instructions obtained at one time in the cache, so that when it is necessary to fetch again in response to the loop count not being zero, the loop instructions are obtained from the loop body in the cache, which can provide a faster fetch speed and help improve the execution efficiency of the zero-overhead hardware loop.

[0030] In one embodiment, the first reduced instruction includes a loop instruction start flag and a loop count; correspondingly, the method of the present invention further includes: extracting the loop count in the first reduced instruction, and decrementing the loop count by one each time the zero-overhead hardware loop is controlled to execute based on the loop body; clearing the loop body to release the cache in response to the loop count being zero. Among them, the specific process of decrementing the loop count by one each time the zero-overhead hardware loop is controlled to execute based on the loop body is that the loop count is decremented by one each time the zero-overhead hardware loop is controlled to execute based on all the loop instructions in the loop body.

[0031] In one embodiment, the second reduced instruction includes a loop instruction end flag; correspondingly, the operation of decrementing the loop count by one each time the zero-overhead hardware loop is executed based on the loop body control in the above embodiment includes: decrementing the loop count by one each time a second reduced instruction containing the instruction end flag is fetched. In the zero-overhead hardware loop method proposed by the invention, the addressing process of loop instructions is simplified, and the structures of the first reduced instruction and the second reduced instruction are simple, which conforms to the design concept of reduced instructions.

[0032] Figure 3 FIG. is a schematic structural diagram of the loop body of the present invention. As Figure 3 shown, each loop body further includes a first reduced instruction (ZOHLS), a plurality of loop instructions, and a second reduced instruction (ZOHLT): wherein, the first reduced instruction includes a loop instruction start flag and a loop count, and is configured to control consecutive instruction fetch operations; the plurality of loop instructions are configured to control the execution of corresponding operations; the second reduced instruction includes a loop instruction end flag and is configured to control the stop of consecutive instruction fetch operations; wherein, the first reduced instruction, the plurality of loop instructions, and the second reduced instruction have consecutive instruction fetch addresses in the reduced instruction set.

[0033] The loop body of the present invention has the characteristics of simple instructions and convenient instruction fetching, which conforms to the design concept of reduced instructions.

[0034] Figure 4 FIG. is a schematic structural diagram of a readable storage medium of the present invention. As Figure 4 shown, the readable storage medium 700 of the present invention includes a runnable computer program 701, and when the computer program 701 is executed, it is used to implement the steps of the zero-overhead hardware loop method in the above embodiments.

[0035] The above are exemplary embodiments disclosed by the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments disclosed by the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein do not need to be executed in any specific order. In addition, although the elements disclosed in the embodiments of the present invention can be described or claimed in an individual form, they can also be understood as plural unless explicitly limited to the singular.

[0036] It should be understood that, as used herein, unless the context clearly supports an exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations including one or more of the associated listed items.

[0037] The serial numbers of the disclosed embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments.

[0038] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope (including the claims) disclosed by the embodiments of the present invention is limited to these examples; under the concept of the embodiments of the present invention, the technical features in the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omissions, modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.

Claims

1. A zero-overhead hardware loop processor, characterized in that, Comprising: An instruction fetch module configured to obtain and forward an instruction in response to an instruction fetch request; A zero-overhead loop module configured to determine whether the obtained instruction is a loop instruction, extract the number of loops in the obtained loop instruction, and load a cache to save the loop instruction as a loop body, and sequentially forward the loop instructions in the loop body, and decrement the number of loops by 1 each time the loop body is forwarded; A decoding module configured to decode the received instruction; And An execution module configured to perform corresponding operations according to the decoded instruction; Wherein, the zero-overhead loop module includes: a loop decoding module configured to determine whether the received instruction is a loop instruction, decode the loop instruction and extract the number of loops therein, and forward the loop instruction and the number of loops; a loop storage module configured to load a cache module and generate a loop body to store the loop instruction; a loop control module configured to manage the number of loops and forward the loop instruction, and decrement the number of loops by 1 each time the loop body is forwarded; and a cache module configured to temporarily save the loop body; the loop decoding module is further configured to, in response to the received instruction being a non-loop instruction, directly forward the non-loop instruction to the loop control module without processing the non-loop instruction.

2. The zero-overhead hardware loop processor according to claim 1, wherein The loop control module is further configured to control the loop storage module to clear the loop body to release the cache when the number of loops is zero.

3. The zero-overhead hardware loop processor according to claim 1, wherein Further comprising: A memory access module configured to access a corresponding data storage module; And A write-back module configured to obtain the result of performing the corresponding operation and write it back to the corresponding register.

4. A zero-overhead hardware loop method, characterized in that, The method includes: Configuring a loop body in a reduced instruction set, the loop body including a first reduced instruction with consecutive fetch addresses, a loop instruction, and a second reduced instruction; Obtaining an instruction from the reduced instruction set in response to the processor receiving an instruction fetch request; In response to the obtained instruction being the first reduced instruction, performing consecutive instruction fetching and storing the obtained loop instruction and the first reduced instruction in the loaded cache; In response to the obtained instruction being the second reduced instruction, stopping the consecutive instruction fetching operation and storing the second reduced instruction in the loaded cache to form a loop body; Controlling the execution of a zero-overhead hardware loop based on the loop instruction in the loop body; The method further includes: extracting the number of loops in the first reduced instruction, and decrementing the number of loops by 1 each time the zero-overhead hardware loop is controlled to execute based on the loop body; clearing the loop body to release the cache in response to the number of loops being zero; and directly forwarding the non-loop instruction without processing the non-loop instruction in response to the received instruction being a non-loop instruction.

5. The zero-overhead hardware loop method according to claim 4, wherein The method further includes: Fetching an instruction from the loop body in the cache in response to the need to fetch an instruction again when the number of loops is not zero.

6. The zero-overhead hardware loop method according to claim 4, wherein The second reduced instruction includes a loop instruction end flag; correspondingly, the decrementing the number of loops by 1 each time the zero-overhead hardware loop is controlled to execute based on the loop body includes: In response to each fetching of a second reduced instruction including the instruction end flag, decrement the loop count by one.

7. A storage medium, characterized in that, A runnable computer program is stored in the storage medium, and when the computer program is executed, it is used to implement the steps of the zero-overhead hardware loop method according to any one of claims 4-6.

Citation Information

Patent Citations

  • Instruction word processor, zero-overhead circulation processing method, electronic equipment and medium

    CN112835624A