Vector instruction processing method, device, electronic device and storage medium

By splitting and executing the internal operations of vector instructions in the instruction sequence queue, the problem of high resource utilization of vector instructions is solved, and the performance stability and processing efficiency of the processor are improved.

CN119861969BActive Publication Date: 2025-06-20BEIJING VCORE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336980.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the prior art, vector instruction processing requires a large number of processor computing resources, resulting in low processor performance stability, high operating costs and manufacturing costs.

Method used

By assigning empty items to vector instructions in the instruction sequence queue, splitting them into multiple vector instructions internal operations, and assigning empty items separately in the vector instructions secondary queue, these internal operations are performed out of order, and the result is finally written back to the original queue.

Benefits of technology

The out-of-order execution and sequential submission of operations corresponding to each data element in a vector instruction is realized, and the processor resources are rationally utilized, the computing resource occupation required for vector instruction processing is reduced, and the processor's performance stability and processing efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119861969B_ABST
    Figure CN119861969B_ABST
Patent Text Reader

Abstract

The present invention provides a vector instruction processing method, apparatus, electronic device and storage medium, relating to the technical field of data processing. The method includes: allocating an empty entry for a vector instruction in an instruction ordering queue, splitting the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, performing register renaming on each vector instruction internal operation, allocating an empty entry for each vector instruction internal operation in the vector instruction ordering queue respectively, executing the vector instruction internal operations out of order, and after writing each vector instruction internal operation back to the entry corresponding to each vector instruction internal operation in the vector instruction ordering queue, writing the vector instruction back to the entry corresponding to the vector instruction in the instruction ordering queue. The present invention provides a vector instruction processing method, apparatus, electronic device and storage medium, which can more reasonably utilize the computing resources of a processor, thereby reducing the computing resources of the processor required for vector instruction processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a vector instruction processing method, device, electronic equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, users have put forward higher and higher requirements on the performance of artificial intelligence algorithms and data analysis applications, and therefore the requirements on computer processing capabilities are also getting higher and higher. Vector processing technology has become a hot technology that is of common concern to academia and industry because it can process multiple data with one instruction, provide data-level parallelism, and thus greatly improve data processing efficiency.

[0003] The bit width of a vector instruction indicates the total width of data that can be processed by a single vector instruction at one time. The bit width of a vector instruction is relatively large, usually up to 128 bits, 256 bits, or 512 bits. The high parallelism and complexity of vector instruction processing result in more processor computing resources being required for vector instruction processing, which has a significant impact on the processor's main frequency, chip area, and power consumption.

[0004] Therefore, how to reduce the processor computing resources required for vector instruction processing, thereby improving the performance stability of the processor and reducing the operating cost and manufacturing cost of the processor, is a technical problem that needs to be urgently solved in this field. Summary of the invention

[0005] The present invention provides a vector instruction processing method, device, electronic device and storage medium, which are used to solve the defects of the prior art in that the high parallelism and complexity of vector instruction processing result in a large number of processor computing resources required for vector instruction processing, which has a significant impact on the processor's main frequency, chip area, power consumption and other performance, and reduces the processor computing resources required for vector instruction processing, thereby improving the performance stability of the processor and reducing the operating cost and manufacturing cost of the processor.

[0006] The present invention provides a vector instruction processing method, comprising the following steps.

[0007] In the case where a vector instruction enters the vector instruction issue queue, allocate a blank entry in the instruction ordering queue as the entry corresponding to the vector instruction in the instruction ordering queue; split the vector instruction into internal vector instruction operations corresponding to each data element in the vector instruction, perform register renaming on the internal vector instruction operations corresponding to each data element, and allocate a blank entry in the vector instruction ordering queue for each internal vector instruction operation corresponding to each data element as the entry corresponding to the internal vector instruction operation corresponding to each data element in the vector instruction ordering queue; execute out of order the internal vector instruction operations corresponding to each data element; after writing back each internal vector instruction operation corresponding to each data element to the entry corresponding to the internal vector instruction operation corresponding to each data element in the vector instruction ordering queue, write back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue.

[0008] According to a vector instruction processing method provided by the present invention, the out-of-order execution of the internal vector instruction operations corresponding to each data element includes: adding the internal vector instruction operations corresponding to each data element to the internal vector instruction issue queue, and when the internal vector instruction issue queue determines that the source operands of the internal vector instruction operations corresponding to each data element are all ready, using an out-of-order issue mechanism to issue the internal vector instruction operations corresponding to each data element, so that the internal vector instruction operations corresponding to each data element read the vector register file to obtain the values of the source operands of the internal vector instruction operations corresponding to each data element; based on the operation types of the internal vector instruction operations corresponding to each data element, respectively allocate the internal vector instruction operations corresponding to each data element to different execution units, so that the different execution units execute the internal vector instruction operations corresponding to each data element based on the values of the source operands of the internal vector instruction operations corresponding to each data element.

[0009] According to a vector instruction processing method provided by the present invention, the splitting of the vector instruction into internal vector instruction operations corresponding to each data element in the vector instruction includes: splitting the vector instruction into internal vector instruction operations corresponding to each data element based on the configuration information of the vector control and status register of the vector instruction and the vector instruction opcode of the vector instruction.

[0010] According to a vector instruction processing method provided by the present invention, after writing back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue, the method further includes: submitting the vector instruction in order through the instruction ordering queue;

[0011] Notify the instruction sequencing queue to release the entry corresponding to the vector instruction in the instruction sequencing queue, and notify the vector instruction sequencing queue to release the entry corresponding to the internal operation of each vector instruction corresponding to the data element in the vector instruction sequencing queue.

[0012] According to a vector instruction processing method provided by the present invention, the method further includes: when a scalar instruction enters the scalar instruction issue queue, allocate an empty entry in the instruction sequencing queue for the scalar instruction as the entry corresponding to the scalar instruction in the instruction sequencing queue; add the scalar instruction to the scalar instruction issue queue, and when the scalar instruction issue queue determines that the source operands of the scalar instruction are ready, issue the scalar instruction for the scalar instruction to read the scalar register file to obtain the values of the source operands of the scalar instruction; allocate the scalar instruction to an execution unit for the execution unit to execute the scalar instruction based on the values of the source operands of the scalar instruction; write back the scalar instruction to the entry corresponding to the scalar instruction in the instruction sequencing queue.

[0013] According to a vector instruction processing method provided by the present invention, the method further includes: when a prediction execution error occurs in the processor, based on the vector instruction sequencing queue and / or the instruction sequencing queue, obtain the order relationship between the error instruction with a prediction execution error and other instructions; based on the order relationship between the error instruction and other instructions, restore the state information of the processor to the state information before executing the error instruction, and the state information of the processor includes the register renaming relationship.

[0014] According to a vector instruction processing method provided by the present invention, after allocating an empty entry in the instruction sequencing queue for the vector instruction as the entry corresponding to the vector instruction in the instruction sequencing queue, the method further includes: determining the identification information of the entry corresponding to the vector instruction in the instruction sequencing queue as the sequence identification information of the vector instruction; before writing back the vector instruction to the entry corresponding to the vector instruction in the instruction sequencing queue, the method further includes: determining the entry corresponding to the vector instruction in the instruction sequencing queue based on the sequence identification information of the vector instruction.

[0015] According to a vector instruction processing method provided by the present invention, after allocating an empty entry for the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue as the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue, the method further includes: determining the identification information of the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue as the sequence identification of the internal operation of each vector instruction corresponding to each data element; before writing back the internal operation of each vector instruction corresponding to each data element to the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue, the method further includes: determining the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue based on the sequence identification of the internal operation of each vector instruction corresponding to each data element.

[0016] The present invention also provides a vector instruction processing apparatus, including the following modules.

[0017] A first ordering module, configured to allocate an empty entry for the vector instruction in the instruction ordering queue as the entry corresponding to the vector instruction in the instruction ordering queue when the vector instruction enters the vector instruction issue queue.

[0018] A second ordering module, configured to split the vector instruction into internal operations of the vector instruction corresponding to each data element, perform register renaming on the internal operations of the vector instruction corresponding to each data element, and allocate an empty entry for the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue as the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue.

[0019] An operation execution module, configured to execute the internal operations of the vector instruction corresponding to each data element out of order.

[0020] An instruction write-back module, configured to write back the internal operation of each vector instruction corresponding to each data element to the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue, and then write back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue.

[0021] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the vector instruction processing method as described in any one of the above when executing the computer program.

[0022] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the vector instruction processing method as described in any one of the above is implemented.

[0023] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the vector instruction processing method as described in any one of the above is implemented.

[0024] The vector instruction processing method, device, electronic device and storage medium provided by the present invention, when a vector instruction enters the vector instruction issue queue, allocate an empty entry in the instruction ordering queue for the vector instruction as the entry corresponding to the vector instruction in the instruction ordering queue, and then split the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, perform register renaming on the vector instruction internal operations corresponding to each data element, allocate an empty entry in the vector instruction ordering queue for each vector instruction internal operation corresponding to each data element as the entry corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction ordering queue. After out-of-order execution of the vector instruction internal operations corresponding to each data element, write back the vector instruction internal operations corresponding to each data element to the entries corresponding to the vector instruction internal operations corresponding to each data element in the vector instruction ordering queue, and then write back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue. Based on the above multi-level ordering queues, it can achieve out-of-order execution and in-order submission of the vector instruction internal operations corresponding to each data element in the vector instruction, can more reasonably utilize the computing resources of the processor, thereby reducing the computing resources of the processor occupied by vector instruction processing, can improve the performance stability of the processor, can reduce the operating cost and manufacturing cost of the processor, can improve the execution efficiency of the vector instruction internal operations corresponding to each data element in the vector instruction, and further can improve the processing efficiency and throughput of the vector instruction, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0026] Figure 1 It is a field distribution diagram of a vector type in the related art.

[0027] Figure 2 It is one of the flow diagrams of the vector instruction processing method provided by the present invention.

[0028] Figure 3It is the second schematic flowchart of the vector instruction processing method provided by the present invention.

[0029] Figure 4 It is the schematic structural diagram of the vector instruction processing device provided by the present invention.

[0030] Figure 5 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0031] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0032] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "mounted", "connected" and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0033] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object may be one or more. In addition, in the description of the present application, " / and / " means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0034] It should be noted that RISC-V (Reduced Instruction Set Computing - Five, the fifth generation of reduced instruction set) is an open-source reduced instruction set that is currently widely used. The vector extension instruction set specification riscv-v-spec (RISC-V Vector Extension, RVV) of the RISC-V instruction set is an important extension of the RISC-V architecture, aiming to improve the parallel processing performance of the processor. RVV supports vector operations with multiple bit widths and allows dynamic adjustment of the vector length to flexibly adapt to different application requirements.

[0035] The RISC-V vector instruction set RVV includes vector control and status registers (Control and Status Register, CSR) and vector registers. The vector CSR is used to store configuration information, including information such as vector length and operation bit width, and the vector registers are used to store vectorized operation data.

[0036] RVV stipulates that on the basis of the basic spec (Specification), 32 vector registers (denoted as v0 - v31, with a bit width of VLEN each) and 7 non-privileged CSRs (vstart, vxsat, vxrm, vcsr, vl, vtype, and vlenb, with a bit width of XLEN each) are extended. Among them, VLEN refers to the length of a vector register (Vector Length, VLEN); XLEN refers to the length of a vector control and status register.

[0037] Among the configuration information stored in the vector CSR, vtype (vector type) is used to indicate the default type of the content of the vector register. The vector type also determines the arrangement of elements in each vector register and how to group multiple vector registers. Figure 1 is the field distribution diagram of the vector type in the related technology. The fields included in the vector type are as Figure 1 shown.

[0038] It should be noted that Figure 1 the vlmul[2:0] in [here] represents vector register grouping. Multiple vector registers can be grouped together so that a single vector instruction can operate on multiple vector registers.

[0039] LMUL refers to the number of logical channels in the vector register (Lane Multiplication Factor). vlmul represents a signed number. In addition to being an integer, the value of LMUL can also be a fraction or a decimal, reducing the number of bits used in a single vector register. LMUL = 2 vlmul[2:0]。The fractional grouping is mainly used for mixed-width instruction operations. Therefore, the values of LMUL can be 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8. The correspondence between LMUL and vlmul is shown in Table 1:

[0040] Table 1 Correspondence Table of LMUL and vlmul

[0041]

[0042] vsew[2:0] represents the width of each element in the vector register. The vector register is regarded as being divided into VLEN / SEW standard-width elements. SEW refers to the standard width of a single data element (Select Element Width, SEW).

[0043] The correspondence between SEW and vsew is shown in Table 2:

[0044] Table 2 Correspondence Table of SEW and vsew

[0045]

[0046] Figure 1 In, reserved indicates that this field is reserved, that is, it is not used in the current version of the instruction set. Reserved fields are usually for future extensions or compatibility considerations; vma (vector mask agnostic) indicates vector mask agnostic; vta (vector tail agnostic) indicates vector tail agnostic; vill being set indicates that an illegal vtype value is set.

[0047] Vector instructions usually include multiple data elements. The data elements in vector instructions are the basic units processed by vector instructions. When the processor processes vector instructions, the processor can perform arithmetic operations, logical operations, or other operations on each data element in the vector instruction simultaneously.

[0048] The number of data elements in a vector instruction is jointly determined by the vector control and status register and the type of vector instruction. The number of data elements in a vector instruction = VLEN × LMUL / SEW. For example, when the length of the vector source register corresponding to a vector instruction is 256bit, the number of logical channels of the above vector source register is 8, and the standard width of the data elements in the above vector source register is 16bit, the number of data elements in the above vector instruction is 128.

[0049] When the processor processes vector instructions, it needs to complete a large number of computing tasks in a short period of time. The high parallelism and complexity of vector instruction processing result in more processor computing resources being required for vector instruction processing, which has a significant impact on the processor's main frequency, chip area, power consumption and other performance.

[0050] For example, in order to support the parallel calculation of multiple data elements in vector instructions, multiple execution units in the processor are required to perform parallel calculations. However, when multiple execution units perform parallel calculations, the dynamic power consumption of the processor increases sharply.

[0051] For another example, after a vector instruction is executed, if the results of multiple data elements in the vector instruction need to be written back to different vector registers, a multi-port and wide-bit-width register stack is required to support parallel writing. A multi-port and wide-bit-width register stack needs to be configured with more transistors, thereby occupying more chip area, increasing static power consumption and chip manufacturing costs. In addition, writing the results of multiple data elements back to different vector registers will also cause a surge in the dynamic power consumption of the processor.

[0052] For example, the bit width of vector instructions is relatively large, which means that the data path in the processor that supports vector instructions requires wider connections (such as ALU or bus, etc.), which takes up more chip area. In addition, the wiring design of wide data paths is difficult, which can easily lead to the extension of data transmission paths and increased wiring complexity, which in turn leads to data transmission delays. The processor's main frequency is limited by the data transmission delay, and in actual operation, the processor may need to reduce the frequency to ensure operational stability.

[0053] Therefore, how to reduce the processor computing resources required for vector instruction processing, thereby improving the performance stability of the processor and reducing the operating cost and manufacturing cost of the processor, is a technical problem that needs to be urgently solved in this field.

[0054] In this regard, the present invention provides a vector instruction processing method. The vector instruction processing method provided by the present invention splits a vector instruction into multiple vector instruction internal operations inside a processor, and through a multi-level sequence queue, it can be realized that the multiple vector instruction internal operations split from the vector instruction can be executed in disorder and efficiently, and the vector instruction and the scalar instruction can also be executed in disorder and efficiently, thereby realizing the disorder execution and sequential submission of the vector instruction and the scalar instruction, and achieving the goals of high performance, high main frequency, low power consumption and low cost of the processor.

[0055] Combine the following Figures 2 - 3 The vector instruction processing method provided by the present invention is described.

[0056] Figure 2 FIG. 1 is one of the flow charts of the vector instruction processing method provided by the present invention.Figure 2 As shown in the figure, the method includes the following: Step 201, when a vector instruction enters the vector instruction issue queue, allocate an empty entry in the instruction ordering queue for the vector instruction as the entry corresponding to the vector instruction in the instruction ordering queue.

[0057] It should be noted that the execution subject of the embodiments of the present invention is a vector instruction processing device, and the vector instruction processing device can be a processor. When the above-mentioned processor executes the vector instruction processing method provided by the present invention, it can process vector instructions.

[0058] It should be noted that before the vector instruction in the embodiments of the present invention enters the vector instruction issue queue, register renaming has been completed in the main pipeline.

[0059] It should be noted that the instruction ordering queue in the embodiments of the present invention is shared by vector instructions and scalar instructions. The above-mentioned instruction ordering queue can ensure that instructions can be submitted in the original order under out-of-order execution.

[0060] Figure 3 This is the second flow chart of the vector instruction processing method provided by the present invention. As Figure 3 shown, after the vector instruction completes register renaming in the main pipeline, it enters the vector instruction issue queue.

[0061] When the vector instruction enters the vector instruction issue queue, an empty entry can be allocated in the instruction ordering queue for the above-mentioned vector instruction as the entry corresponding to the above-mentioned vector instruction in the instruction ordering queue, and the identification information of the entry corresponding to the above-mentioned vector instruction in the instruction ordering queue is determined as the sequence identification information of the above-mentioned vector instruction.

[0062] It should be noted that in the embodiments of the present invention, an empty entry can be allocated in the instruction ordering queue for the above-mentioned vector instruction in the order of the empty entries in the instruction ordering queue as the entry corresponding to the above-mentioned vector instruction in the instruction ordering queue.

[0063] Step 202, split the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, perform register renaming on the vector instruction internal operations corresponding to each data element, and allocate an empty entry in the vector instruction ordering queue for each vector instruction internal operation corresponding to each data element as the entry corresponding to each vector instruction internal operation corresponding to each data element in the vector instruction ordering queue.

[0064] Specifically, in the embodiments of the present invention, the above-mentioned vector instruction can be split into vector instruction internal operations corresponding to each data element in the above-mentioned vector instruction based on each data element in the above-mentioned vector instruction.

[0065] As an optional embodiment, splitting a vector instruction into internal operations of the vector instruction corresponding to each data element in the vector instruction includes: splitting the vector instruction into internal operations of the vector instruction corresponding to each data element based on the configuration information of the vector control and status register of the vector instruction and the vector instruction opcode of the vector instruction.

[0066] It should be noted that the vector instruction opcode of the vector instruction is used to identify the vector operation type of the vector instruction, such as load (i.e., fetching data) and store (i.e., storing data), as well as the data elements in the vector instruction.

[0067] It should be noted that the RISC-V vector extension introduces a series of vector control and status registers, which are used to configure the execution environment of vector instructions, manage the status information of vector processing, and provide fine-grained control over vector operations. The above vector control and status registers usually include a vector length register (VLEN), a vector type register (VTYPE), a vector start register (VSTART), etc. The above vector control and status registers together determine the behavior and performance of vector instructions. Therefore, the configuration information of the vector control and status register of the vector instruction can be used to indicate the vector length and type, the starting position of the vector operation, the execution environment of the vector operation, and exception and interrupt handling.

[0068] Based on the configuration information of the vector control and status register of the vector instruction and the vector instruction opcode of the vector instruction, the above vector instruction can be split into internal operations of the vector instruction corresponding to each data element in the above vector instruction.

[0069] After splitting the above vector instruction into internal operations of the vector instruction corresponding to each data element in the above vector instruction, register renaming can be performed on the internal operations of the vector instruction corresponding to each data element in the above vector instruction, find corresponding physical registers for the source register and destination register of the internal operations of the vector instruction corresponding to each data element in the above vector instruction, and map the destination register of the internal operations of the vector instruction corresponding to each data element in the above vector instruction to an idle physical register.

[0070] After performing register renaming on the internal operations of the vector instruction corresponding to each data element in the above vector instruction, an empty entry can be allocated respectively in the vector instruction ordering queue for the internal operations of the vector instruction corresponding to each data element in the above vector instruction, as the entry corresponding to the internal operations of the vector instruction corresponding to each data element in the above vector instruction in the vector instruction ordering queue, and determine the identification information of the entry corresponding to the internal operations of the vector instruction corresponding to each data element in the above vector instruction in the vector instruction ordering queue as the order identification information of the internal operations of the vector instruction corresponding to each data element in the above vector instruction.

[0071] It should be noted that the vector instruction ordering queue in the embodiments of the present invention is dedicated to vector instructions. The above vector instruction ordering queue can ensure that vector instructions are submitted in the original order under the condition of out-of-order execution.

[0072] It should be noted that in the embodiments of the present invention, an empty entry in the vector instruction ordering queue can be allocated to the internal operation of each vector instruction corresponding to each data element in the above vector instruction according to the arrangement order of the empty entries in the vector instruction ordering queue, as the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the above vector instruction in the vector instruction ordering queue.

[0073] Step 203: Execute the internal operations of the vector instructions corresponding to each data element out of order.

[0074] Specifically, after determining the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue, the internal operations of the vector instructions corresponding to each data element in the above vector instruction can be executed out of order to obtain the result data of the internal operation of each vector instruction corresponding to each data element in the above vector instruction.

[0075] As an optional embodiment, executing the internal operations of the vector instructions corresponding to each data element out of order includes: adding the internal operation of each vector instruction corresponding to each data element to the internal operation issue queue of the vector instruction. When the internal operation issue queue of the vector instruction determines that the source operands of the internal operations of the vector instructions corresponding to each data element are all ready, it issues the internal operations of the vector instructions corresponding to each data element by using an out-of-order issue mechanism, so that the internal operations of the vector instructions corresponding to each data element can read the vector register file to obtain the values of the source operands of the internal operations of the vector instructions corresponding to each data element.

[0076] Based on the operation types of the internal operations of the vector instructions corresponding to each data element, the internal operations of the vector instructions corresponding to each data element are respectively allocated to different execution units, so that different execution units execute the internal operations of the vector instructions corresponding to each data element based on the values of the source operands of the internal operations of the vector instructions corresponding to each data element.

[0077] Step 204: After writing back the internal operation of each vector instruction corresponding to each data element to the entry corresponding to the internal operation of each vector instruction corresponding to each data element in the vector instruction ordering queue, write back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue.

[0078] Specifically, after performing the internal operations of the vector instructions corresponding to each data element, based on the sequence identification information of the internal operations of the vector instructions corresponding to each data element in the above vector instructions, the internal operations of the vector instructions corresponding to each data element in the above vector instructions can be written back to the corresponding items of the internal operations of the vector instructions corresponding to each data element in the vector instruction ordering queue.

[0079] In the case where it is determined that the internal operations of the vector instructions corresponding to each data element have been written back to the corresponding items of the internal operations of the vector instructions corresponding to each data element in the vector instruction ordering queue, based on the sequence identification information of the above vector instructions, the vector instructions can be written back to the corresponding items of the vector instructions in the instruction ordering queue.

[0080] As an optional embodiment, after writing the vector instructions back to the corresponding items of the vector instructions in the instruction ordering queue, the method further includes: submitting the vector instructions in order through the instruction ordering queue.

[0081] Notify the instruction ordering queue to release the corresponding items of the vector instructions in the instruction ordering queue.

[0082] As an optional embodiment, after writing the vector instructions back to the corresponding items of the vector instructions in the instruction ordering queue, the method further includes: notifying the vector instruction ordering queue to release the corresponding items of the internal operations of the vector instructions corresponding to each data element in the vector instruction ordering queue.

[0083] It can be understood that any internal operation of the vector instructions occupies the items in the vector instruction ordering queue from after register renaming to before the vector instructions are submitted. After the vector instructions are register-renamed in the main pipeline, they occupy the items in the instruction ordering queue until the above vector instructions are submitted.

[0084] In the embodiment of the present invention, when a vector instruction enters the vector instruction issue queue, an empty entry is allocated in the instruction ordering queue for the vector instruction as the entry corresponding to the vector instruction in the instruction ordering queue. Then, the vector instruction is split into internal operations of the vector instruction corresponding to each data element in the vector instruction. Register renaming is performed on the internal operations of the vector instruction corresponding to each data element. An empty entry is allocated in the vector instruction ordering queue for each internal operation of the vector instruction corresponding to each data element as the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue. After out-of-order execution of the internal operations of the vector instruction corresponding to each data element, after writing back the internal operations of the vector instruction corresponding to each data element to the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue, the vector instruction is written back to the entry corresponding to the vector instruction in the instruction ordering queue. Based on the above multi-level ordering queue, out-of-order execution and in-order submission of the internal operations of the vector instruction corresponding to each data element in the vector instruction can be achieved, the computing resources of the processor can be utilized more reasonably, thereby reducing the computing resources of the processor required for processing vector instructions, improving the performance stability of the processor, reducing the operating cost and manufacturing cost of the processor, improving the execution efficiency of the internal operations of the vector instruction corresponding to each data element in the vector instruction, and further improving the processing efficiency and throughput of the vector instruction, having broad application prospects.

[0085] As an optional embodiment, the method further includes: when a scalar instruction enters the scalar instruction issue queue, an empty entry is allocated in the instruction ordering queue for the scalar instruction as the entry corresponding to the scalar instruction in the instruction ordering queue.

[0086] It should be noted that the scalar instruction in the embodiment of the present invention has completed register renaming in the main pipeline before entering the scalar instruction issue queue.

[0087] When the scalar instruction enters the scalar instruction issue queue, an empty entry can be allocated in the instruction ordering queue for the above scalar instruction as the entry corresponding to the above scalar instruction in the instruction ordering queue, and the identification information of the entry corresponding to the above scalar instruction in the instruction ordering queue is determined as the sequence identification information of the above scalar instruction.

[0088] It should be noted that in the embodiment of the present invention, an empty entry can be allocated in the instruction ordering queue for the above scalar instruction as the entry corresponding to the above scalar instruction in the instruction ordering queue according to the arrangement order of the empty entries in the instruction ordering queue.

[0089] The scalar instruction is added to the scalar instruction issue queue. When the scalar instruction issue queue determines that the source operands of the scalar instruction are ready, the scalar instruction is issued for the scalar instruction to read the scalar register file to obtain the values of the source operands of the scalar instruction.

[0090] Allocate scalar instructions to execution units for the execution units to execute the scalar instructions based on the values of the source operands of the scalar instructions.

[0091] Write the scalar instruction back to the entry corresponding to the scalar instruction in the instruction ordering queue.

[0092] In the embodiments of the present invention, the instruction ordering queue maintains the order among all instructions, including the order between vector instructions, between vector instructions and scalar instructions, and between scalar instructions. The vector instruction ordering queue maintains the order among the internal operations of vector instructions corresponding to multiple data elements in the vector instruction, ensuring that all internal operations of vector instructions split from one vector instruction are written back, which can make more reasonable use of the computing resources of the processor. By uniformly maintaining the order of the two types of instructions through the instruction ordering queue, it can better avoid execution conflicts or data inconsistencies, maximize the parallelism of instruction execution on the premise of ensuring program correctness, and simplify the hardware design and improve the scalability and maintainability of the processor by separating the order management of scalar instructions and vector instructions.

[0093] As an optional embodiment, the method further includes: in the case of a prediction execution error occurring in the processor, obtaining the order relationship between the error instruction with a prediction execution error and other instructions based on the vector instruction ordering queue and / or the instruction ordering queue.

[0094] Based on the order relationship between the error instruction and other instructions, restore the state information of the processor to the state information before executing the error instruction. The state information of the processor includes the register renaming relationship, and the processor is the execution subject of the method.

[0095] It should be noted that in the related art, in order to improve performance, the processor usually adopts prediction mechanisms such as branch prediction, memory access address prediction, and value prediction when processing instructions. The above prediction mechanisms allow the processor to execute subsequent instructions in advance when the results are not yet determined, so as to make full use of the pipeline and hardware resources. Among them, branch prediction can predict the direction (jump or not jump) of branch instructions (such as conditional jumps); memory access address prediction can predict the memory address or data correlation accessed by memory access instructions (such as Load / Store). Value prediction can predict the result values of certain instructions (such as Load instructions).

[0096] A prediction execution error occurs in a processor when the prediction of the processor does not match the actual result during the processing of an instruction, resulting in the processor executing the wrong instruction or using the wrong data. Common prediction execution errors include branch prediction errors, memory access instruction-related prediction errors, and memory access instruction value prediction errors. Among them, a branch prediction error means that the processor predicts the wrong direction of a branch instruction, resulting in the execution of the wrong instruction path. A memory access instruction-related prediction error means that the processor predicts the address or data correlation of a memory access instruction (such as Load / Store) incorrectly, resulting in data conflicts or errors. A memory access instruction value prediction error means that the processor predicts the result value of a memory access instruction (such as a Load instruction) incorrectly, resulting in subsequent instructions using the wrong data.

[0097] In an embodiment of the present invention, when a prediction execution error occurs in a processor, based on the order relationship between the internal operations of vector instructions corresponding to data elements in the vector instructions maintained by the vector instruction ordering queue, and / or the order relationship between the instructions maintained by the instruction ordering queue, the order relationship between the wrong instruction with the prediction execution error and other instructions can be obtained. Furthermore, based on the above order relationship, the state information of the processor can be restored to the state information before the execution of the wrong instruction.

[0098] Optionally, the state information of the processor in an embodiment of the present invention may further include a program counter.

[0099] In an embodiment of the present invention, by restoring the processor state in the vector instruction ordering queue and / or the instruction ordering queue when a prediction execution error occurs in the processor, the accuracy of the processor state restoration can be ensured, the impact of the prediction execution error on the subsequent instruction execution can be better predicted, the pipeline stall duration can be reduced, thereby improving the overall execution efficiency of the processor, and the fault tolerance of the processor when dealing with prediction execution errors can be enhanced, ensuring the continuous and stable operation of the processor.

[0100] Figure 4 It is a schematic structural diagram of a vector instruction processing device provided by the present invention. The following combines Figure 4 to describe the vector instruction processing device provided by the present invention. The vector instruction processing device described below can be mutually corresponding and referred to the vector instruction processing method provided by the present invention described above. As Figure 4 shown, the device includes: a first ordering module 401, a second ordering module 402, an operation execution module 403, and an instruction write-back module 404.

[0101] The first ordering module 401 is configured to, when a vector instruction enters the vector instruction issue queue, allocate an empty entry in the instruction ordering queue as the entry corresponding to the vector instruction in the instruction ordering queue.

[0102] A second ordering module 402, configured to split a vector instruction into internal operations of the vector instruction corresponding to each data element in the vector instruction, perform register renaming on the internal operations of the vector instruction corresponding to each data element, and allocate an empty entry in the vector instruction ordering queue for each internal operation of the vector instruction corresponding to each data element, as the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue.

[0103] An operation execution module 403, configured to execute out-of-order the internal operations of the vector instruction corresponding to each data element.

[0104] An instruction write-back module 404, configured to write back the internal operations of the vector instruction corresponding to each data element to the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue, and then write back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue.

[0105] Specifically, the first ordering module 401, the second ordering module 402, the operation execution module 403, and the instruction write-back module 404 are electrically connected.

[0106] In the vector instruction processing device according to the embodiment of the present invention, when a vector instruction enters the vector instruction issue queue, an empty entry is allocated in the instruction ordering queue for the vector instruction, as the entry corresponding to the vector instruction in the instruction ordering queue. Furthermore, the vector instruction is split into internal operations of the vector instruction corresponding to each data element in the vector instruction, register renaming is performed on the internal operations of the vector instruction corresponding to each data element, an empty entry is allocated in the vector instruction ordering queue for each internal operation of the vector instruction corresponding to each data element, as the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue. After out-of-order executing the internal operations of the vector instruction corresponding to each data element, the internal operations of the vector instruction corresponding to each data element are written back to the entry corresponding to the internal operation of the vector instruction corresponding to each data element in the vector instruction ordering queue, and then the vector instruction is written back to the entry corresponding to the vector instruction in the instruction ordering queue. It can achieve out-of-order execution and in-order submission of the internal operations of the vector instruction corresponding to each data element based on the above multi-level ordering queue, can more reasonably utilize the computing resources of the processor, thereby reducing the computing resources of the processor required for vector instruction processing, can improve the performance stability of the processor, can reduce the operating cost and manufacturing cost of the processor, can improve the execution efficiency of the internal operation of the vector instruction corresponding to each data element, and further can improve the processing efficiency and throughput of the vector instruction, and has broad application prospects.

[0107] Figure 5 Illustrates a schematic physical structure diagram of an electronic device, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a vector instruction processing method, which includes: when a vector instruction enters the vector instruction issue queue, allocating a blank entry in the instruction ordering queue as the entry corresponding to the vector instruction in the instruction ordering queue; splitting the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, performing register renaming on the vector instruction internal operations corresponding to each data element, and respectively allocating a blank entry in the vector instruction ordering queue as the entry corresponding to the vector instruction internal operation corresponding to each data element; executing out-of-order the vector instruction internal operations corresponding to each data element; after writing back the vector instruction internal operations corresponding to each data element to the entries corresponding to the vector instruction internal operations corresponding to each data element in the vector instruction ordering queue, writing back the vector instruction to the entry corresponding to the vector instruction in the instruction ordering queue.

[0108] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0109] Optionally, the electronic device in the embodiments of the present invention may be a processor.

[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the vector instruction processing method provided by each of the above methods. The method includes: when a vector instruction enters a vector instruction issue queue, allocate a blank entry in the instruction sequencing queue for the vector instruction as the entry corresponding to the vector instruction in the instruction sequencing queue; split the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, perform register renaming on the vector instruction internal operations corresponding to each data element, and allocate a blank entry in the vector instruction sequencing queue for each vector instruction internal operation corresponding to each data element as the entry corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue; execute out-of-order the vector instruction internal operations corresponding to each data element; after writing back each vector instruction internal operation corresponding to each data element to the entry corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue, write back the vector instruction to the entry corresponding to the vector instruction in the instruction sequencing queue.

[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the vector instruction processing method provided by each of the above methods. The method includes: when a vector instruction enters a vector instruction issue queue, allocate a blank entry in the instruction sequencing queue for the vector instruction as the entry corresponding to the vector instruction in the instruction sequencing queue; split the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, perform register renaming on the vector instruction internal operations corresponding to each data element, and allocate a blank entry in the vector instruction sequencing queue for each vector instruction internal operation corresponding to each data element as the entry corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue; execute out-of-order the vector instruction internal operations corresponding to each data element; after writing back each vector instruction internal operation corresponding to each data element to the entry corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue, write back the vector instruction to the entry corresponding to the vector instruction in the instruction sequencing queue.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A vector instruction processing method, characterized in that: include: In the case where a vector instruction enters the vector instruction issue queue, an empty item is allocated to the vector instruction in the instruction sequencing queue as the item corresponding to the vector instruction in the instruction sequencing queue; Splitting the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, performing register renaming on the vector instruction internal operations corresponding to each data element, and respectively allocating an empty item to each vector instruction internal operation corresponding to the data element in the vector instruction sequencing queue as an item corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue; Execute the vector instruction internal operations corresponding to each of the data elements out of order; After writing the vector instruction internal operation corresponding to each of the data elements back to the entry corresponding to the vector instruction internal operation corresponding to each of the data elements in the vector instruction sequencing queue, the vector instruction is written back to the entry corresponding to the vector instruction in the instruction sequencing queue.

2. The vector instruction processing method according to claim 1, characterized in that: The out-of-order execution of the vector instruction internal operations corresponding to the data elements includes: Adding the vector instruction internal operation corresponding to each of the data elements into a vector instruction internal operation emission queue, and the vector instruction internal operation emission queue, when determining that the source operands of the vector instruction internal operation corresponding to each of the data elements are ready, uses a disorderly emission mechanism to emit the vector instruction internal operation corresponding to each of the data elements, so that the vector instruction internal operation corresponding to each of the data elements reads a vector register stack and obtains the value of the source operand of the vector instruction internal operation corresponding to each of the data elements; Based on the operation type of the vector instruction internal operation corresponding to each of the data elements, the vector instruction internal operation corresponding to each of the data elements is respectively allocated to different execution units, so that the different execution units can execute the vector instruction internal operation corresponding to each of the data elements based on the value of the source operand of the vector instruction internal operation corresponding to each of the data elements.

3. The vector instruction processing method according to claim 1, characterized in that: The step of splitting the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction includes: Based on the configuration information of the vector control and status register of the vector instruction and the vector instruction operation code of the vector instruction, the vector instruction is split into vector instruction internal operations corresponding to each of the data elements.

4. The vector instruction processing method according to claim 1, characterized in that: After writing the vector instruction back to the item corresponding to the vector instruction in the instruction sequencing queue, the method further includes: Submitting the vector instructions in order through the instruction sequencing queue; The instruction sequencing queue is notified to release the item corresponding to the vector instruction in the instruction sequencing queue, and the vector instruction sequencing queue is notified to release the item corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue.

5. The vector instruction processing method according to any one of claims 1 to 4, characterized in that: The method further comprises: In the case where a scalar instruction enters the scalar instruction issue queue, allocating an empty entry for the scalar instruction in the instruction sequencing queue as the entry corresponding to the scalar instruction in the instruction sequencing queue; The scalar instruction is added to a scalar instruction issuance queue, and the scalar instruction issuance queue, when determining that a source operand of the scalar instruction is ready, issues the scalar instruction, so that the scalar instruction reads a scalar register file to obtain a value of the source operand of the scalar instruction; Dispatching the scalar instruction to an execution unit so that the execution unit executes the scalar instruction based on a value of a source operand of the scalar instruction; The scalar instruction is written back to the entry corresponding to the scalar instruction in the instruction sequencing queue.

6. The vector instruction processing method according to claim 5, characterized in that: The method further comprises: When a prediction execution error occurs in the processor, obtaining a sequence relationship between an erroneous instruction having the prediction execution error and other instructions based on the vector instruction sequencing queue and / or the instruction sequencing queue; Based on the sequential relationship between the erroneous instruction and other instructions, the state information of the processor is restored to the state information before the erroneous instruction is executed, the state information of the processor includes a register renaming relationship, and the processor is the execution subject of the method.

7. The vector instruction processing method according to claim 1, characterized in that: After allocating an empty entry for the vector instruction in the instruction sequencing queue as the entry corresponding to the vector instruction in the instruction sequencing queue, the method further comprises: Determine the identification information of the item corresponding to the vector instruction in the instruction sequencing queue as the sequence identification information of the vector instruction; Before writing the vector instruction back to the item corresponding to the vector instruction in the instruction sequencing queue, the method further includes: Based on the sequence identification information of the vector instruction, an item corresponding to the vector instruction is determined in the instruction sequencing queue.

8. The vector instruction processing method according to claim 1, characterized in that: After allocating an empty item in the vector instruction sequencing queue for each vector instruction internal operation corresponding to the data element as an item corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue, the method further comprises: Determine the identification information of the item corresponding to the vector instruction internal operation corresponding to each of the data elements in the vector instruction sequencing queue as the sequence identification of the vector instruction internal operation corresponding to each of the data elements; Before writing the vector instruction internal operation corresponding to each of the data elements back to the entry corresponding to the vector instruction internal operation corresponding to each of the data elements in the vector instruction sequencing queue, the method further includes: Based on the sequence identifier of the vector instruction internal operation corresponding to each of the data elements, an item corresponding to the vector instruction internal operation corresponding to each of the data elements is determined in the vector instruction sequencing queue.

9. A vector instruction processing device, characterized in that: include: A first sequencing module is used for allocating an empty item for the vector instruction in the instruction sequencing queue as the item corresponding to the vector instruction in the instruction sequencing queue when the vector instruction enters the vector instruction issuance queue; A second sequencing module is used to split the vector instruction into vector instruction internal operations corresponding to each data element in the vector instruction, perform register renaming on the vector instruction internal operations corresponding to each data element, and allocate an empty item to each vector instruction internal operation corresponding to the data element in the vector instruction sequencing queue as an item corresponding to the vector instruction internal operation corresponding to each data element in the vector instruction sequencing queue; An operation execution module, used for executing the internal operations of the vector instructions corresponding to each of the data elements out of order; The instruction write-back module is used to write the vector instruction internal operation corresponding to each of the data elements back to the item corresponding to the vector instruction internal operation corresponding to each of the data elements in the vector instruction sequencing queue, and then write the vector instruction back to the item corresponding to the vector instruction in the instruction sequencing queue.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the vector instruction processing method according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vector instruction processing method according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the vector instruction processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Distribution execution method and device for multiple instructions of out-of-order execution queue in out-of-order processor

    CN113805944A

  • Instruction processing method and device for out-of-order multi-transmission processor

    CN118885218A