Floating-point unordered add operation method, apparatus, and processor

CN122044660BActive Publication Date: 2026-09-04GUANGDONG LEAPFIVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610487676.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-14
Publication Date
2026-09-04
Estimated Expiration
2046-04-14

AI Technical Summary

Technical Problem

然而,采用vfredusum.vs指令进行无序归约求和所需的硬件结构复杂、时序路径长且扩展性差,难以兼顾高吞吐与低延迟的性能需求

Benefits of technology

[0015]The floating-point unordered addition method provided in this application obtains the vector source operands and writes them into a logical vector register composed of at least one physical vector register. This allows for dynamic adjustment of the number of physical vector registers used based on the vector length multiplier, achieving flexible storage and organization of the vector source operands. Furthermore, it performs vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain the vector reduction result. This reduces computational depth and increases parallelism. Simultaneously, it adheres to the semantic specifications of the RISC-V vector extension instruction set, accumulating the vector reduction result with the least significant bit in the scalar register to obtain the computation result of the vector source operands. This ensures the consistency of instruction semantics, avoids interference with the parallel reduction tree structure, reduces hardware complexity, and effectively solves the technical problem of the performance bottleneck of vector reduction operations in the RISC-V vector extension instruction set. Therefore, it provides a superior hardware implementation solution for high-performance computing and AI acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122044660B_ABST
    Figure CN122044660B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a floating point unordered addition operation method, device and processor. The method comprises the following steps: obtaining a vector source operand, and writing the vector source operand into a logical vector register; the vector source operand comprises a plurality of floating point elements, and the logical vector register comprises at least one physical vector register; performing vector-level unordered summation reduction on the plurality of floating point elements in the logical vector register to obtain a vector reduction result of the vector source operand; and accumulating the vector reduction result and the element with the lowest bit in a scalar register to obtain an operation result of the vector source operand. The technical problem of the performance bottleneck of the vector reduction operation in the RISC-V vector extension instruction set is effectively solved, and a more optimal hardware implementation scheme is provided for high-performance computing and AI acceleration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a floating-point unordered addition operation method, apparatus and processor. Background Technology

[0002] In the field of parallel computing, vector reduction is a fundamental operation in high-performance computing, artificial intelligence, and scientific computing, and is widely used in matrix operations, statistical analysis, and neural network training. Vector reduction achieves efficient computation of operations such as addition, maximization, and minimization by merging multiple operands hierarchically according to a tree structure.

[0003] In the RISC-V Vector Extension instruction set (RVV), the `vfredusum.vs` instruction, as an unordered reduction summation instruction, allows the implementation of a reduction tree structure that does not guarantee the order of operations, thereby reducing computation depth and improving parallelism. However, the hardware structure required for unordered reduction summation using the `vfredusum.vs` instruction is complex, the timing path is long, and the scalability is poor, making it difficult to simultaneously meet the performance requirements of high throughput and low latency. Summary of the Invention

[0004] Based on this, this application provides a floating-point unordered addition operation method, apparatus and processor, which can effectively solve the technical problem of the performance bottleneck of vector reduction operation in the RISC-V vector extension instruction set.

[0005] Firstly, this application provides a floating-point unordered addition operation method, including: Obtain the vector source operand and write it into the logical vector register; the vector source operand includes multiple floating-point elements, and the logical vector register includes at least one physical vector register; Perform vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operands; The vector reduction result is accumulated with the least significant element in the scalar register to obtain the result of the vector source operand operation.

[0006] Furthermore, in the floating-point unordered addition method provided in this application, before writing the vector source operand into the logic vector register, the method further includes: Determine the vector length ratio of the control logic vector register based on the vector bit width of the vector source operand and the physical bit width of the physical vector register; Based on the vector length ratio, at least one physical vector register is selected to form a logical vector register.

[0007] Furthermore, in the floating-point unordered addition method provided in this application, writing the vector source operand into the logic vector register includes: If the vector length ratio is less than the preset length ratio threshold, the vector source operand is written to a physical vector register; If the vector length multiplier is greater than or equal to the length multiplier threshold, the vector source operand is written to multiple physical vector registers.

[0008] Furthermore, in the floating-point unordered addition method provided in this application, the vector source operands are written to multiple physical vector registers, including: Based on the vector length ratio, the vector source operands are split into multiple groups of floating-point numbers; each group of floating-point numbers includes multiple floating-point elements; Each set of floating-point numbers is input into a physical vector register; each set of floating-point numbers corresponds to one physical vector register.

[0009] Furthermore, in the floating-point unordered addition method provided in this application, multiple floating-point elements in the logic vector register are subjected to vector-level unordered summation and reduction to obtain the vector reduction result of the vector source operands, including: If the vector length ratio is less than the length ratio threshold, a multi-level parallel floating-point adder is used to perform vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result. If the vector length multiplier is greater than or equal to the length multiplier threshold, perform a floating-point unordered addition operation on the two floating-point elements at the same position between the first physical vector register and the second physical vector register in the logical vector register to obtain the first vector operand. The first vector operand is written into the first buffer register to perform unordered summation and reduction at the vector level, and the vector reduction result is obtained.

[0010] Furthermore, in the floating-point unordered addition method provided in this application, the first vector operand is written into the first buffer register to perform unordered summation and reduction at the vector level, obtaining a vector reduction result, including: Write the first vector operand to the first buffer register; If the vector length multiplier is equal to the length multiplier threshold, perform unordered summation and reduction on the first vector operand in the first buffer register to obtain the vector reduction result; If the vector length ratio is greater than the length ratio threshold, perform a floating-point unordered addition operation on the two floating-point elements at the same position between the first buffer register and the third physical vector register in the logical vector register to obtain the second vector operand. The first vector operand is written into the second buffer register to perform unordered summation and reduction at the vector level, resulting in a vector reduction result.

[0011] Furthermore, in the floating-point unordered addition method provided in this application, the first physical vector register and the second physical vector register are two adjacent physically connected registers in the logical vector registers; the first physical vector register or the second physical vector register is adjacent to the third physical vector register in the logical vector registers; or / and, The floating-point number of a floating-point element is any one of 8 bits, 16 bits, 32 bits, 64 bits, and 128 bits; or / and, The physical bit width is any one of 128bit, 256bit, or 512bit, and the physical bit widths of the first physical vector register, the second physical vector register, and the third physical vector register are equal.

[0012] Furthermore, in the floating-point unordered addition method provided in this application, before performing vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operands, the method further includes: The vector source operands are preprocessed to obtain the preprocessed vector source operands.

[0013] Secondly, this application also provides a floating-point unordered addition arithmetic device, comprising: The write unit is used to obtain the vector source operand and write the vector source operand into the logical vector register; the vector source operand includes multiple floating-point elements, and the logical vector register includes at least one physical vector register. The processing unit is used to perform vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operands; The accumulation unit is used to accumulate the vector reduction result with the least significant element in the scalar register to obtain the result of the vector source operand operation.

[0014] Thirdly, this application also provides a processor that uses RISC-V vector unordered reduction floating-point summation instructions to execute the floating-point unordered addition method provided in the first aspect.

[0015] The floating-point unordered addition method provided in this application obtains the vector source operands and writes them into a logical vector register composed of at least one physical vector register. This allows for dynamic adjustment of the number of physical vector registers used based on the vector length multiplier, achieving flexible storage and organization of the vector source operands. Furthermore, it performs vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain the vector reduction result. This reduces computational depth and increases parallelism. Simultaneously, it adheres to the semantic specifications of the RISC-V vector extension instruction set, accumulating the vector reduction result with the least significant bit in the scalar register to obtain the computation result of the vector source operands. This ensures the consistency of instruction semantics, avoids interference with the parallel reduction tree structure, reduces hardware complexity, and effectively solves the technical problem of the performance bottleneck of vector reduction operations in the RISC-V vector extension instruction set. Therefore, it provides a superior hardware implementation solution for high-performance computing and AI acceleration. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic flowchart of a floating-point unordered addition operation method provided in an embodiment of this application; Figure 2 This is a first scenario architecture diagram for floating-point unordered addition operations provided in an embodiment of this application; Figure 3 This is a second scenario architecture diagram for floating-point unordered addition operations provided in an embodiment of this application; Figure 4 This is a third scenario architecture diagram for floating-point unordered addition operations provided in an embodiment of this application; Figure 5 A schematic block diagram of a floating-point unordered addition arithmetic device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0020] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0021] It should also be further understood that the term “and / or” as used in this application specification and the appended claims is to mean any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0022] Furthermore, in this application, unless otherwise explicitly specified or limited in the embodiments, the terms "installation," "connection," "joining," and "fixing" appearing in the embodiments should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral part; it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication between two components, or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific implementation.

[0023] In related technologies, floating-point parallel addition tree structures are used to implement vector reduction operations. The core idea is to add N inputs pairwise to form a balanced binary tree. The parallel depth is log2N.

[0024] The floating-point parallel addition tree structure contains multiple levels of parallel floating-point addition units, connected by pipelined registers, for parallel reduction and summation of input floating-point operands. Each floating-point addition node contains submodules for exponent comparison, mantissa alignment, mantissa addition, normalization, and rounding. To comply with the IEEE 754 standard, it is typically designed with 3-5 pipelined stages. Furthermore, the floating-point parallel addition tree structure can be implemented using either an ordered addition tree (first sorted by exponent) or an unordered addition tree (direct pairwise addition) depending on requirements.

[0025] However, the hardware structure required for out-of-order reduction summation using the vfredusum.vs instruction is complex, the timing path is long, and the scalability is poor, making it difficult to balance the performance requirements of high throughput and low latency.

[0026] To address this, this application provides a floating-point unordered addition method. By acquiring the vector source operand and writing it into a logical vector register composed of at least one physical vector register, the number of physical vector registers used can be dynamically adjusted according to the vector length ratio, achieving flexible storage and organization of the vector source operand. Furthermore, vector-level unordered summation and reduction are performed on multiple floating-point elements in the logical vector register to obtain the vector reduction result. This reduces computational depth and increases parallelism. Simultaneously, it adheres to the semantic specifications of the RISC-V vector extension instruction set, accumulating the vector reduction result with the least significant bit in the scalar register to obtain the computation result of the vector source operand. This ensures the consistency of instruction semantics, avoids interference with the parallel reduction tree structure, reduces hardware complexity, and effectively solves the technical problem of the performance bottleneck of vector reduction operations in the RISC-V vector extension instruction set. Therefore, it provides a better hardware implementation solution for high-performance computing and AI acceleration.

[0027] The instruction prefetching method provided in this application will be described in detail below.

[0028] like Figure 1 As shown, the method includes the following steps S110 to S130.

[0029] S110. Obtain the vector source operand and write it into the logical vector register; the vector source operand includes multiple floating-point elements, and the logical vector register includes at least one physical vector register.

[0030] In this embodiment, the vector source operand can be understood as all floating-point elements in the source vector register group vs2 specified by the RISC-V vector instruction vfredusum.vs, which are valid within the currently active vector length range.

[0031] Floating-point elements can be floating-point numbers in 8-bit, 16-bit, 32-bit, 64-bit, or 128-bit floating-point formats conforming to the IEEE 754 standard.

[0032] Logical vector registers are abstract registers from a software perspective. They have variable bit widths and are composed of one or more consecutively numbered physical vector registers, each with a fixed bit width (e.g., 128 bits, 256 bits, or 512 bits). The physical mapping size of logical vector registers is variable, determined by the LMUL (Length MULtiplier) in the vector type register vtype. LMUL=1 corresponds to a single physical vector register, while LMUL≥2 corresponds to LMUL consecutive physical vector registers.

[0033] Specifically, during the process of writing vector source operands, this application can mask invalid elements based on the current vl value and vector mask vmask, and write valid floating-point elements into the data field of the corresponding physical vector register in logical order, thereby providing a complete, ordered, and addressable data foundation for subsequent vector-level unordered reduction.

[0034] As an example, this application can trigger the reading of the vs2 register group based on the vector instruction decoding result, load floating-point elements in parallel into the write buffer queue via the vector register file read port, and then distribute the elements to the corresponding physical vector registers according to the LMUL configuration information.

[0035] As another example, this application can also dynamically generate address decoding signals based on the vtype value configured by the vstvli instruction, and control the vector register file to write floating-point elements of a specified range in vs2 to the target physical vector register by byte offset.

[0036] In addition, this application can also perform position alignment and order reordering on floating-point elements read across registers through the vector rearrangement unit, and then write them into the set of physical vector registers mapped by the logical vector register.

[0037] For example, when LMUL=1, SEW=32, and VLEN=512, the vector source operand vs2 contains 16 FP32 elements, all of which are written into a single 512-bit physical vector register v2, with element indices 0 to 15 corresponding to register bits [31:0] to [511:480] respectively; when LMUL=2, SEW=32, and VLEN=512, vs2 contains 32 FP32 elements, which are split into two groups of 16 elements each, and written into two consecutive 512-bit physical vector registers v2 and v3 respectively, where v2 carries elements 0 to 15 and v3 carries elements 16 to 31.

[0038] In some embodiments, before writing the vector source operand into the logic vector register, the method further includes: determining the vector length multiplier of the control logic vector register based on the vector bit width of the vector source operand and the physical bit width of the physical vector register; and selecting at least one physical vector register to form the logic vector register based on the vector length multiplier.

[0039] The vector length multiplier (LMUL) is a key configuration parameter in the RISC-V Vector Extension (RVV) architecture used to define the mapping relationship between logical vector registers and physical vector registers. LMUL can be understood as the number of physical vector registers occupied by the logical vector registers, and its value is determined by the ratio of the vector bit width to the physical vector register bit width.

[0040] In this embodiment, the vector bit width is equal to the product of the effective vector length (VL) and the standard element width (SEW), and the physical bit width of the physical vector register is VLEN. Therefore, LMUL = (VL × SEW) / VLEN, where the ratio is an integer or a fraction and is encoded in the vtype control register in the RVV specification. The relationship between LMUL and the physical vector register is shown in Table 1. Table 1

[0041] Among them, LMUL can be determined during the instruction decoding stage and used as a pre-control signal for hardware resource configuration to drive subsequent register address generation, data path width configuration and reduction structure selection.

[0042] Physical vector registers are actual register units in the processor with a fixed bit width (e.g., 128-bit, 256-bit, or 512-bit), numbered v0 to v31.

[0043] When LMUL=1, the logical vector register consists of a single physical vector register, for example, vs2 is mapped to v2; when LMUL=2, the logical vector register consists of two consecutively numbered physical vector registers, for example, vs2 is mapped to v2 and v3; when LMUL=4, the logical vector register consists of four consecutively numbered physical vector registers, for example, vs2 is mapped to v2, v3, v4 and v5; when LMUL is a fraction (such as 1 / 2 or 1 / 4), the logical vector register consists of a portion of a single physical vector register after subword partitioning. In this case, it is still considered as selecting at least one physical vector register, but only a portion of its bit width is used.

[0044] For example, before the RISC-V processor executes the vfredusum.vs instruction, the LMUL is configured to 2 using the vsetvli t0, a0,e32,m2 instruction. At this time, the vector width is 1024 bits (VLEN=512 bits × LMUL=2). It can be determined that vs2 needs to be jointly carried by two 512-bit physical vector registers, v2 and v3. Subsequently, during the instruction execution phase, the hardware address decoding module automatically distributes the read request of vs2 to v2 and v3 and enables dual-path parallel data.

[0045] When the `vsetvli t0, a0, e16, m4` instruction configures LMUL to 4, it is recognized that four physical vector registers (v2 to v5) are needed to form a 2048-bit logical vector register, and then four parallel read paths and cross-register alignment buffer units are configured. When the `vsetvli t0, a0, e8, mf2` instruction configures LMUL to 1 / 2, it is recognized that a single 512-bit physical vector register v2 can be divided into 64 8-bit elements. Only the first VL elements are used in the operation, and the rest are masked or set to zero.

[0046] In this application, a configurable mapping relationship between logical vector registers and physical resources is established by explicitly determining the LMUL parameters and dynamically selecting the physical vector registers. This avoids hardware redundancy of configuring separate wiring for each LMUL and ensures compatibility and scalability for the full range of RVV LMUL modes.

[0047] In some embodiments, writing vector source operands into logical vector registers includes: if the vector length ratio is less than a preset length ratio threshold, writing the vector source operands into one physical vector register; if the vector length ratio is greater than or equal to the length ratio threshold, writing the vector source operands into multiple physical vector registers.

[0048] In this embodiment, the vector length multiplier is less than the length multiplier threshold, which can be understood as LMUL<2. As a result, a physical vector register can simultaneously hold at least one logical vector, and the vector source operands can be mapped as a whole and written into a single physical vector register without the need for cross-register layout.

[0049] When the vector length multiplier is greater than or equal to the length multiplier threshold, it can be understood as LMUL≥2, which means that the logical vector bit width is greater than or equal to the bit width of two physical vector registers. In this case, one physical vector register cannot hold all vector elements. Therefore, the vector source operands need to be split according to the logical order of the elements and written sequentially into multiple consecutively numbered physical vector registers.

[0050] For example, when LMUL=2, the vector source operand is written to two physical vector registers, v2 and v3; when LMUL=4, it is written to four physical vector registers, v2, v3, v4 and v5; each register carries the same number of floating-point elements, and the element numbers are kept consecutive in the logical vector.

[0051] Meanwhile, this application can generate a register address increment sequence based on the LMUL value, divide the vector source operand into units of 512 bits, and write them sequentially into the starting register and its subsequent consecutively numbered registers; or, this application can also linearly expand the logical vector address space into the address space of multiple physical vector registers through the address mapping unit, and synchronously drive the write enable signals of multiple registers during the writing phase.

[0052] For example, when SEW=32bit, VLEN=512bit, and LMUL=2, the effective vector bit width is 1024bit, corresponding to 32 FP32 elements. The vector source operand is split into two groups of 16 FP32 elements each, which are written to v2 (carrying elements 0–15) and v3 (carrying elements 16–31) respectively. During the writing process, the hardware automatically verifies the continuity of the register numbers of v2 and v3 and ensures that the two register write operations are completed in the same instruction cycle or meet strict timing constraints to maintain the data consistency of the logical vector.

[0053] In this application, when LMUL < 2, single-register compact writing is used to reduce memory access overhead; when LMUL ≥ 1, multi-register continuous writing is used to ensure the integrity of ultra-wide vector data. This not only adapts to the RVV specification's definition of LMUL, but also provides a clear and stable data organization premise for key operations such as cross-register parallel addition and hierarchical reduction in subsequent embodiments 4–6, thereby supporting the efficient and scalable hardware implementation of the vfredusum.vs instruction in high LMUL scenarios.

[0054] In some embodiments, writing vector source operands to multiple physical vector registers includes: splitting the vector source operands into multiple groups of floating-point numbers based on a vector length multiplier; each group of floating-point numbers includes multiple floating-point elements; inputting each group of floating-point numbers into a physical vector register; each group of floating-point numbers corresponds to a physical vector register.

[0055] In this embodiment, during the process of splitting the vector source operands into multiple groups of floating-point numbers, the original consecutively arranged vector source operands can be divided into equal-length segments based on the value of the vector length multiplier, using the total number of floating-point elements that the physical vector register can accommodate as the unit.

[0056] For example, when LMUL=2, SEW=32bit, and VLEN=512bit, the vector source operand contains a total of 32 FP32 elements. These are split into two groups of 16 elements each: the first group contains elements 0 to 15, and the second group contains elements 16 to 31. The first group is written to the physical vector register vs2, and the second group is written to the physical vector register vs3. The two groups still logically constitute the same vector source operand vs2, and the element number i is mapped to vs2[i] and vs3[i-16] in vs2 and vs3 respectively, ensuring that elements at the same logical position can be synchronously addressed and processed in subsequent calculations.

[0057] In this application, by logically splitting the vector source operands according to the vector length multiple, and strictly mapping each group of floating-point numbers to consecutively numbered physical vector registers, the structured organization of vector data in high LMUL scenarios can be achieved. At the same time, floating-point elements at the same index position in each physical vector register can be spatially aligned and synchronously read in time, avoiding the problem of relying on complex rearrangement buffers or dynamic address calculation units to restore the logical order. This reduces the hardware implementation complexity and timing pressure, and improves the scalability and execution efficiency of vector reduction instructions under different LMUL configurations.

[0058] S120. Perform vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operands.

[0059] Specifically, this application can perform vector-level unordered summation and reduction on multiple floating-point elements in a logical vector register based on LMUL to obtain the vector reduction result of the vector source operands. The vector-level unordered summation and reduction can be understood as performing a tree-like summation operation on all valid floating-point elements in the logical vector register through multi-stage parallel floating-point adders, without guaranteeing the order of floating-point addition. The result satisfies IEEE 754 floating-point semantics and is consistent with the precise and rounded results under any legal assortment order.

[0060] The vector level indicates that the reduction operation operates on the entire set of floating-point elements carried by the entire logical vector register, rather than being limited to the internal structure of a single physical vector register.

[0061] Unordered operation specifically refers to the fact that the input order, execution timing, and hierarchical allocation of each addition node during the operation process are not constrained by the program order of the original vector elements, allowing the hardware to be freely scheduled to optimize the critical path and timing convergence.

[0062] The vector reduction result can be understood as a single floating-point value output by the root node of the reduction tree, and its data format is consistent with the SEW (Standard Element Width) of the floating-point elements in the vector source operands.

[0063] As an example, this application can divide the N floating-point elements in the logic vector register into N / 2 pairs according to the multi-level parallel floating-point adder structure. Each pair of elements is input to an independent floating-point adder unit in the same cycle to perform addition, generating N / 2 intermediate sums. Then, the N / 2 intermediate sums are paired up and input to the next level adder unit, and so on, until a unique reduction result is generated.

[0064] As another example, this application can also use a balanced binary tree topology to organize the addition units, with the number of adders in each stage being half that of the previous stage. Pipeline registers are inserted between each stage to increase the operating frequency. The delay of the first stage is mainly determined by exponent alignment, while the delay of subsequent stages is dominated by mantissa addition and normalization.

[0065] In addition, this application can also dynamically disable some addition unit inputs based on the vector mask vmask, so that the elements that are masked and set to zero are equivalent to the addition identity element (i.e., 0) during the reduction process, thereby ensuring that the reduction result only reflects the sum of the valid elements.

[0066] For example, when vs2={1.0,2.0,3.0,4.0} (FP32, VL=4) and vmask=4'b1111, the reduction process is as follows: Level-0 executes 1.0+2.0=3.0, 3.0+4.0=7.0; Level-1 executes 3.0+7.0=10.0, resulting in a vector reduction result of 10.0; when vmask=4'b0111, vs2 is equivalent to {0.0,2.0,3.0,4.0}, and the reduction process is as follows: Level-0 executes 0.0+2.0=2.0, 3.0+4.0=7.0; Level-1 executes 2.0+7.0=9.0, resulting in a vector reduction result of 9.0.

[0067] In some embodiments, performing vector-level unordered summation and reduction on multiple floating-point elements in a logical vector register to obtain a vector reduction result for the vector source operand includes: if the vector length ratio is less than a length ratio threshold, using a multi-stage parallel floating-point adder to perform vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain a vector reduction result; if the vector length ratio is greater than or equal to the length ratio threshold, performing unordered floating-point addition on two floating-point elements at the same position between the first physical vector register and the second physical vector register in the logical vector register to obtain a first vector operand; and writing the first vector operand into a first buffer register for vector-level unordered summation and reduction to obtain a vector reduction result.

[0068] In this embodiment, the multi-stage parallel floating-point adder is a multi-stage pipelined floating-point addition unit group organized according to the reduction tree structure. Each stage pairs the input floating-point elements to perform IEEE 754-compatible unordered floating-point addition operations and outputs intermediate reduction results.

[0069] like Figure 2 As shown, the multi-level parallel floating-point adder is configured to operate on all floating-point elements carried by a single physical vector register. Its input source is the vector source operands already written in the physical vector register, and its output is the target vector reduction result. It does not involve data reading or element alignment operations across physical vector registers.

[0070] As an example, such as Figure 2 As shown, this application can perform pairwise addition of all floating-point elements in a single physical vector register by means of a multi-level parallel floating-point adder to generate a reduction result.

[0071] As another example, this application can also reduce a masked subset of valid floating-point elements using a multi-stage parallel floating-point adder, wherein invalid elements are set to zero or replaced with additive identity elements.

[0072] For example, when LMUL=1, SEW=32bit, and VLEN=512bit, the vector source operand contains 16 FP32 floating-point elements, all located in the v2 register; the multi-level parallel floating-point adder is expanded into 5 levels from Level-0 to Level-4. Level-0 receives 16 inputs and generates 8 partial sums, Level-1 generates 4 partial sums, Level-2 generates 2 partial sums, Level-3 generates 1 partial sum, and Level-4 accumulates the partial sum with the scalar register vs1[0] and outputs it to vd[0]. The entire process does not access other physical vector registers, which meets the judgment condition that the vector length multiple is less than the length multiple threshold.

[0073] Two consecutive physical vector registers can be understood as a pair of physical vector registers with adjacent numbers (e.g., v2 and v3), which logically carry different segments of the same vector source operand.

[0074] Two floating-point elements at the same position can be understood as floating-point elements with the same index number in their respective registers (e.g., v2[i] and v3[i], i∈[0,N)), with the same data width and satisfying the IEEE 754 floating-point format requirements.

[0075] Floating-point unordered addition can be understood as single-cycle or multi-cycle floating-point addition operations conforming to the IEEE 754 standard, supporting exponent alignment, mantissa addition, normalization, and dynamic rounding.

[0076] In this embodiment, the unordered floating-point addition operation is configured to be executed in vector-level parallelism, that is, one operation processes N pairs of elements with the same position simultaneously and outputs N floating-point results, which constitute the first vector operand. The unordered floating-point addition operation does not depend on the reduction tree structure, nor does it introduce order constraints. It belongs to the data-level parallel addition behavior and provides a compressed intermediate vector representation for subsequent reduction.

[0077] As an example, this application can synchronously perform addition on the floating-point elements at corresponding index positions in the first physical vector register and the second physical vector register using the means of vector-level floating-point addition units to generate N-dimensional first vector operands.

[0078] As another example, this application can use the vector-level floating-point addition unit to perform masking and filtering of the floating-point elements involved in the operation before execution, perform addition only on valid elements, and output the result of setting the position of invalid elements to zero.

[0079] For example, when LMUL=4, SEW=32bit, and VLEN=512bit, the vector source operands are distributed in four consecutive physical vector registers v2, v3, v4, and v5, each containing 16 FP32 elements. First, the vector-level floating-point addition unit is started, and the elements with indices 0 to 15 in v2 and v3 are added respectively to generate the first vector operand R1 containing 16 FP32 elements, where R1[i]=v2[i]+v3[i]. This operation completes all 16 parallel additions in a single cycle without the need to build a giant reduction tree covering v2 to v5, significantly reducing the first-cycle latency.

[0080] The first buffer register can be understood as a register used to temporarily store intermediate results of cross-register reduction. Its bit width is consistent with that of a single physical vector register (e.g., 512 bits). It supports reading and writing at the vector element level. After completing the first-level cross-register vector addition, it can provide a unified, aligned, and reusable data carrier for subsequent reduction operations.

[0081] During the vector-level unordered summation reduction of all floating-point elements written to the first buffer register, standard reduction tree calculation is performed on N floating-point elements in a single buffer register to reuse reduction hardware resources under the path where LMUL < length multiple threshold, thus avoiding linear expansion of hardware size with LMUL.

[0082] As an example, this application can write the first vector operand completely into its storage space according to the write interface and timing control logic of the first buffer register, and trigger the reduction control signal after the write is completed to start the reduction tree operation.

[0083] As another example, this application can synchronously load vector mask information during the writing process based on the mask enable port of the first buffer register, so that the reduction phase automatically ignores the positions of elements that are turned off by the mask.

[0084] In some embodiments, writing the first vector operand into a first buffer register for vector-level unordered summation and reduction to obtain a vector reduction result includes: writing the first vector operand into the first buffer register; if the vector length ratio is equal to a length ratio threshold, performing vector-level unordered summation and reduction on the first vector operand in the first buffer register to obtain a vector reduction result; if the vector length ratio is greater than the length ratio threshold, performing floating-point unordered addition on two floating-point elements at the same position between the first buffer register and the third physical vector register in the logical vector register to obtain a second vector operand; and writing the first vector operand into the second buffer register for vector-level unordered summation and reduction to obtain a vector reduction result.

[0085] In this embodiment, if the vector length multiplier is equal to the length multiplier threshold, the vector reduction result can be obtained by calling a multi-level parallel floating-point adder to perform pairwise pairing and stepwise merging of all floating-point elements in the first buffer register.

[0086] In addition, this application can also obtain vector reduction results by enabling the reduction tree control logic and driving the output port of the first buffer register to be connected to the floating-point addition unit array in a hierarchical timing sequence.

[0087] For example, such as Figure 3As shown, when LMUL=2 and the length multiplier threshold is set to 2, the first buffer register has fully carried the sum of all corresponding elements of vs2 and vs3; a 4-level reduction tree is started for the 16 FP32 elements (VLEN=512, SEW=32) in the buffer: the 0th level generates 8 partial sums (R2[0], R2[1], R2[2], R2[3], R2[4], R2[5], R2[6], R2[7] respectively), the 1st level generates 4 (R3[0], R3[1], R3[2], R3[3] respectively), the 2nd level generates 2 (R4[0], R4[1] respectively), and the 3rd level generates 1 final reduction value r (i.e. R5[0]); this value r is then accumulated with vs1[0] and output to vd[0].

[0088] The third physical vector register can be understood as the next consecutive physical vector register that constitutes the current logical vector register and is numbered after the first and second physical vector registers; it must exist when LMUL≥3, and has the same element bit width and alignment as the first buffer register.

[0089] As an example, this application can obtain the second vector operand by simultaneously sending the corresponding elements of the first buffer register and the third physical vector register into the input port of the same set of floating-point addition units, and performing a floating-point addition under the drive of a control signal.

[0090] As another example, this application can also obtain the second vector operand by first broadcasting the contents of the first buffer register to multiple addition units, and then adding them in parallel with each element of the third physical vector register.

[0091] For example, under the LMUL=4 configuration, vs2, vs3, vs4, and vs5 constitute a logical vector; the first buffer register already contains R1=vs2+vs3; the third physical vector register is vs4; R1[i]+vs4[i] is executed to obtain R2[i], forming the second vector operand R2; after this process is completed, R2 is written to the second buffer register to prepare for subsequent superposition with vs5.

[0092] The second buffer register can be understood as a register with the same function as the first buffer register but physically independent. Its bit width, read / write interface, and timing characteristics are consistent with the first buffer register.

[0093] This application writes the first vector operand into the second buffer register. This does not mean repeatedly writing the original R1. Instead, after forming the second vector operand R2, the second vector operand R2 is written into the second buffer register to perform unordered summation and reduction at the vector level, and obtain the vector reduction result.

[0094] In this application, a structured temporary storage of intermediate results across registers is achieved by writing the first vector operand into the first buffer register. A third physical vector register is introduced to perform floating-point unordered addition to generate a second vector operand, which is then written into the second buffer register. This constructs a scalable, cascaded reduction chain, enabling the unified mapping of multiple physical vector register inputs under any LMUL configuration into a single-buffered vector form, which is then processed by an unordered reduction unit of fixed depth. This avoids customizing the reduction tree structure for different LMUL values, significantly reducing hardware complexity and timing risks, while ensuring the semantic integrity and execution efficiency of the vfredusum.vs instruction.

[0095] In some embodiments, such as Figure 4 As shown, the first physical vector register and the second physical vector register are two adjacent physical connected registers in the logical vector register; the first physical vector register or the second physical vector register is adjacent to the third physical vector register in the logical vector register.

[0096] In this embodiment, the first physical vector register can be understood as the physical vector register with the smallest number among the group of physical vector registers mapped by the current logical vector register, which is used to carry the starting part of the vector source operand; the second physical vector register can be understood as the register in the group of physical vector registers whose number is adjacent to the first physical vector register, and which is physically adjacent to and directly interconnected with it in the hardware layout; the third physical vector register can be understood as the register in the group of physical vector registers whose number is adjacent to the second physical vector register, and which shares a local interconnect bus with the second physical vector register at the chip routing level.

[0097] Specifically, there are no other intermediate physical vector registers allocated to the same logical vector register between the first physical vector register and the second physical vector register, and there are also no other intermediate physical vector registers allocated to the same logical vector register between the second physical vector register and the third physical vector register. This ensures that during cross-register floating-point unordered addition, the data path delay of the corresponding floating-point elements is consistent, the address generation logic is reusable, and the interconnection wiring length is controllable. This avoids the additional timing margin overhead introduced by non-adjacent register access and supports the stable operation of the multi-stage pipeline reduction structure at high frequencies.

[0098] In some embodiments, the number of floating-point bits of the floating-point element is any one of 8 bits, 16 bits, 32 bits, 64 bits, and 128 bits.

[0099] In this embodiment, the number of floating-point bits can be understood as the number of bits occupied by a single floating-point element in memory or register, corresponding to the SEW (Scalar Element Width) parameter in the RISC-V RVV specification.

[0100] The number of floating-point bits can be any of 8 bits (such as FP8 format), 16 bits (such as FP16 or bfloat16), 32 bits (IEEE 754 single-precision), 64 bits (IEEE 754 double-precision) or 128 bits (IEEE 754 quadruple-precision).

[0101] In some embodiments, the physical bit width is any one of 128 bits, 256 bits, and 512 bits, and the physical bit widths of the first physical vector register, the second physical vector register, and the third physical vector register are equal.

[0102] The physical bit width can be understood as the total number of bits that a single physical vector register can store; the physical bit width can be any of 128 bits, 256 bits, or 512 bits, corresponding to the VLEN configuration commonly found in mainstream RISC-V vector processor implementations.

[0103] Specifically, the physical bit widths of the first, second, and third physical vector registers are equal, which ensures that the floating-point elements output by each physical vector register have the same data organization format and alignment in cross-register element-level parallel addition operations. This avoids delays caused by splitting / reassembling operations, additional shift logic, or format conversion due to heterogeneous bit widths, allowing vector-level element-by-element addition to be completed on a unified data path, thus guaranteeing the timing convergence and area efficiency of the reduction link.

[0104] In some embodiments, before performing vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain the vector reduction result of the vector source operand, the method further includes: preprocessing the vector source operand to obtain the preprocessed vector source operand.

[0105] In this embodiment, the vector source operands can be preprocessed before the vector-level unordered summation and reduction is executed. Specifically, data validity checks and semantic alignment operations can be performed on each floating-point element in the original vector source operands to ensure that all floating-point elements input to the subsequent reduction process meet the semantic constraints of the IEEE 754 floating-point arithmetic specification and the RISC-V vector instruction vfredusum.vs. This avoids errors in the reduction results, hardware anomalies, or unreproducible instruction behavior due to invalid data. Thus, without increasing the complexity of the reduction core, the robustness, specification compliance, and industrial usability of the entire floating-point unordered addition operation method in a real processor environment can be significantly improved.

[0106] As an example, this application can perform a zeroing operation on the floating-point elements at the corresponding positions in the vector source operand based on the mask bit state of the vector mask vmask, so as to mask the elements that are turned off by the mask.

[0107] As another example, this application can also perform an addendum identity operation on the floating-point elements at the corresponding positions in the vector source operands according to the mask bit state of the vector mask vmask, replacing the masked elements with addendum identity (i.e., 0.0) in the floating-point format that matches the current SEW.

[0108] For example, when the vector source operand is v2={1.0,2.0,3.0,4.0}, the vector mask is vmask=4'b0111, VL=4, and SEW=32, the preprocessing process first sets v2[0] to 0.0 according to vmask to obtain the intermediate vector {0.0,2.0,3.0,4.0}; then, according to VL, it is confirmed that all four elements are within the valid range and no truncation is required; then, it is checked that each element is a normalized FP32 number and no normalization conversion is required; finally, the preprocessed vector source operand is output as {0.0,2.0,3.0,4.0}, which will be written into the logic vector register and used as the input for vector-level unordered summation and reduction of multiple floating-point elements in the logic vector register in this application.

[0109] S130. The vector reduction result is added to the least significant element in the scalar register to obtain the result of the vector source operand operation.

[0110] A scalar register can be understood as a register in a RISC-V integer register file or floating-point register file used to carry the scalar source operand vs1, and its contents are used as the initial accumulation value in the vfredusum.vs instruction.

[0111] The least significant element is the floating-point element with index 0 in the vector register, namely vs1[0], whose data format is consistent with the SEW of the floating-point element in the vector source operand.

[0112] Accumulation can be understood as performing an independent, non-reduction path floating-point unordered addition operation, with the input being the vector reduction result obtained above and vs1[0], and the output being the final operation result.

[0113] The result of the operation is the final value that conforms to the semantics of the vfredusum.vs instruction, i.e., vd[0]=vs1[0]+Σvs2[i],i∈[0,VL).

[0114] Specifically, this application can send the output vector reduction result to one input of a dedicated floating-point adder, and at the same time send the scalar register vs1[0] output through the scalar read port to the other input of the adder, so as to complete the floating-point addition and output the result in the same cycle.

[0115] For example, when the vector reduction result is 10.0 and vs1={1.0,0.0,0.0,0.0} (FP32), 10.0+1.0=11.0 is executed, and the result 11.0 is obtained and written to vd[0]; when vs1={-2.0,0.0,0.0,0.0}, 10.0+(-2.0)=8.0 is executed, and the result 8.0 is obtained.

[0116] In this application, by writing the vector source operands into a logical vector register consisting of at least one physical vector register, a unified abstraction and flexible mapping of vector data under different LMUL configurations is achieved. By using vector-level unordered summation and reduction, high-latency operations such as exponential sorting are avoided, and the reduction critical path is significantly compressed. The least significant element in the scalar register is used as an independent post-accumulation term, which satisfies the mandatory semantic requirement of the vfredusum.vs instruction for vs1[0] to participate in the final result, and prevents it from intervening in the reduction tree in advance, thereby destroying the unorderedness, increasing scheduling complexity and hardware overhead. Thus, this scheme can ensure that the IEEE 754 floating-point precision and RVV instruction semantics are strictly consistent, while taking into account high throughput, low latency and strong scalability, providing a solid technical foundation for subsequent dynamic path selection based on vector length ratio.

[0117] In the floating-point unordered addition method provided in this application, the vector source operand is obtained and written into a logical vector register composed of at least one physical vector register. This allows for dynamic adjustment of the number of physical vector registers used based on the vector length multiplier, achieving flexible storage and organization of the vector source operand. Furthermore, vector-level unordered summation and reduction are performed on multiple floating-point elements in the logical vector register to obtain the vector reduction result. This reduces computational depth and increases parallelism. Simultaneously, it adheres to the semantic specifications of the RISC-V vector extension instruction set, accumulating the vector reduction result with the least significant element in the scalar register to obtain the computation result of the vector source operand. This ensures the consistency of instruction semantics while avoiding interference with the parallel reduction tree structure, reducing hardware complexity. It effectively solves the technical problem of the performance bottleneck of vector reduction operations in the RISC-V vector extension instruction set, thus providing a superior hardware implementation solution for high-performance computing and AI acceleration.

[0118] In some embodiments, such as Figure 5 As shown, this application also provides a floating-point unordered addition operation device 200, including: a writing unit 210, a processing unit 220 and an accumulation unit 230.

[0119] The write unit 210 is used to obtain the vector source operand and write it into the logical vector register; the vector source operand includes multiple floating-point elements, and the logical vector register includes at least one physical vector register; the processing unit 220 is used to perform vector-level unordered summation and reduction on the multiple floating-point elements in the logical vector register to obtain the vector reduction result of the vector source operand; the accumulation unit 230 is used to accumulate the vector reduction result with the least significant bit in the scalar register to obtain the operation result of the vector source operand.

[0120] In the floating-point unordered addition apparatus 200 provided in this application, by acquiring the vector source operand and writing it into a logical vector register composed of at least one physical vector register, the number of physical vector registers used can be dynamically adjusted according to the vector length ratio, realizing flexible storage and organization of the vector source operand. Furthermore, vector-level unordered summation and reduction are performed on multiple floating-point elements in the logical vector register to obtain the vector reduction result, which reduces the computation depth and improves parallelism. Simultaneously, following the semantic specifications of the RISC-V vector extension instruction set, the vector reduction result is accumulated with the least significant element in the scalar register to obtain the computation result of the vector source operand. This ensures the consistency of instruction semantics, avoids interference with the parallel reduction tree structure, reduces hardware complexity, and effectively solves the technical problem of the performance bottleneck of vector reduction operation in the RISC-V vector extension instruction set, thus providing a better hardware implementation scheme for high-performance computing and AI acceleration.

[0121] Each module in the aforementioned floating-point unordered addition device 200 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0122] In some embodiments, this application also provides a processor that uses RISC-V vector unordered reduction floating-point summation instructions to execute the floating-point unordered addition method provided in this application.

[0123] Specifically, the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0124] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A floating-point unordered addition operation method, characterized in that, Using RISC The implementation of the vfredusum.vs instruction in the V-vector extension instruction set RVV includes the following methods: Obtain the vector source operand and write it into the logical vector register; the vector source operand includes multiple floating-point elements, and the logical vector register includes at least one physical vector register; the floating-point elements are floating-point numbers in 8-bit, 16-bit, 32-bit, 64-bit, or 128-bit floating-point formats conforming to the IEEE 754 standard. The vector-level unordered summation and reduction of multiple floating-point elements in the logical vector register are performed to obtain the vector reduction result of the vector source operand; The vector reduction result is added to the least significant element in the scalar register to obtain the result of the vector source operand operation; Before writing the vector source operand into the logic vector register, the method further includes: The vector length multiplier controlling the logic vector register is determined based on the vector bit width of the vector source operand and the physical bit width of the physical vector register. Based on the vector length multiplier, at least one of the physical vector registers is selected to form the logical vector register; The step of performing vector-level unordered summation and reduction on the multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operand includes: If the vector length multiplier is less than a preset length multiplier threshold, a multi-level parallel floating-point adder is used to perform vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain the vector reduction result. If the vector length multiplier is greater than or equal to the length multiplier threshold, perform a floating-point unordered addition operation on the two floating-point elements at the same position between the first physical vector register and the second physical vector register in the logical vector register to obtain the first vector operand. The first vector operand is written into the first buffer register to perform unordered summation and reduction at the vector level, thereby obtaining the vector reduction result.

2. The floating-point unordered addition method according to claim 1, characterized in that, The step of writing the vector source operand into the logic vector register includes: If the vector length multiplier is less than a preset length multiplier threshold, the vector source operand is written to one of the physical vector registers; If the vector length multiplier is greater than or equal to the length multiplier threshold, the vector source operand is written to one of the physical vector registers.

3. The floating-point unordered addition method according to claim 2, characterized in that, The step of writing the vector source operand to the plurality of physical vector registers includes: Based on the vector length multiplier, the vector source operand is split into multiple groups of floating-point numbers; each group of floating-point numbers includes multiple floating-point elements; Each set of floating-point numbers is input into one of the physical vector registers; each set of floating-point numbers corresponds to one physical vector register.

4. The floating-point unordered addition method according to claim 1, characterized in that, The step of writing the first vector operand into the first buffer register for unordered summation and reduction at the vector level to obtain the vector reduction result includes: Write the first vector operand to the first buffer register; If the vector length multiplier is equal to the length multiplier threshold, perform vector-level unordered summation and reduction on the first vector operands in the first buffer register to obtain the vector reduction result; If the vector length multiplier is greater than the length multiplier threshold, perform a floating-point unordered addition operation on the two floating-point elements at the same position between the first buffer register and the third physical vector register in the logical vector register to obtain the second vector operand; The first vector operand is written into the second buffer register to perform unordered summation and reduction at the vector level, thereby obtaining the vector reduction result.

5. The floating-point unordered addition method according to claim 4, characterized in that, The first physical vector register and the second physical vector register are two adjacent physically connected registers in the logical vector registers; the first physical vector register or the second physical vector register is adjacent to the third physical vector register in the logical vector registers; or / and, The physical bit width is any one of 128bit, 256bit, or 512bit, and the physical bit widths of the first physical vector register, the second physical vector register, and the third physical vector register are equal.

6. The floating-point unordered addition method according to any one of claims 1-5, characterized in that, Before performing vector-level unordered summation and reduction on the plurality of floating-point elements in the logical vector register to obtain the vector reduction result of the vector source operand, the method further includes: The vector source operands are preprocessed to obtain preprocessed vector source operands.

7. A floating-point unordered addition arithmetic device, characterized in that, Using RISC The device, which implements the vfredusum .vs instruction in the V-vector extension instruction set RVV, includes: The write unit is used to obtain vector source operands and write the vector source operands into a logical vector register; the vector source operands include multiple floating-point elements, and the logical vector register includes at least one physical vector register; the floating-point elements are floating-point numbers in 8-bit, 16-bit, 32-bit, 64-bit, or 128-bit floating-point formats conforming to the IEEE 754 standard. The processing unit is used to perform vector-level unordered summation and reduction on multiple floating-point elements in the logic vector register to obtain the vector reduction result of the vector source operand; An accumulation unit is used to accumulate the vector reduction result with the least significant element in the scalar register to obtain the operation result of the vector source operand; The floating-point unordered addition device is further configured to determine the vector length multiplier of the logic vector register based on the vector bit width of the vector source operand and the physical bit width of the physical vector register; and to select at least one of the physical vector registers to form the logic vector register based on the vector length multiplier. The processing unit is further configured to: if the vector length multiplier is less than a preset length multiplier threshold, use a multi-level parallel floating-point adder to perform vector-level unordered summation and reduction on multiple floating-point elements in the logical vector register to obtain the vector reduction result; if the vector length multiplier is greater than or equal to the length multiplier threshold, perform floating-point unordered addition on two floating-point elements at the same position between the first physical vector register and the second physical vector register in the logical vector register to obtain a first vector operand; and write the first vector operand into a first buffer register to perform vector-level unordered summation and reduction to obtain the vector reduction result.

8. A processor, characterized in that, The floating-point unordered addition operation method according to any one of claims 1-6 is executed using RISC-V vector unordered reduction floating-point summation instructions.

Citation Information

Patent Citations

  • Device and method for floating point complex number parallel addition and subtraction

    CN104866278A

  • Instruction execution method and device, electronic equipment and storage medium

    CN116243976A