RISC-V Vector Segment Write Instruction Processing Method and Apparatus

By storing segment data in temporary architecture registers, the performance reduction and power consumption increase caused by frequent memory access of RISC-V vector segment write instructions is solved, achieving more efficient CPU performance and lower power consumption.

CN118981337BActive Publication Date: 2025-06-03CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411116915.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2025-06-03
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

When implementing RISC-V vector segment writing instructions in the prior art, due to frequent memory access, the CPU system performance is reduced and power consumption is increased.

Method used

By assembling the metadata in multiple source registers into segment data in response to vector segment write instructions, and temporarily storing the segment data in the temporary architecture register, ultimately writing the segment data in the temporary architecture register to memory or cache.

Benefits of technology

Reduces cache/memory access times, improves CPU core performance, and reduces power consumption of accessing cache/memory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981337B_ABST
    Figure CN118981337B_ABST
Patent Text Reader

Abstract

This specification relates to the field of computer technologies, and specifically discloses a method and apparatus for processing RISC-V vector segment write instructions. The method includes: in response to a vector segment write instruction, assembling metadata in a plurality of source registers into segment data, and temporarily storing the segment data in a temporary architectural register; the segment data includes a plurality of metadata; the temporary architectural register is an architectural register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within an instruction; writing the segment data in the temporary architectural register into memory or a cache. The above solution converts segment data by using a temporary architectural register, writes data into memory in terms of segment data rather than metadata, and reduces the number of cache / memory accesses at the cost of increasing data transfer between vector registers, thereby improving the performance of the CPU core and reducing the power consumption of accessing the cache / memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technologies, and particularly to a method and apparatus for processing RISC-V vector segment write instructions. Background Art

[0002] The RISC-V vector instruction set is one of many standard instruction set extensions of RISC-V, which is used to support the single instruction multiple data (SIMD) computing mode to improve the computing power and efficiency of the corresponding CPU. The register width used by SIMD instructions is generally 128 bits or wider, and each vector register contains multiple integer metadata with a width of 8 / 16 / 32 / 64 bits or floating-point metadata with a width of 16 / 32 / 64 bits. In common SIMD instruction sets, the registers and data widths corresponding to each instruction are fixed. However, a RISC-V vector instruction can be used for different registers and data widths. The advantage of this is that an application using RISC-V vector instructions can run on multiple CPUs with different register widths without any modification.

[0003] There is a type of RISC-V vector data read / write instruction called the segment store instruction. The segment store instruction supports the above-mentioned several address calculation methods, but a block of segment data in memory corresponds to metadata at the same position from multiple registers. A block of segment data contains NFIELD blocks of metadata, where NFIELD can be an integer selected from 1 to 8. NFIELD = 1 is a normal, non-segment vector data read / write instruction. NFIELD>1 is a segment store instruction. NFIELD is used as a 3-bit field (nf[2:0]) in the opcode of the data read / write instruction.

[0004] For the segment store instruction, the general implementation method is to write each of its metadata as a data unit into memory. Although the function is correct, due to the large number of memory accesses and the small amount of data written to memory each time, it will affect the overall performance of the CPU and increase the power consumption of the CPU.

[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0006] Embodiments of this specification provide a method and apparatus for processing RISC-V vector segment write instructions to solve the problem in the prior art that the performance of the CPU system is reduced and the power consumption is increased due to frequent memory accesses when implementing the segment store instruction.

[0007] The embodiment of this specification provides a method for processing RISC-V vector segment write instructions, including:

[0008] In response to a vector segment write instruction, assemble the metadata in multiple source registers into segment data, and temporarily store the segment data in a temporary architectural register; the segment data includes multiple metadata; the temporary architectural register is an architectural register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction;

[0009] Write the segment data in the temporary architectural register into memory or cache.

[0010] In one embodiment, the method further includes:

[0011] Obtain multiple parameters of the vector segment write instruction; the multiple parameters include: address calculation mode, register group length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core;

[0012] Decode the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions; the multiple microinstructions include: integer value read microinstruction, segment transpose microinstruction, segment write address calculation microinstruction, and segment write data read microinstruction;

[0013] Correspondingly, assembling the metadata in multiple source registers into segment data and temporarily storing the segment data in a temporary architectural register includes:

[0014] Execute the segment transpose microinstruction to read the metadata from multiple source registers and store the read metadata in segments into a second temporary architectural register;

[0015] Correspondingly, writing the segment data in the temporary architectural register into memory or cache includes:

[0016] Execute the integer value read microinstruction to read the address base value or the address base value and offset value from the integer architectural register and store them in a first temporary architectural register;

[0017] Execute the segment write address calculation microinstruction to read the address information from the first temporary architectural register or the vector register according to the address calculation mode, perform address generation and address translation access based on the read address information to obtain physical address data, and write the physical address data into the first storage queue entry in the storage queue;

[0018] Execute the segment write data read microinstruction to read the segment data from the second temporary architectural register and write it into the second storage queue entry in the storage queue;

[0019] After the vector segment write instruction retires, write the segment data in the second storage queue entry into memory or cache according to the physical address data in the first storage queue entry.

[0020] In one embodiment, after writing the segment data in the second storage queue entry into memory or cache according to the physical address data in the first storage queue entry, it further includes:

[0021] Release the first storage queue entry and the second storage queue entry in the storage queue.

[0022] In one embodiment, decode the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions, including:

[0023] Determine a first quantity according to the address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions;

[0024] Determine a second quantity according to the register bank length and the number of metadata included in the segment data to generate a second quantity of segment transpose microinstructions;

[0025] Determine a third quantity based on the address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate a third quantity of segment write address calculation microinstructions and a third quantity of segment write data read instructions.

[0026] In one embodiment, the address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode;

[0027] Correspondingly, determine a first quantity according to the address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions, including:

[0028] When the address calculation mode is the unit stride address calculation mode or the index address calculation mode, generate one integer value read microinstruction to read the address base value and store it in a first temporary architecture register;

[0029] When the address calculation mode is the stride address calculation mode, generate two integer value read microinstructions to read the address base value and the offset value respectively and store them in two first temporary architecture registers respectively.

[0030] In one embodiment, determine a second quantity according to the register bank length and the number of metadata included in the segment data, including:

[0031] If the product of the number of metadata contained in the segment data and the length of the register bank is less than or equal to 4, then the second quantity is equal to the product of the number of metadata contained in the segment data and the length of the register bank; otherwise, the second quantity is equal to twice the product of the number of metadata contained in the segment data and the length of the register bank.

[0032] In one embodiment, the address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode;

[0033] Correspondingly, determining a third quantity based on the address calculation mode, the length of the register bank, the number of metadata contained in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core includes:

[0034] When the vector segment write instruction is in the unmasked mode and the address calculation mode is the unit stride address calculation mode, the third quantity is the product of the vector register width supported by the CPU core, the length of the register bank, and the number of metadata contained in the segment data divided by the cache read width supported by the CPU core;

[0035] When the vector segment write instruction is in the masked mode and the address calculation mode is the unit stride address calculation mode, or when the address calculation mode is the stride address calculation mode or the index address calculation mode, determine whether the minimum value between the product of the number of metadata contained in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the third quantity is the product of the vector register width supported by the CPU core, the length of the register bank, and the number of metadata contained in the segment data divided by the minimum value; otherwise, the third quantity is the product of the vector register width supported by the CPU core, the length of the register bank, and the number of metadata contained in the segment data divided by the minimum value and then plus the product of the length of the register bank and the number of metadata contained in the segment data minus 1.

[0036] In one embodiment, after decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions, it further includes:

[0037] Mapping the first temporary architecture register and the second temporary architecture register to physical registers; renaming the source register to the corresponding physical register; the segment write address calculation microinstruction has a data dependency on the integer value read microinstruction through the first temporary architecture register; the segment write data read microinstruction has a data dependency on the segment transpose microinstruction through the second temporary architecture register;

[0038] Distribute the microinstructions for reading integer values, the segment transpose microinstructions, the segment write address calculation microinstructions, and the segment write data reading microinstructions.

[0039] In one embodiment, when the vector segment write instruction is in mask mode, when executing the segment write address calculation microinstruction, if the mask value of the segment write address calculation microinstruction is zero, set the NOP bit in the first storage queue entry corresponding to the segment write address calculation microinstruction to 1, otherwise set it to 0;

[0040] When the vector segment write instruction is in mask mode, when executing the segment write data reading microinstruction, if the mask value of the segment write data reading microinstruction is zero, convert the segment write data reading microinstruction into a no-operation and do not write any data to the corresponding second storage queue entry.

[0041] The embodiments of this specification also provide a RISC-V vector segment write instruction processing device, including:

[0042] A temporary storage module, configured to assemble metadata in multiple source registers into segment data in response to a vector segment write instruction, and temporarily store the segment data in a temporary architecture register; the segment data includes multiple metadata; the temporary architecture register is an architecture register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction;

[0043] A writing module, configured to write the segment data in the temporary architecture register into memory or cache.

[0044] The embodiments of this specification provide a method for processing RISC-V vector segment write instructions. In response to a vector segment write instruction, metadata in multiple source registers can be assembled into segment data, and the segment data can be temporarily stored in a temporary architecture register. The segment data includes multiple metadata. The temporary architecture register is an architecture register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction. Then, the segment data in the temporary architecture register is written into memory or cache. The above solution converts segment data by using a temporary architecture register, writes data into memory in units of larger segment data (instead of multiple metadata within segment data), and reduces the number of cache / memory accesses at the cost of increasing data transfer between vector registers, thereby improving the performance of the CPU core and reducing the power consumption of accessing cache / memory.

[0045] Specific embodiments of the present invention are disclosed in detail with reference to the following description and the accompanying drawings, indicating the ways in which the principles of the present invention can be employed. It should be understood that the embodiments of the present invention are not limited in scope thereby. Features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or substituted for features in other embodiments.

[0046] It should be emphasized that the term "comprising / including" as used herein refers to the presence of features, components, steps, or assemblies, but does not preclude the presence or addition of one or more other features, components, steps, or assemblies. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present specification, form a part of the present specification, and do not limit the present specification. In the drawings:

[0048] Figure 1 A schematic diagram showing the format of the RISC-V vector segment write instruction;

[0049] Figure 2 Shows the data correspondence of the RISC-V vector segment write instruction with NFIELD being 3;

[0050] Figure 3 Shows the flowchart of the method for processing the RISC-V vector segment write instruction in an embodiment of the present specification;

[0051] Figure 4 Shows an example diagram of the method for processing the RISC-V vector segment write instruction in an embodiment of the present specification;

[0052] Figure 5 Shows a schematic diagram of the data dependency relationship between the microinstructions generated by a RISC-V vector segment write instruction in an embodiment of the present specification;

[0053] Figure 6 Shows a schematic diagram of the RISC-V vector segment write instruction processing apparatus in an embodiment of the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The principles and spirit of the present specification will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present specification, and do not limit the scope of the present specification in any way. On the contrary, these embodiments are provided to make the disclosure of the present specification more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0055] Those skilled in the art know that the embodiments of this specification can be implemented as a system, device, equipment, method, or computer program product. Therefore, the disclosure of this specification can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0056] In the RISC-V vector instruction set, the metadata width is represented by SEW (selected element width). One feature of the RISC-V vector instruction set is register grouping, and its setting is called LMUL (register group length). LMUL is used to indicate that the RISC-V vector architecture supports reading and writing multiple registers (referred to as a register group) or a part of a register with one instruction. When LMUL > 1, one instruction can read and write multiple consecutive vector registers. LMUL is configured by the application program at runtime, and the selectable values are [1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8].

[0057] Most RISC-V vector instructions support masked computing, that is, the instruction only writes the metadata with a mask (masked, that is, the corresponding mask bit is 1) in the destination register, and the value of the unmasked (corresponding mask bit is 0) metadata remains unchanged. For these instructions, the vm bit in the opcode is used to identify whether masked computing is performed. vm = 1 indicates unmasked computing, and all metadata calculation results will be written into the register. vm = 0 indicates masked computing, and the mask value comes from the architecture register v0. At this time, each bit starting from the low bit in v0 identifies a piece of metadata. When v0[i] is 1, the calculation result can be written to the corresponding i-th piece of metadata, otherwise it is not written.

[0058] Many CPU cores implement register renaming, which maps an architectural register to a physical register in the core to avoid false register dependencies such as write-after-read and write-after-write, and to improve the instruction level parallelism (ILP) and performance of the CPU. For such CPU cores that implement register renaming, masked instructions need to copy the unmasked metadata from the old target physical register A to the new target physical register B mapped to the same architectural register to ensure that the values of these metadata do not change. At this time, a masked instruction needs to read the values of four different physical registers: 2 sources, 1 mask, and one currently mapped target physical register.

[0059] RISC-V vector data load / store instructions have the following various address calculation methods.

[0060] Unit-stride (unit stride address calculation mode): The memory addresses corresponding to consecutive metadata are also consecutive, that is, the memory address interval between two consecutive metadata is the byte width of one metadata. Unit-stride load / store uses the base from an integer register to identify the address of the first metadata.

[0061] Strided (stride address calculation mode): The memory addresses of consecutive metadata have a fixed offset. This offset is not equal to the byte width of the corresponding metadata. Strided load / store calculates the address of each metadata in the form of base + N × offset. Both the base and offset values come from integer registers.

[0062] Indexed (indexed address calculation mode): Consecutive metadata are not consecutive in memory. The offset value corresponding to the metadata of an instruction comes from another (or a group of) vector registers. Indexed data load / store instructions have two subclasses: ordered and unordered. Indexed load / store calculates the address of each metadata in the form of base + V[i]. The base value comes from an integer register, while the offset value V[i] comes from the vector register (group).

[0063] Figure 1 Shows the format of the RISC-V vector store instruction. Figure 1The upper figure in it is the VS*unit-stride instruction, that is, a vector data storage instruction with the address calculation mode of unit-stride. Figure 1 The middle figure in it is the VSS*strided instruction, that is, a vector data storage instruction with the address calculation mode of strided. Figure 1 The lower figure in it is the VSX*indexed instruction, that is, a vector data storage instruction with the address calculation mode of indexed. Among them, nf is used to indicate that a segment of data in the segment write instruction contains (nf + 1) metadata. mew and width jointly determine the data width of the metadata and the index data, as shown in Table 1 below. vm is used to indicate whether the instruction is masked, as shown in Table 2 below. The mop field is used to identify the address calculation method. Among them, the mop field is used to identify the address calculation method. The corresponding relationship between the values of the mop field and the address calculation mode is shown in Table 3 below. sumop is used to further subdivide the unit-stride store instruction as follows: wholeregister store is a subclass of unit-stride store, which writes all the values in one or a group of registers into memory, as shown in Table 4 below. rs1 is the base address. vs3 is used to represent the store data. rs2 is the stride. vs2 is the address offset.

[0064] Table 1

[0065]

[0066] Table 2

[0067] vm Description 0 Vector result, only where vO.mask[i] = 1 1 Non-mask

[0068] Table 3

[0069]

[0070] Table 4

[0071]

[0072]

[0073] Figure 2 Shows the data correspondence of the segment strided store with NFIELD being 3. The RISC-V standard requires LMUL × NFIELD <= 8. The segment store instruction may also be masked or maskless, and the mask is at the segment granularity. Figure 2The upper line in [description] is the storage location of segment data in memory, and the following three lines are the corresponding locations in the source vector registers of the data written to memory. For the segment store instruction, the general implementation method is to write to memory / cache in terms of metadata granularity. This will result in excessive memory access for a single segment store instruction. For example, when the register width is 128 bits (VLEN = 128), a normal unit-stride store instruction with LMUL = 3 and SEW = 16 bits only needs to access memory 3 times to write the data of 3 registers to memory / cache. While a corresponding segment unit-stride store instruction with LMUL = 1, NFIELD = 3, and SEW = 16 bits needs to access memory 24 (i.e., 128×3 / 16) times. Compared with the data movement between registers, the bandwidth between the CPU core and cache / memory is limited, the latency is longer, and the power consumption for access is greater. Therefore, frequent memory access will reduce the CPU system performance and increase power consumption.

[0074] An embodiment of this specification provides a method for processing RISC-V vector segment write instructions. Figure 3 The flowchart of the method for processing RISC-V vector segment write instructions in an embodiment of this specification is shown. Although this specification provides method operation steps or device structures as shown in the following embodiments or drawings, based on routine or non-creative labor, more or fewer operation steps or module units may be included in the method or device. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of this specification and shown in the drawings. When the method or module structure is applied to an actual device or terminal product, it can be executed sequentially or in parallel according to the method or module structure connection shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, or even a distributed processing environment).

[0075] Specifically, as Figure 3 shown, the method for processing RISC-V vector segment write instructions provided by an embodiment of this specification may include the following steps.

[0076] Step S301, in response to a vector segment write instruction, assemble the metadata in multiple source registers into segment data, and temporarily store the segment data in a temporary architecture register; the segment data includes multiple metadata; the temporary architecture register is an architecture register defined in the CPU core implementation and is only used to temporarily store and transfer data within one instruction.

[0077] Step S302, write the segment data in the temporary architecture register to memory or cache.

[0078] In this embodiment, first, the metadata in multiple source registers is assembled into segment data, and these segment data are stored in temporary architectural registers (TARs), and then copied from the TARs to the cache / memory. Since this solution replaces multiple memory write microinstructions with register copy microinstructions inside the vector execution unit, both performance and power consumption are improved. In this embodiment, the difference between the TAR and the normal architectural register is that the TAR is only used inside one instruction to temporarily store and transfer data between its multiple microinstructions, while the normal architectural register needs to transfer data between multiple instructions.

[0079] Since one RISC-V instruction needs to read four register values (two source registers, mask, and the initial value of the destination register), it can be reasonably assumed that in a CPU core, each RISC-V vector instruction execution unit is allocated 4 vector register file read ports (that is, 4 different register values can be read from the vector register file in one clock cycle).

[0080] The vector segment write instruction includes two operations: write address calculation and obtaining the data to be written to memory from the register. In this solution, these two operations can generate a group of microinstructions: the segment write address calculation microinstruction VSTA for calculating the write address to memory and the segment write data read microinstruction VSTD for reading the data to be written to memory. A pair of VSTA and VSTD microinstructions complete an operation of writing data to memory. In this embodiment, VSTD sends the segment data from the TAR to the store queue to be finally written to the cache / memory.

[0081] For a high-performance CPU core with out-of-order execution, to ensure the correctness of program execution, the corresponding data of the vector segment write instruction will only be written to the cache / memory after retirement. In this case, there are multiple clock cycles from the execution to the retirement of a vector segment write instruction, during which its address and data are both stored in the store queue. Since the vector segment write instruction must write data to the memory / cache in order, the allocation of its store queue entries is generally also sequential. Therefore, a vector segment write instruction is generally allocated one or more store queue entries before it enters the dispatch stage.

[0082] In the above embodiment, by using the temporary architectural register to transform the segment data, writing data to memory in units of larger segment data rather than multiple metadata within the segment data, at the cost of increasing data transfer between vector registers, the cache / memory access times are reduced, thereby improving the performance of the CPU core and reducing the power consumption of accessing the cache / memory.

[0083] In some embodiments of the present specification, the method may further include: obtaining a plurality of parameters of a vector segment write instruction; the plurality of parameters may include: an address calculation mode, a register bank length, the number of metadata included in segment data, a metadata width, a vector register width supported by a CPU core, and a cache read width supported by the CPU core; decoding the vector segment write instruction based on the plurality of parameters to obtain a plurality of microinstructions; the plurality of microinstructions may include: an integer value read microinstruction, a segment transpose microinstruction, a segment write address calculation microinstruction, and a segment write data read microinstruction. Correspondingly, assembling the metadata in a plurality of source registers into segment data and temporarily storing the segment data in a temporary architecture register may include: executing the segment transpose microinstruction to read the metadata from the plurality of source registers and store the read metadata in segments into a second temporary architecture register. Correspondingly, writing the segment data in the temporary architecture register into memory or a cache may include: executing the integer value read microinstruction to read an address base value or an address base value and an offset value from an integer architecture register and store them in a first temporary architecture register; executing the segment write address calculation microinstruction to read address information from the first temporary architecture register or a vector register according to the address calculation mode, perform address generation and address translation access according to the read address information to obtain physical address data, and write the physical address data into a first storage queue entry in a storage queue; executing the segment write data read microinstruction to read the segment data from the second temporary architecture register and write it into a second storage queue entry in the storage queue; after the vector segment write instruction retires, writing the segment data in the second storage queue entry into memory or a cache according to the physical address data in the first storage queue entry.

[0084] In this embodiment, multiple parameters of the vector segment write instruction can be obtained. The multiple parameters may include: address calculation mode, register bank length, number of metadata included in segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core. Among them, the address calculation mode refers to the address calculation method of the vector instruction. The address calculation mode includes one of the following: Unit-stride address calculation mode, Strided address calculation mode, and Indexed address calculation mode. The register bank length LMUL means that the RISC-V vector architecture supports an instruction to read and write multiple registers (referred to as a register bank) or a part of a register. This is controlled by the LMUL field in the vtype vector VCSR, and the optional values of this field are [1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8]. The metadata included in the segment data refers to the metadata from the same position in multiple registers included in a segment data. The number of metadata NFIELD included in the segment data is the number of metadata included in a segment data. In the RISC-V vector architecture, a vector register contains multiple metadata. The size of each metadata (i.e., metadata width SEW) is controlled by the SEW field in the vtype vector VCSR. SEW can be 8 / 16 / 32 / 64-bit integer, or 32 / 64-bit floating point. The vector register width VLEN supported by the CPU core is the number of bits of the vector register used by a RISC-V CPU core, which needs to be a multiple of 2 and >= 128. The cache read width supported by the CPU core is DLEN, and generally DLEN is less than or equal to VLEN.

[0085] The vector segment write instruction can be decoded based on the multiple parameters to obtain multiple microinstructions. Specifically, the type, quantity, parameters, etc. of the microinstructions to be generated can be determined according to the multiple parameters to decode the vector segment write instruction and obtain multiple microinstructions. The multiple microinstructions include: integer value read microinstruction, segment transpose microinstruction, segment write address calculation microinstruction, and segment write data read microinstruction. Among them, the integer value read microinstruction is used to read the address base value or the address base value and offset value from the integer architecture register and store them in the first temporary architecture register. The segment transpose microinstruction is used to read the metadata from multiple source registers and store the read metadata in segments in the second temporary architecture register. The segment write address calculation microinstruction is used to read the address information in the first temporary architecture register or the vector register according to the address calculation mode, perform address generation and address translation access based on the read address information to obtain the physical address data, and write the physical address data into the first storage queue entry in the storage queue. The segment write data read microinstruction is used to read the segment data from the second temporary architecture register and write it into the second storage queue entry in the storage queue.

[0086] In one embodiment, 10 TAR registers can be set. Among them, TAR0 and TAR1 store the integer values required for address generation, while the other 8 TARs (TAR2 to TAR9) support the segment store instruction when LMUL×NFIELD is at most 8. It can be understood that other numbers of TAR registers can also be set. For example, 3 TAR registers can be set. Among them, registers TAR0 and TAR1 store the integer values required for address generation, and the other TAR2 is used to support the segment store instruction. In this case, more conversions and more microinstructions may be required. Another example is that 8 TAR registers can be set. Among them, registers TAR0 and TAR1 store the integer values required for address generation, and TAR2 to TAR7 are used to support the segment store instruction.

[0087] After decoding multiple microinstructions, the integer value reading microinstruction can be executed to read the address base value or the address base value and the offset value from the integer architecture register and store them in the first temporary architecture register. The segment transpose microinstruction can be executed to read the metadata from multiple source registers and store the read metadata in segments in the second temporary architecture register. The segment write address calculation microinstruction can be executed to read the address information in the first temporary architecture register or the vector register according to the address calculation mode, and perform address generation and address translation access based on the read address information to obtain the physical address data, and write the physical address data into the first storage queue entry in the storage queue. The segment write data reading microinstruction can also be executed to read the segment data from the second temporary architecture register and write it into the second storage queue entry in the storage queue. After the vector segment write instruction retires, the segment data in the second storage queue entry is written into the memory or cache according to the physical address data in the first storage queue entry. In the above manner, a series of microinstructions are used to assemble the metadata in multiple source registers into segment data, and these segment data are stored in the TAR, and then copied from the TAR to the cache / memory.

[0088] In some embodiments of this specification, after writing the segment data in the second storage queue entry into the memory or cache according to the physical address data in the first storage queue entry, it may further include: releasing the first storage queue entry and the second storage queue entry in the storage queue. After the segment data in the second storage queue entry is written into the memory or cache, the first storage queue entry and the second storage queue entry can be released.

[0089] In some embodiments of this specification, decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions may include: determining a first quantity according to the address calculation mode of the vector segment write instruction to generate the first quantity of integer value read microinstructions; determining a second quantity according to the register bank length and the number of metadata included in the segment data to generate the second quantity of segment transpose microinstructions; determining a third quantity based on the address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate the third quantity of segment write address calculation microinstructions and the third quantity of segment write data read instructions.

[0090] Specifically, the integer value read microinstruction is used to read the address base value or the address base value and the offset value, and the address-related parameters are related to the address calculation mode. The first quantity can be determined according to the vector read address calculation mode of the vector segment write instruction to generate the first quantity of integer value read microinstructions. The segment transpose microinstruction is used to read the metadata from multiple source registers and store the read metadata in the second temporary architecture register by segment. The second quantity can be determined according to the register bank length and the number of metadata included in the segment data to generate the second quantity of segment transpose microinstructions. The segment write address calculation microinstruction and the segment write data read microinstruction are respectively used to write the physical address data into the first storage queue entry in the storage queue and read the segment data from the second temporary architecture register and write it into the second storage queue entry in the storage queue. The number of segment write address calculation microinstructions and segment write data read microinstructions is the same. The third quantity can be determined based on the vector read address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate the third quantity of segment write address calculation microinstructions and segment write data read microinstructions. In the above manner, the number of microinstructions to be generated can be determined.

[0091] In some embodiments of this specification, the address calculation mode may include one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining a first quantity according to the address calculation mode of the vector segment write instruction to generate the first quantity of integer value read microinstructions may include: in the case where the address calculation mode is the unit stride address calculation mode or the index address calculation mode, generating one integer value read microinstruction to read the address base value and store it in one first temporary architecture register; in the case where the address calculation mode is the stride address calculation mode, generating two integer value read microinstructions to respectively read the address base value and the offset value and store them in two first temporary architecture registers.

[0092] In this embodiment, when the address calculation mode is the unit stride address calculation mode or the index address calculation mode, an integer value read micro-instruction is generated to read the address base value and store it in a first temporary architecture register. That is, a first temporary architecture register is used to store the address base value. When the address calculation mode is the stride address calculation mode, two integer value read micro-instructions are generated to read the address base value and the offset value respectively and store them in two first temporary architecture registers respectively. That is, a first temporary architecture register is used to store the address base value, and another first temporary architecture register is used to store the offset value.

[0093] In some embodiments of this specification, determining the second quantity according to the register bank length and the number of metadata included in the segment data may include: if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 4, then the second quantity is equal to the product of the number of metadata included in the segment data and the register bank length; otherwise, the second quantity is equal to twice the product of the number of metadata included in the segment data and the register bank length.

[0094] In this embodiment, a segment transpose micro-instruction VSEG copies metadata from at most 4 source registers to a TAR, and the data is stored in segments in the TAR, so that the data write operation can use these TARs as source registers and operate at the segment granularity. Whether in the masked or unmasked mode, the segment transpose micro-instruction VSEG ignores the mask, selects metadata from at most 4 source registers, and writes the merged segment data into a TAR. The masking operation is completed by the subsequent VSTA or VSTD micro-instructions. If LMUL × NFIELD <= 4, then one segment store instruction needs to generate LMUL × NFIELD VSEG micro-instructions, and each micro-instruction writes one TAR register. If LMUL × NFIELD > 4, then two segment transpose micro-instructions VSEG are required to fill the metadata in one TAR register, so 2 × LMUL × NFIELD segment transpose micro-instructions VSEG need to be generated.

[0095] In some embodiments of this specification, the address calculation mode may include one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining a third quantity based on the address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core may include: when the vector segment write instruction is in the maskless mode and the address calculation mode is the unit stride address calculation mode, the third quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the cache read width supported by the CPU core; when the vector segment write instruction is in the masked mode and the address calculation mode is the unit stride address calculation mode, or when the address calculation mode is the stride address calculation mode or the index address calculation mode, determining whether the minimum value of the product of the number of metadata included in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the third quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value; otherwise, the third quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value and then plus the product of the register bank length and the number of metadata included in the segment data minus 1.

[0096] In this embodiment, according to the different address calculation modes of a vector segment write instruction, the segment write address calculation micro-instruction VSTA reads TAR0, TAR1, or a vector register. According to the different address calculation modes of a segment store, the number of generated VSTAs is also different. In the mask calculation mode, the mask register itself is also a source of the VSTA instruction.

[0097] For the Maskless unit-stride segment store instruction (unit stride vector segment write instruction in the maskless mode), DLEN bits of data are written each time, so (VLEN × LMUL × NFIELD) / DLEN VSTA micro-instructions are required.

[0098] For vector segment write instructions in other cases (including masked unit-stride segment store, masked / maskless stride segment store, and masked / maskless indexed segment store), each write data is no more than min(DLEN, NFIELD×SEW) bits. If min(DLEN, NFIELD×SEW) is a power of 2, no write operation crosses the TAR register, so a total of (VLEN×LMUL×NFIELD) / min(DLEN, NFIELD×SEW) VSTA microinstructions are required. If min(DLEN, NFIELD×SEW) is not a power of 2, then (VLEN×LMUL×NFIELD) / min(DLEN, NFIELD×SEW)+LMUL×(NFIELD–1) VSTA microinstructions are needed. The extra LMUL×(NFIELD - 1) microinstructions are because some segment data needs to be stored in two TAR registers and thus needs to be split into two VSTA. The number of segment write data read microinstructions VSTD is the same as the number of corresponding segment write address calculation microinstructions VSTA. One VSTD reads the corresponding segment data from a TAR register and thus needs to include the data size and its offset in that TAR register.

[0099] In some embodiments of this specification, after decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions, it may further include: mapping the first temporary architecture register and the second temporary architecture register to physical registers; renaming the source register to the corresponding physical register; the segment write address calculation microinstruction has a data dependency on the integer value read microinstruction through the first temporary architecture register; the segment write data read microinstruction has a data dependency on the segment transpose microinstruction through the second temporary architecture register; distributing the integer value read microinstruction, the segment transpose microinstruction, the segment write address calculation microinstruction, and the segment write data read microinstruction.

[0100] In this embodiment, after decoding, the I2V and VSEG microinstructions have TAR target registers, and new physical registers are allocated to all these TARs. The source registers of these microinstructions are also renamed to the corresponding physical registers. The VSTA / VSTD microinstructions have data dependencies on the corresponding I2V / VSEG microinstructions through these TAR registers. Similar to ordinary instructions, when the source registers and other hardware resources (such as execution units and store queues, etc.) are ready, the microinstructions can be dispatched and executed. Among the microinstructions generated by decoding the same vector segment write instruction, the VSTA / VSTD microinstructions have data dependencies on the corresponding I2V / VSEG microinstructions through these TAR registers. Therefore, a VSTA must be executed after the I2V it depends on has been executed, and a VSTD must be executed after the VSEG it depends on has been executed.

[0101] In some embodiments of this specification, when the vector segment write instruction is in mask mode, when executing the segment write address calculation microinstruction, if the mask value of the segment write address calculation microinstruction is zero, the NOP bit in the first store queue entry corresponding to the segment write address calculation microinstruction is set to 1, otherwise it is set to 0; when the vector segment write instruction is in mask mode, when executing the segment write data read microinstruction, if the mask value of the segment write data read microinstruction is zero, the segment write data read microinstruction is converted into a no-operation, and no data is written to the corresponding second store queue entry.

[0102] When executing the segment write address calculation microinstruction VSTA, similar to ordinary data address calculation microinstructions, VSTA needs to go through steps such as address generation, address translation (virtual address to physical address) access, etc. VSTA writes the address part to the corresponding store queue entry. If the mask value of a VSTA in mask mode is 0, this VSTA sets the NOP bit of the corresponding store queue entry to 1, otherwise its NOP bit is set to 0. When executing the segment write data read microinstruction VSTD, the corresponding data is read from the TAR register and written into the corresponding store queue entry. If the mask value of a VSTD in mask mode is 0, this VSTD is converted into a no-operation (NOP), and no data is written into its corresponding store queue entry.

[0103] The retirement of the vector segment write instruction is the same as that of other write instructions. Only after a vector segment write instruction retires, the data in the store queue entry with its corresponding NOP bit being 0 will be written into the cache / memory. Then its store queue entry will be released.

[0104] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. Specifically, reference can be made to the descriptions of the relevant embodiments of the foregoing related processing, and details will not be repeated here.

[0105] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] The above method will be described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating this specification and does not constitute an improper limitation to this specification.

[0107] A RISC-V vector segment write instruction processing method is provided in this specific embodiment. This solution uses a series of microinstructions to assemble the metadata in multiple source registers into segment data, saves these segment data in temporary architecture registers, and then copies them from the temporary architecture registers to the cache / memory. Since this solution replaces multiple memory write microinstructions with register copy microinstructions inside the vector execution unit, both performance and power consumption are improved.

[0108] The difference between the temporary architecture register TAR and the normal architecture register is that the TAR is only used within one instruction to temporarily store and transfer data between its multiple microinstructions, while the normal architecture register needs to transfer data between multiple instructions. In this embodiment, 10 TAR registers are set, where TAR0 and TAR1 store the integer values required for address generation, and the other 8 TARs (TAR2 to TAR9) support the segment store instruction (vector segment write instruction) when LMUL×NFIELD is at most 8.

[0109] Since a RISC-V instruction needs to read four register values (two source registers, mask, and the initial value of the destination register), it can be reasonably assumed that in a CPU core, 4 vector register file read ports are allocated to each RISC-V vector instruction execution unit (i.e., 4 different register values can be read from the vector register file within one clock cycle).

[0110] A vector segment write instruction involves two operations: calculating the write address and obtaining the data to be written to memory from a register. In this solution, a set of microinstructions is generated for each of these two operations: VSTA calculates the address to be written to memory, while VSTD reads the data to be written to memory. A pair of VSTA and VSTD microinstructions completes an operation of writing data to memory. In this solution, the VSTD microinstruction sends the segment data from TAR to the store queue for eventual writing to the cache / memory.

[0111] In the address calculation of the vector store instruction, both the base (address base value) and the offset of the stride store come from integer registers. High-performance CPU cores generally have separate physical register files for integer and vector registers. When a vector instruction uses an integer register, corresponding I2V microinstructions are generated to read the values of the corresponding source registers from the integer register file and write them to the TAR target vector register. Other microinstructions in the same instruction obtain the integer values from this TAR. Since the RISC-V vector write instruction uses at most two integer registers, this solution reserves two TAR registers to temporarily store the integer values.

[0112] For a high-performance CPU core with out-of-order execution, to ensure the correctness of program execution, the corresponding data of a vector segment write instruction will only be written to the cache / memory after retirement. In this case, a vector segment write instruction has multiple clock cycles from execution to retirement, during which its address and data are stored in the store queue. Since the vector segment write instruction must write data to the memory / cache in order, the allocation of its store queue entries is generally also sequential. Therefore, a vector segment write instruction is generally allocated one or more store queue entries before it enters the dispatch stage.

[0113] Figure 4 The figure shows a schematic diagram of the RISC-V vector segment write instruction processing method in this specific embodiment. As Figure 4 shown is a maskless segment stride store instruction (strided vector segment write instruction in non-mask mode) with VLEN = 128b, NFIELD = 3, LMUL = 1, and SEW = 16b. In this optimized solution, this instruction is decomposed into 3 VSEG microinstructions to copy the metadata from the source register to the TAR register organized by segment data, and 10 (instead of 24 in the prior art solution) segment write address calculation microinstructions VSTA and 10 (instead of 24 in the prior art solution) segment write data read microinstructions VSTD.

[0114] This solution includes the following steps.

[0115] Decoding step. According to the LMUL, NFIELD, and SEW values of a vector segment write instruction and the vector register width VLEN supported by the CPU core, the instruction is decoded into multiple VSEG, VSTA, and VSTD microinstructions. The cache read width supported by the CPU core is DLEN, and generally DLEN <= VLEN.

[0116] The integer value read microinstruction I2V reads an integer architecture register and writes the result into a vector TAR. The Unit-stride and Indexed store instructions generate only one I2V to read the base value and put it into TAR0. The Stride store instruction needs to generate two I2Vs to read the base and offset values and put them into TAR0 and TAR1.

[0117] The segment transpose microinstruction VSEG copies metadata from up to 4 source registers into a TAR, and the data is stored in segments in the TAR, so that data write operations can use these TARs as source registers and operate at the segment granularity. Whether in masked or unmasked mode, one VSEG microinstruction ignores the mask, selects metadata from up to 4 source registers, combines the segment data, and writes it into a TAR. The masking operation is completed by the subsequent VSTA / VSTD microinstructions. If LMUL × NFIELD <= 4, one vector segment write instruction needs to generate LMUL × NFIELD VSEG microinstructions, and each microinstruction writes one TAR register. If LMUL × NFIELD > 4, two VSEG microinstructions are required to fill the metadata in one TAR register, so 2 × LMUL × NFIELD VSEG microinstructions need to be generated.

[0118] According to the different segment store address calculation modes, the segment write address calculation microinstruction VSTA reads TAR0, TAR1, or a vector register. According to the different address calculation modes of a vector segment write instruction, the number of generated VSTAs is also different. In masked mode, the mask register itself is also a source of the VSTA instruction.

[0119] For the Maskless unit-stride segment store instruction (unit-stride vector segment write instruction in unmasked mode), DLEN bits of data are written each time, so (VLEN × LMUL × NFIELD) / DLEN VSTA microinstructions are required.

[0120] For vector segment write instructions in other cases (including masked unit-stride segment store, masked / maskless stride segment store, and masked / maskless indexed segment store), each write data does not exceed min(DLEN, NFIELD×SEW) bits. If min(DLEN, NFIELD×SEW) is a power of 2, there is no store operation crossing the TAR register, so a total of (VLEN×LMUL×NFIELD) / min(DLEN, NFIELD×SEW) VSTA microinstructions are required. If min(DLEN, NFIELD×SEW) is a power of 2, then (VLEN×LMUL×NFIELD) / min(DLEN, NFIELD×SEW) + LMUL×(NFIELD–1) VSTA microinstructions are required. The extra LMUL×(NFIELD - 1) microinstructions are because some segment data needs to be stored in two registers and thus needs to be split into two VSTAs.

[0121] The number of microinstructions for segment write data reading VSTD: is the same as the number of corresponding microinstructions for segment write address calculation VSTA. One VSTD reads the corresponding segment data from a TAR register, so it needs to include the data size and its offset in that TAR register. In the masked mode, the mask register is also a source of the VSTA instruction.

[0122] Register renaming step. In a segment store, the I2V and VSEG microinstructions have TAR target registers, and new physical registers are allocated to these TARs. The source registers of these microinstructions are also renamed to the corresponding physical registers. The VSTA / VSTD microinstructions have data dependencies on the corresponding I2V / VSEG microinstructions through these TAR registers.

[0123] Microinstruction distribution step. Similar to ordinary instructions, when the source registers and other hardware resources (such as execution units and store queues, etc.) are ready, the microinstructions can be distributed and executed. Among the microinstructions generated by decoding the same vector segment write instruction, the VSTA / VSTD microinstructions have data dependencies on the corresponding I2V / VSEG microinstructions through these TAR registers. Therefore, a VSTA must be executed after the I2V it depends on has been executed, and a VSTD must be executed after the VSEG it depends on has been executed.

[0124] Micro-instruction execution steps. I2V copies the base / offset value from the integer physical register to TAR0 / TAR1. VSEG copies all relevant data (regardless of mask = 0 or 1) from the source register to TAR. Similar to the ordinary data address calculation micro-instruction, VSTA needs to go through steps such as address generation, address translation (virtual address to physical address) access, etc. VSTA writes the address part to the corresponding store queue entry. If the mask value of a VSTA in a mask mode is 0, then this VSTA sets the NOP bit of the corresponding store queue entry to 1, otherwise its NOP bit is set to 0. VSTD reads the corresponding data from the TAR register and writes it into the corresponding store queue entry. If the mask value of a VSTD in a mask mode is 0, then this VSTD is converted into a no-operation (NOP), and no data is written into its corresponding store queue entry.

[0125] Vector segment write instruction retirement steps. Similar to other write instructions, only after a vector segment write instruction retires, the data in the store queue entry with its corresponding NOP bit being 0 will be written into the cache / memory. Then multiple of its store queue entries will be released.

[0126] Figure 5 Shows the data dependency relationships between the micro-instructions generated by a vector segment write instruction, and how such data dependencies are established between them through the TAR register. Due to these dependency relationships, a VSTA must be executed after the I2V it depends on has been executed, and a VSTD must be executed after the VSEG it depends on has been executed. Figure 5 In it, TAR[0,1] are the first temporary architecture registers TAR0 and TAR1. TAR[2 - 9] are the second temporary architecture registers TAR2 to TAR9.

[0127] Since the metadata size in the vector register can vary, from 1 byte to 8 bytes. At the same time, different metadata in the same register may be set by different micro-instructions. And before a micro-instruction is executed, it must be ensured that every piece of metadata in each source register it uses has been obtained (ready). Therefore, each vector register file entry, in addition to the data of the VLEN bit, also has VLEN / 8 valid bits (valid bits), and each valid bit corresponds to one byte of data of this entry. For a masked segment store instruction, its VSEG micro-instruction writes all the read metadata into the TAR physical register and sets the corresponding valid bits. And a vector micro-instruction will only be dispatched and executed after all the valid bits of its source register have been set.

[0128] In addition, in the register renaming unit, since 10 new temporary architectural registers (TARs) are added, the corresponding architectural / predicted register mapping table also needs to add 10 entries to include these TARs.

[0129] For RISC-V vector segment write instructions, this technical solution uses temporary architectural registers to convert and save segment data. The memory write operation is granular with larger segment data (instead of multiple metadata within the segment data). At the cost of increasing data transfer between vector registers, the number of cache / memory writes is reduced, thereby improving the CPU core performance and reducing the power consumption of accessing the cache / memory.

[0130] In this solution, the release of the physical registers mapped by the temporary architectural registers is the same as that of ordinary architectural registers. When an architectural register mapping is written into the architectural mapping table, the physical register corresponding to the replaced item is released and can be used for new mappings. Since the temporary architectural registers only transfer data between multiple micro-instructions within an instruction, the physical registers they map can be released when the instruction retires. In this way, the utilization rate of physical registers can be improved. However, optimizing specifically for TAR register mapping will increase the logical complexity of register renaming.

[0131] In this embodiment, 8 TARs are reserved to temporarily store segment data. A possible solution may only need to reserve fewer TARs, such as only 4. The advantage of this solution is that the register mapping table entries of register renaming can be fewer, but the disadvantage is that when LMUL×NFIELD is greater than the number of reserved TARs, the renaming of a TAR to a physical register needs to be performed multiple times, and it may be necessary to generate more VSTD micro-instructions.

[0132] Based on the same inventive concept, an embodiment of this specification also provides a RISC-V vector segment write instruction processing device as described in the following embodiments. Since the principle of the RISC-V vector segment write instruction processing device for solving problems is similar to that of the RISC-V vector segment write instruction processing method, the implementation of the RISC-V vector segment write instruction processing device can refer to the implementation of the RISC-V vector segment write instruction processing method, and the repeated parts will not be elaborated. Hereinafter, the term "unit" or "module" may be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated. Figure 6 is a structural block diagram of the RISC-V vector segment write instruction processing device according to an embodiment of this specification, as Figure 6 shown, including: a temporary storage module 601 and a writing module 602. The following describes this structure.

[0133] The temporary storage module 601 is configured to assemble the metadata in multiple source registers into segment data in response to a vector segment write instruction, and temporarily store the segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction.

[0134] The writing module 602 is configured to write the segment data in the temporary architecture register into memory or cache.

[0135] In some embodiments of the present specification, the device further includes a decoding module, and the decoding module is specifically configured to: obtain multiple parameters of the vector segment write instruction; the multiple parameters include: address calculation mode, register group length, the number of metadata included in the segment data, metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core; decode the vector segment write instruction based on the multiple parameters to obtain multiple micro-instructions; the multiple micro-instructions include: integer value read micro-instruction, segment transpose micro-instruction, segment write address calculation micro-instruction, and segment write data read micro-instruction; correspondingly, the temporary storage module is specifically configured to: execute the segment transpose micro-instruction to read the metadata from multiple source registers and store the read metadata in segments into a second temporary architecture register; correspondingly, the writing module is specifically configured to: execute the integer value read micro-instruction to read the address base value or the address base value and the offset value from the integer architecture register and store them in a first temporary architecture register; execute the segment write address calculation micro-instruction to read the address information from the first temporary architecture register or the vector register according to the address calculation mode, and perform address generation and address translation access based on the read address information to obtain physical address data, and write the physical address data into the first storage queue entry in the storage queue; execute the segment write data read micro-instruction to read the segment data from the second temporary architecture register and write it into the second storage queue entry in the storage queue; after the vector segment write instruction retires, write the segment data in the second storage queue entry into memory or cache according to the physical address data in the first storage queue entry.

[0136] In some embodiments of the present specification, the device further includes a release module, and the release module is specifically configured to: after the writing module writes the segment data in the second storage queue entry into memory or cache according to the physical address data in the first storage queue entry, release the first storage queue entry and the second storage queue entry in the storage queue.

[0137] In some embodiments of this specification, the decoding module is specifically configured to: determine a first quantity according to the address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions; determine a second quantity according to the register bank length and the number of metadata included in the segment data to generate a second quantity of segment transpose microinstructions; determine a third quantity based on the address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate a third quantity of segment write address calculation microinstructions and a third quantity of segment write data read instructions.

[0138] In some embodiments of this specification, the address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining a first quantity according to the address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions includes: when the address calculation mode is the unit stride address calculation mode or the index address calculation mode, generating one integer value read microinstruction to read the address base value and store it in a first temporary architecture register; when the address calculation mode is the stride address calculation mode, generating two integer value read microinstructions to respectively read the address base value and the offset value and store them in two first temporary architecture registers.

[0139] In some embodiments of this specification, determining the second quantity according to the register bank length and the number of metadata included in the segment data includes: if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 4, the second quantity is equal to the product of the number of metadata included in the segment data and the register bank length; otherwise, the second quantity is equal to twice the product of the number of metadata included in the segment data and the register bank length.

[0140] In some embodiments of the present specification, the address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining the third quantity based on the address calculation mode of the vector segment write instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core includes: when the vector segment write instruction is in the unmasked mode and the address calculation mode is the unit stride address calculation mode, the third quantity is the product of the vector register width, the register bank length, and the number of metadata included in the segment data supported by the CPU core divided by the cache read width supported by the CPU core; when the vector segment write instruction is in the masked mode and the address calculation mode is the unit stride address calculation mode, or when the address calculation mode is the stride address calculation mode or the index address calculation mode, determining whether the minimum value of the product of the number of metadata included in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the third quantity is the product of the vector register width, the register bank length, and the number of metadata included in the segment data supported by the CPU core divided by the minimum value; otherwise, the third quantity is the product of the vector register width, the register bank length, and the number of metadata included in the segment data supported by the CPU core divided by the minimum value and then added to the product of the register bank length and the number of metadata included in the segment data minus 1.

[0141] In some embodiments of the present specification, the device further includes a distribution module, and the distribution module is specifically configured to: map the first temporary architecture register and the second temporary architecture register to physical registers; rename the source register to the corresponding physical register; the segment write address calculation micro-instruction has a data dependency on the integer value read micro-instruction through the first temporary architecture register; the segment write data read micro-instruction has a data dependency on the segment transpose micro-instruction through the second temporary architecture register; distribute the integer value read micro-instruction, the segment transpose micro-instruction, the segment write address calculation micro-instruction, and the segment write data read micro-instruction.

[0142] In some embodiments of this specification, when the vector segment write instruction is in mask mode, when executing the segment write address calculation micro-instruction, if the mask value of the segment write address calculation micro-instruction is zero, set the NOP bit in the first storage queue entry corresponding to the segment write address calculation micro-instruction to 1, otherwise set it to 0; when the vector segment write instruction is in mask mode, when executing the segment write data read micro-instruction, if the mask value of the segment write data read micro-instruction is zero, convert the segment write data read micro-instruction into a no-operation and do not write any data to the corresponding second storage queue entry.

[0143] Obviously, those skilled in the art should understand that the above-mentioned modules or steps in the embodiments of this specification can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of this specification are not limited to any specific combination of hardware and software.

[0144] It should be understood that the above description is for illustrative purposes rather than for limitation. By reading the above description, many embodiments and many applications other than the provided examples will be obvious to those skilled in the art. Therefore, the scope of this specification should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents of these claims.

[0145] The above are only the preferred embodiments of this specification and are not used to limit this specification. For those skilled in the art, the embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.

Claims

1. A RISC-V vector segment write instruction processing method, characterized in that: include: In response to a vector segment write instruction, assembling metadata in a plurality of source registers into segment data, and temporarily storing the segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in a CPU core implementation, and is only used to temporarily store and transfer data within an instruction; Writing the segment data in the temporary architectural register into a memory or cache; Wherein, the method further comprises: Acquire multiple parameters of a vector segment write instruction; the multiple parameters include: an address calculation mode, a register group length, a number of metadata contained in the segment data, a metadata width, a vector register width supported by the CPU core, and a cache read width supported by the CPU core; Decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions; the multiple microinstructions include: integer value read microinstructions, segment transpose microinstructions, segment write address calculation microinstructions and segment write data read microinstructions; Accordingly, assembling metadata in a plurality of source registers into segment data, and temporarily storing the segment data in a temporary architecture register, includes: Executing the segment transpose microinstruction to read metadata from a plurality of source registers and store the read metadata into a second temporary architecture register by segment; Accordingly, writing the segment data in the temporary architecture register into the memory or cache includes: Executing the integer value read microinstruction to read the address base value or the address base value and the offset value from the integer architecture register and store them in the first temporary architecture register; Execute the segment write address calculation microinstruction to read the address information in the first temporary architecture register or the vector register according to the address calculation mode, perform address generation and address translation access according to the read address information, obtain physical address data, and write the physical address data into a first storage queue item in the storage queue; executing the segment write data read microinstruction to read segment data from the second temporary architectural register and write into a second storage queue item in the storage queue; After the vector segment write instruction is retired, the segment data in the second storage queue entry is written into the memory or cache according to the physical address data in the first storage queue entry.

2. The RISC-V vector segment write instruction processing method according to claim 1, characterized in that: After writing the segment data in the second storage queue item into the memory or cache according to the physical address data in the first storage queue item, the method further includes: A first storage queue entry and a second storage queue entry in the storage queue are released.

3. The RISC-V vector segment write instruction processing method according to claim 1, characterized in that: Decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions, including: Determining a first quantity according to an address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions; Determine a second number according to the register group length and the number of metadata included in the segment data to generate a second number of segment transposition microinstructions; A third quantity is determined based on the address calculation mode of the vector segment write instruction, the register group length, the number of metadata contained in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate a third number of segment write address calculation microinstructions and a third number of segment write data read instructions.

4. The RISC-V vector segment write instruction processing method according to claim 3, characterized in that: The address calculation mode includes one of the following: a unit stride address calculation mode, a stride address calculation mode, and an index address calculation mode; Accordingly, determining a first quantity according to the address calculation mode of the vector segment write instruction to generate a first quantity of integer value read microinstructions includes: When the address calculation mode is a unit stride address calculation mode or an index address calculation mode, generating an integer value read microinstruction to read the address base value and store it in a first temporary architecture register; When the address calculation mode is the stride address calculation mode, two integer value read microinstructions are generated to read the address base value and the offset value respectively and store them in two first temporary architecture registers respectively.

5. The RISC-V vector segment write instruction processing method according to claim 3, characterized in that: Determining the second quantity according to the register group length and the number of metadata included in the segment data includes: If the product of the number of metadata contained in the segment data and the length of the register group is less than or equal to 4, the second number is equal to the product of the number of metadata contained in the segment data and the length of the register group; otherwise, the second number is equal to twice the product of the number of metadata contained in the segment data and the length of the register group.

6. The RISC-V vector segment write instruction processing method according to claim 3, characterized in that: The address calculation mode includes one of the following: a unit stride address calculation mode, a stride address calculation mode, and an index address calculation mode; Accordingly, the third quantity is determined based on the address calculation mode of the vector segment write instruction, the register group length, the number of metadata contained in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core, including: When the vector segment write instruction is in a maskless mode and the address calculation mode is a unit stride address calculation mode, the third number is a product of a vector register width supported by the CPU core, a register group length, and a number of metadata included in the segment data divided by a cache read width supported by the CPU core; When the vector segment write instruction is in mask mode and the address calculation mode is unit stride address calculation mode, or when the address calculation mode is stride address calculation mode or index address calculation mode, determine whether the minimum value of the product of the number of metadata contained in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the third number is the product of the vector register width supported by the CPU core, the register group length and the number of metadata contained in the segment data divided by the minimum value; otherwise, the third number is the product of the vector register width supported by the CPU core, the register group length and the number of metadata contained in the segment data divided by the minimum value plus the product of the register group length and the number of metadata contained in the segment data minus 1.

7. The RISC-V vector segment write instruction processing method according to claim 3, characterized in that: After decoding the vector segment write instruction based on the multiple parameters to obtain multiple microinstructions, the method further includes: The first temporary architectural register and the second temporary architectural register are mapped to physical registers; the source register is renamed to a corresponding physical register; the segment write address calculation microinstruction has a data dependency relationship with the integer value read microinstruction through the first temporary architectural register; the segment write data read microinstruction has a data dependency relationship with the segment transpose microinstruction through the second temporary architectural register; The integer value reading microinstruction, the segment transposition microinstruction, the segment write address calculation microinstruction and the segment write data reading microinstruction are distributed.

8. The RISC-V vector segment write instruction processing method according to claim 3, characterized in that: In the case where the vector segment write instruction is in mask mode, when executing the segment write address calculation microinstruction, if the mask value of the segment write address calculation microinstruction is zero, then the NOP bit in the first storage queue item corresponding to the segment write address calculation microinstruction is set to 1, otherwise it is set to 0; In the case where the vector segment write instruction is in mask mode, when executing the segment write data read microinstruction, if the mask value of the segment write data read microinstruction is zero, the segment write data read microinstruction is converted into a no operation.

9. A RISC-V vector segment write instruction processing device, characterized in that: include: A temporary storage module, configured to assemble metadata in a plurality of source registers into segment data in response to a vector segment write instruction, and temporarily store the segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in the CPU core implementation, and is only used to temporarily store and transfer data within an instruction; A writing module, used for writing the segment data in the temporary architecture register into a memory or a cache; The device further comprises an acquisition module, wherein the acquisition module is used to acquire multiple parameters of a vector segment write instruction; the multiple parameters include: an address calculation mode, a register group length, a number of metadata contained in segment data, a metadata width, a vector register width supported by a CPU core, and a cache read width supported by a CPU core; based on the multiple parameters, the vector segment write instruction is decoded to obtain multiple microinstructions; the multiple microinstructions include: an integer value read microinstruction, a segment transpose microinstruction, a segment write address calculation microinstruction, and a segment write data read microinstruction; Accordingly, the temporary storage module is specifically used to: execute the segment transposition microinstruction to read metadata from multiple source registers and store the read metadata in the second temporary architecture register by segment; Correspondingly, the write module is specifically used to: execute the integer value read microinstruction to read the address base value or the address base value and the offset value from the integer architecture register and store it in the first temporary architecture register; execute the segment write address calculation microinstruction to read the address information in the first temporary architecture register or the vector register according to the address calculation mode, and perform address generation and address translation access according to the read address information to obtain physical address data, and write the physical address data into the first storage queue item in the storage queue; execute the segment write data read microinstruction to read the segment data from the second temporary architecture register and write it into the second storage queue item in the storage queue; after the vector segment write instruction retires, write the segment data in the second storage queue item into the memory or cache according to the physical address data in the first storage queue item.

Citation Information

Patent Citations

  • Memory system and operating method of memory system

    CN115481056A

  • Memory access method and device, electronic equipment and readable storage medium

    CN116909946A