RISC-V Segment Data Reading Vector Instruction Processing Method and Apparatus
By using temporary architectural registers to handle segment data in RISC-V vector instructions, the method addresses performance and power consumption issues associated with frequent memory accesses, enhancing CPU efficiency through reduced access counts.
Patent Information
- Application Number
- CN202411114126.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-14
AI Technical Summary
In the prior art, frequent memory or cache access of RISC-V segment data read instructions leads to problems such as degradation of CPU performance and increased power consumption.
The temporary architecture register is used to temporarily store segment data, and copy the data from the temporary architecture register to the target register through a series of micro-instructions, reducing the number of memory or cache accesses.
Improves CPU core performance, reduces cache/memory access power consumption, and improves data transfer efficiency.
Smart Images

Figure CN118981336B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a method and apparatus for processing RISC-V segment data read vector instructions. Background Art
[0002] The RISC-V vector instruction set is one of many standard instruction set extensions of RISC-V, used to support the single instruction multiple data (SIMD) computing mode to improve the computing power and efficiency of the corresponding CPU. The register width used by SIMD instructions is generally 128 bits or wider, and each vector register contains multiple integer metadata with a width of 8 / 16 / 32 / 64 bits or floating-point metadata with a width of 16 / 32 / 64 bits. In common SIMD instruction sets, the registers and data widths corresponding to each instruction are fixed. However, a RISC-V vector instruction can be used for different registers and data widths. The advantage of this is that an application using RISC-V vector instructions can run on multiple CPUs with different register widths without any modification.
[0003] There is a class of RISC-V vector data read and write instructions called segment load / store instructions. Segment load / store instructions support the above-mentioned several address calculation methods, but a block of segment data in memory or cache corresponds to metadata from the same position of multiple registers. A block of segment data contains NFIELD blocks of metadata, where NFIELD can be an integer between 1 and 8. NFIELD = 1 is a normal, non-segment vector data read and write instruction. NFIELD > 1 is a segment load / store instruction. NFIELD is a 3-bit field in the opcode (operation code) of the data read and write instruction.
[0004] For the segment load instruction, a general implementation method is to access memory or cache with each metadata as a data read unit. Although the function is correct, due to the large number of memory or cache accesses and the small amount of data read each time, it will affect the overall performance of the CPU and increase the power consumption of the CPU.
[0005] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] The embodiments of this specification provide a method and apparatus for processing RISC-V segment data read vector instructions to solve the problem in the prior art that when implementing the segment load instruction, the CPU system performance is reduced and the power consumption is increased due to frequent memory or cache accesses.
[0007] The embodiments of this specification provide a method for processing RISC-V segment data read vector instructions, including:
[0008] In response to a segment data read vector instruction, read segment data from memory or cache and temporarily store the read segment data in a temporary architectural register; the segment data includes multiple metadata; the temporary architectural register is an architectural register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction;
[0009] Write the multiple metadata in the segment data in the temporary architectural register into multiple target registers respectively.
[0010] In one embodiment, the method further includes:
[0011] Obtain multiple parameters of the segment data read vector instruction; the multiple parameters include: vector read address calculation mode, register group length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core;
[0012] Decode the segment data read vector instruction based on the multiple parameters to obtain multiple microinstructions; the multiple microinstructions include: integer value read microinstructions, memory read microinstructions, and data movement microinstructions;
[0013] Correspondingly, in response to a segment data read vector instruction, read segment data from memory or cache and temporarily store the read segment data in a temporary architectural register, including:
[0014] Execute the integer value read microinstructions to read the address base value or the address base value and offset value from the integer architectural register and store them in the corresponding first temporary architectural register;
[0015] Execute the memory read microinstructions to read the address base value or the address base value and offset value from the first temporary architectural register, generate address data according to the address base value or the address base value and offset value, and read segment data from memory or cache based on the address data and write it into the corresponding second temporary architectural register;
[0016] Correspondingly, writing the multiple metadata in the segment data in the temporary architectural register into multiple target registers respectively includes:
[0017] Execute the data movement micro-instruction to copy multiple metadata in the segment data in the second temporary architecture register to multiple target registers respectively.
[0018] In one embodiment, decode the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions, including:
[0019] Determine a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read micro-instructions;
[0020] Determine a second quantity based on the vector read address calculation mode of the segment data read vector instruction, the register set length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core to generate a second quantity of memory read micro-instructions;
[0021] Determine a third quantity according to the register set length and the number of metadata included in the segment data to generate a third quantity of data movement micro-instructions.
[0022] In one embodiment, the vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode;
[0023] Correspondingly, determine a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read micro-instructions, including:
[0024] When the vector read address calculation mode is the unit stride address calculation mode or the index address calculation mode, generate an integer value read micro-instruction to read the address base value and store it in a first temporary architecture register;
[0025] When the vector read address calculation mode is the stride address calculation mode, generate two integer value read micro-instructions to read the address base value and the offset value respectively and store them in two first temporary architecture registers respectively.
[0026] In one embodiment, the vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode;
[0027] Correspondingly, determine a second quantity based on the vector read address calculation mode of the segment data read vector instruction, the register set length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core, including:
[0028] When the vector read address calculation mode is the unit stride address calculation mode, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata contained in the segment data, divided by the cache read width supported by the CPU core;
[0029] When the vector read address calculation mode is the stride address calculation mode or the index address calculation mode, determine whether the minimum value of the product of the number of metadata contained in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata contained in the segment data, divided by the minimum value; otherwise, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata contained in the segment data, divided by the minimum value, and then added with the product of the register bank length and the number of metadata contained in the segment data minus 1.
[0030] In one embodiment, determining a third quantity according to the register bank length and the number of metadata contained in the segment data includes:
[0031] When the segment data read vector instruction is in the unmasked mode, if the product of the number of metadata contained in the segment data and the register bank length is less than or equal to 4, the third quantity is equal to the number of metadata contained in the segment data; otherwise, the third quantity is equal to twice the number of metadata contained in the segment data;
[0032] When the segment data read vector instruction is in the masked mode, if the product of the number of metadata contained in the segment data and the register bank length is less than or equal to 3, the third quantity is equal to the number of metadata contained in the segment data; if the product of the number of metadata contained in the segment data and the register bank length is greater than 3 and less than or equal to 6, the third quantity is equal to twice the number of metadata contained in the segment data; if the product of the number of metadata contained in the segment data and the register bank length is greater than 6, the third quantity is equal to three times the number of metadata contained in the segment data.
[0033] In one embodiment, the data movement micro-instruction includes the position of the first metadata to be copied in the second temporary architecture register and the position of the first metadata in the target register.
[0034] In one embodiment, decoding the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions further includes:
[0035] When the segment data read vector instruction is in the mask mode, generate a fourth quantity of vector move micro-instructions, where the fourth quantity is the product of the register bank length and the number of metadata included in the segment data; the vector move micro-instructions are used to copy the unmasked calculated metadata in the source register to the target register.
[0036] In one embodiment, after decoding the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions, it further includes:
[0037] Map the first temporary architecture register and the second temporary architecture register to physical registers;
[0038] Allocate physical registers to the target registers corresponding to the integer value read micro-instruction, the memory read micro-instruction, and the data move micro-instruction; the source register of the vector move micro-instruction uses the initially mapped physical register of the target architecture register, and the target register of the vector move micro-instruction uses the physical register allocated to the data move micro-instruction;
[0039] Distribute the integer value read micro-instruction, the memory read micro-instruction, the data move micro-instruction, and the vector move micro-instruction; the memory read micro-instruction has a data dependency on the integer value read micro-instruction; the data move micro-instruction has a data dependency on the memory read micro-instruction, and the vector move micro-instruction has no dependency on the integer value read micro-instruction, the memory read micro-instruction, and the data move micro-instruction.
[0040] An embodiment of this specification also provides a RISC-V segment data read vector instruction processing device, including:
[0041] A reading module, configured to read segment data from a memory or a cache in response to a segment data read vector instruction, and temporarily store the read segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction;
[0042] A writing module, configured to write the plurality of metadata in the segment data in the temporary architecture register into a plurality of target registers respectively.
[0043] In an embodiment of this specification, a method for processing a RISC-V segment data read vector instruction is provided. In response to a segment data read vector instruction, segment data can be read from memory or cache and temporarily stored in a temporary architectural register; the segment data includes multiple metadata; the temporary architectural register is an architectural register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within an instruction; multiple metadata in the segment data in the temporary architectural register are respectively written into multiple target registers. This solution reduces the number of cache / memory accesses at the cost of increasing data transfer between vector registers by using a temporary architectural register to save segment data, with the memory or cache access granularity being larger segment data (instead of multiple metadata within the segment data), thereby improving the CPU core performance and reducing the power consumption of accessing the cache / memory. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and do not limit this specification. In the drawings:
[0045] Figure 1 A schematic diagram showing the format of a RISC-V vector data read instruction;
[0046] Figure 2 A diagram showing the data correspondence of a RISC-V segment data read vector instruction with NFIELD being 3;
[0047] Figure 3 A flowchart showing the method for processing a RISC-V segment data read vector instruction in an embodiment of this specification;
[0048] Figure 4 An example diagram showing the method for processing a RISC-V segment data read vector instruction in an embodiment of this specification;
[0049] Figure 5 A schematic diagram showing the data dependency relationship between microinstructions generated by a RISC-V segment data read vector instruction in an embodiment of this specification;
[0050] Figure 6 A schematic diagram showing a RISC-V segment data read vector instruction processing device in an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The principles and spirit of this specification will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement this specification, and do not limit the scope of this specification in any way. On the contrary, these embodiments are provided to make the disclosure of this specification more thorough and complete, and to be able to convey the scope of this disclosure fully to those skilled in the art.
[0052] Those skilled in the art know that the embodiments of this specification can be implemented as a system, device, equipment, method, or computer program product. Therefore, the disclosure of this specification can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0053] In the RISC-V vector instruction set, the metadata width is represented by SEW (selected element width). One feature of the RISC-V vector instruction set is register grouping, and its setting is called LMUL (register group length). LMUL is used to indicate that the RISC-V vector architecture supports reading and writing multiple registers (referred to as a register group) or a part of a register by one instruction. When LMUL > 1, one instruction can read and write multiple consecutive vector registers. LMUL is configured by the application program at runtime, and the selectable values are [1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8].
[0054] Most RISC-V vector instructions support masked computing, that is, the instruction only writes the metadata with a mask (i.e., the corresponding mask bit is 1) in the destination register, and the value of the metadata without a mask (i.e., the corresponding mask bit is 0) remains unchanged. For these instructions, the vm bit in the opcode is used to identify whether masked computing is performed. vm = 1 indicates unmasked computing, and all metadata calculation results will be written into the register. vm = 0 indicates masked computing, and the mask value comes from the architecture register v0. At this time, each bit starting from the low bit in v0 identifies a metadata. When v0[i] is 1, the calculation result can be written into the corresponding i-th metadata, otherwise it is not written.
[0055] Many CPU cores implement register renaming, that is, mapping an architectural register to a physical register in the core to avoid false register dependencies such as write-after-read and write-after-write, so as to improve the instruction level parallelism (ILP) and performance of the CPU. For such CPU cores that implement register renaming, masked instructions need to copy the unmasked metadata from the old target physical register A to the new target physical register B mapped by the same architectural register to ensure that the values of these metadata do not change. At this time, a masked instruction needs to read the values of four different physical registers: 2 sources, 1 mask, and one currently mapped target physical register.
[0056] RISC-V vector data load / store instructions have the following multiple address calculation methods.
[0057] Unit-stride (unit stride address calculation mode): The memory addresses corresponding to consecutive metadata are also consecutive, that is, the memory address interval between two consecutive metadata is the byte width of one metadata. Unit-stride load / store uses the base from an integer register to identify the address of the first metadata.
[0058] Strided (stride address calculation mode): The memory addresses of consecutive metadata have a fixed offset. This offset is not equal to the byte width of the corresponding metadata. Strided load / store calculates the address of each metadata in the form of base + N × offset. The values of base and offset both come from integer registers.
[0059] Indexed (indexed address calculation mode): Consecutive metadata are not consecutive in memory. The offset value corresponding to the metadata of an instruction comes from another (or a group of) vector registers. Indexed data load / store instructions have two subclasses: ordered and unordered. Indexed load / store calculates the address of each metadata in the form of base + V[i]. The value of base comes from an integer register, and the offset value V[i] comes from the vector register (group).
[0060] Figure 1 Shows the format of the RISC-V vector data load instruction. Figure 1The upper figure in it is the VL*unit-stride instruction, that is, the vector data reading instruction with the address calculation mode of unit-stride. Figure 1 The middle figure in it is the VLS*strided instruction, that is, the vector data reading instruction with the address calculation mode of strided. Figure 1 The lower figure in it is the VLX*indexed instruction, that is, the vector data reading instruction with the address calculation mode of indexed.
[0061] Among them, nf is used to indicate that a segment of data in the segment read / write instruction contains (nf + 1) metadata. mew and width jointly determine the data width of the metadata and the index data, as shown in Table 1 below.
[0062] vm is used to indicate whether this instruction is in the mask mode, as shown in Table 2 below. The mop field is used to identify the address calculation method. The corresponding relationship between the values of the mop field and the address calculation mode is shown in Table 3 below. rs1 is the address base value (baseaddress). vd is used to represent the destination register of the load (destination of load), rs2 is the stride, and vs2 is the address offset value (address offsets). Among them, the mop field is used to identify the address calculation method.
[0063] Lumop is used to further subdivide the unit-stride load instruction as follows: whole register load is a subclass of unit-stride load, which reads all the values in one or a group of registers from memory or cache, as shown in Table 4 below.
[0064] Table 1
[0065]
[0066] Table 2
[0067] vm Description 0 Vector result, only where vO.mask[i] = 1 1 Non-mask
[0068] Table 3
[0069]
[0070] Table 4
[0071]
[0072]
[0073] Figure 2 Shows the data correspondence of the segment strided load with NFIELD being 3.Figure 2 The upper line is the storage location of segment data in memory, and the following three lines are the corresponding locations of the read data in the target vector register. The RISC-V standard requires that LMUL × NFIELD <= 8. The segment load instruction can also be masked or maskless, and the mask is at the segment granularity. For the segment load instruction, the general implementation method is to access memory at the metadata granularity. This causes excessive memory access for a segment load instruction. For example, when the register width is 128 bits (VLEN = 128), a normal unit-stride load instruction with LMUL = 3 and SEW = 16 bits only needs to access memory 3 times to read the data of 3 registers. While a corresponding segment unit-stride load instruction with LMUL = 1, NFIELD = 3, and SEW = 16 bits needs to access memory 24 (i.e., 128 × 3 / 16) times. Compared with the data movement between registers, the CPU core to cache / memory bandwidth is limited, the latency is longer, and the power consumption for access is greater. Therefore, frequent memory access will reduce the CPU system performance and increase the power consumption.
[0074] To solve the above problems, the embodiments of this specification provide a method for processing RISC-V segment data read vector instructions. Figure 3 The flowchart of the method for processing RISC-V segment data read vector instructions in an embodiment of this specification is shown. Although this specification provides the method operation steps or device structures as shown in the following embodiments or drawings, based on routine or non-creative labor, more or fewer operation steps or module units may be included in the method or device. In steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure described in the embodiments of this specification and shown in the drawings. When the method or module structure is applied to the actual device or terminal product, it can be executed sequentially or in parallel according to the method or module structure connection shown in the embodiments or drawings (for example, in an environment of parallel processors or multi-threaded processing, or even a distributed processing environment).
[0075] Specifically, as Figure 3 shown, the method for processing RISC-V segment data read vector instructions provided by an embodiment of this specification may include the following steps.
[0076] Step S301: In response to a segment data read vector instruction, read segment data from memory or cache and temporarily store the read segment data in a temporary architectural register; the segment data includes multiple metadata; the temporary architectural register is an architectural register defined in the CPU core implementation and is only used to temporarily store and transfer data within one instruction.
[0077] Step S302: Write the multiple metadata in the segment data in the temporary architectural register into multiple target registers respectively.
[0078] In this embodiment, a temporary architectural register (TAR) is used to temporarily store the segment data obtained from memory or cache, and then a series of microinstructions are used to copy the data from the TAR to the target registers. Since this solution replaces the memory read microinstructions with register copy microinstructions inside the vector execution unit, both performance and power consumption are improved. In this embodiment, the difference between the TAR and a normal architectural register is that the TAR is only used within one instruction to temporarily store and transfer data between its multiple microinstructions. However, a normal architectural register needs to transfer data between multiple instructions.
[0079] Since a RISC-V segment data read instruction needs to read four register values (two source registers, a mask, and the initial value of the target register), it can be reasonably assumed that in a CPU core, each RISC-V vector instruction execution unit is allocated 4 vector register file read ports (i.e., 4 different register values can be read from the vector register file within one clock cycle).
[0080] In this embodiment, by using a temporary architectural register to save segment data, the memory or cache access is at the granularity of larger segment data (instead of multiple metadata within the segment data). At the cost of increasing data transfer between vector registers, the number of cache / memory accesses is reduced, thereby improving the CPU core performance and reducing the power consumption of accessing the cache / memory.
[0081] In some embodiments of this specification, the method further includes: obtaining multiple parameters of a segment data read vector instruction; the multiple parameters include: a vector read address calculation mode, a register bank length, the number of metadata included in the segment data, a metadata width, a vector register width supported by a CPU core, and a cache read width supported by the CPU core; decoding the segment data read vector instruction based on the multiple parameters to obtain multiple microinstructions; the multiple microinstructions include: an integer value read microinstruction, a memory read microinstruction, and a data movement microinstruction; correspondingly, in response to the segment data read vector instruction, reading segment data from a memory or a cache and temporarily storing the read segment data in a temporary architecture register, including: executing the integer value read microinstruction to read an address base value or an address base value and an offset value from an integer architecture register and store them in a corresponding first temporary architecture register; executing the memory read microinstruction to read the address base value or the address base value and the offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read the segment data from the memory or the cache based on the address data and write it into a corresponding second temporary architecture register; correspondingly, writing multiple metadata in the segment data in the temporary architecture register into multiple target registers respectively, including: executing the data movement microinstruction to copy the multiple metadata in the segment data in the second temporary architecture register into multiple target registers respectively.
[0082] In this embodiment, multiple parameters of the segment data read vector instruction can be obtained. The multiple parameters may include: a vector read address calculation mode, a register bank length, the number of metadata included in the segment data, a metadata width, a vector register width supported by the CPU core, and a cache read width supported by the CPU core. Among them, the vector read address calculation mode refers to the address calculation method of the vector instruction. The vector read address calculation mode includes one of the following: Unit-stride address calculation mode, Strided address calculation mode, and Indexed address calculation mode. The register bank length LMUL means that the RISC-V vector architecture supports an instruction to read and write multiple registers (referred to as a register bank) or a part of a register. This is controlled by the LMUL field in the vtype vector VCSR, and the optional values of this field are [1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8]. The metadata included in the segment data refers to the metadata from the same position in multiple registers included in a segment data. The number of metadata NFIELD included in the segment data is the number of metadata included in a segment data. In the RISC-V vector architecture, a vector register contains multiple metadata. The size of each metadata is controlled by the SEW field in the vtype vector VCSR. The metadata width SEW can be 8 / 16 / 32 / 64-bit integer, or 32 / 64-bit floating point. The vector register width VLEN supported by the CPU core is the number of bits of the vector register used by a RISC-V CPU core, which needs to be a multiple of 2 and >= 128. The cache read width supported by the CPU core is DLEN, and generally DLEN <= VLEN.
[0083] The segment data read vector instruction can be decoded based on the multiple parameters to obtain multiple microinstructions. Specifically, the type, quantity, parameters, etc. of the microinstructions to be generated can be determined according to the multiple parameters to decode the end data read vector instruction to obtain multiple microinstructions. The multiple microinstructions include: integer value read microinstructions, memory read microinstructions, and data movement microinstructions. Among them, the integer value read microinstructions are used to read the address base value or the address base value and the offset value from the integer architecture register and store them in the corresponding first temporary architecture register. The memory read microinstructions are used to read the address base value or the address base value and the offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read the segment data from the memory or cache based on the address data and write it into the corresponding second temporary architecture register. The data movement microinstructions are used to copy multiple metadata in the segment data in the second temporary architecture register to multiple target registers respectively.
[0084] In one embodiment, 10 TAR registers may be set, where TAR0 and TAR1 store the integer values required for address generation, and the other 8 TARs (TAR2 to TAR9) support the segment load instruction in the case where LMUL×NFIELD is at most 8. It can be understood that other numbers of TAR registers may also be set. For example, 3 TAR registers may be set, where registers TAR0 and TAR1 store the integer values required for address generation, and the other TAR2 is used to support the segment load instruction. In this case, more conversions and more microinstructions may be required.
[0085] The integer value reading microinstruction may be executed to read the address base value or the address base value and the offset value from the integer architecture register and store them in the corresponding first temporary architecture register. The memory reading microinstruction may also be executed to read the address base value or the address base value and the offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read the segment data from the memory or cache based on the address data and write it into the corresponding second temporary architecture register. The data movement microinstruction may also be executed to copy multiple metadata in the segment data in the second temporary architecture register to multiple target registers respectively. In the above manner, the temporary architecture register is used to temporarily store the segment data obtained from the memory or cache, and then a series of microinstructions are used to copy the data from the TAR to the target register, replacing the memory reading microinstruction with the register copy microinstruction inside the vector execution unit, and both the performance and power consumption are improved.
[0086] In some embodiments of the present specification, the segment data reading vector instruction is decoded based on the multiple parameters to obtain multiple microinstructions, including: determining a first quantity according to the vector reading address calculation mode of the segment data reading vector instruction to generate a first quantity of integer value reading microinstructions; determining a second quantity based on the vector reading address calculation mode of the segment data reading vector instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache reading width supported by the CPU core to generate a second quantity of memory reading microinstructions; determining a third quantity according to the register bank length and the number of metadata included in the segment data to generate a third quantity of data movement microinstructions.
[0087] Specifically, the integer value reading micro-instruction is used to read the base address value or the base address value and the offset value, and the address-related parameters are related to the vector reading address calculation mode. Therefore, the first quantity can be determined according to the vector reading address calculation mode of the segment data reading vector instruction to generate the first quantity of integer value reading micro-instructions. The memory reading micro-instruction is used to read the end data from the memory or cache. Therefore, the second quantity can be determined based on the vector reading address calculation mode of the segment data reading vector instruction, the length of the register bank, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache reading width supported by the CPU core to generate the second quantity of memory reading micro-instructions. The data movement micro-instruction is used to read the metadata in the segment data from the second temporary architecture register. Therefore, the third quantity can be determined according to the length of the register bank and the number of metadata included in the segment data to generate the third quantity of data movement micro-instructions. In the above manner, the quantity of micro-instructions to be generated can be determined.
[0088] In some embodiments of the present specification, the vector reading address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining the first quantity according to the vector reading address calculation mode of the segment data reading vector instruction to generate the first quantity of integer value reading micro-instructions includes: when the vector reading address calculation mode is the unit stride address calculation mode or the index address calculation mode, generating one integer value reading micro-instruction to read the base address value and store it in a first temporary architecture register; when the vector reading address calculation mode is the stride address calculation mode, generating two integer value reading micro-instructions to respectively read the base address value and the offset value and store them in two first temporary architecture registers respectively.
[0089] When the vector reading address calculation mode is the unit stride address calculation mode or the index address calculation mode, generating one integer value reading micro-instruction to read the base address value and store it in a first temporary architecture register. That is, one first temporary architecture register is used to store the base address value. When the vector reading address calculation mode is the stride address calculation mode, generating two integer value reading micro-instructions to respectively read the base address value and the offset value and store them in two first temporary architecture registers respectively. That is, one first temporary architecture register is used to store the base address value, and the other first temporary architecture register is used to store the offset value.
[0090] In some embodiments of the present specification, the vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode. Correspondingly, determining a second quantity based on the vector read address calculation mode of the segment data read vector instruction, the register set length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core includes: when the vector read address calculation mode is the unit stride address calculation mode, the second quantity is the product of the vector register width supported by the CPU core, the register set length, and the number of metadata included in the segment data divided by the cache read width supported by the CPU core. When the vector read address calculation mode is the stride address calculation mode or the index address calculation mode, determine whether the minimum value among the product of the number of metadata included in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the second quantity is the product of the vector register width supported by the CPU core, the register set length, and the number of metadata included in the segment data divided by the minimum value; otherwise, the second quantity is the product of the vector register width supported by the CPU core, the register set length, and the number of metadata included in the segment data divided by the minimum value and then plus the product of the register set length and the number of metadata included in the segment data minus 1. By the above method, the number of memory read microinstructions to be generated in different address calculation modes can be determined.
[0091] In some embodiments of the present specification, determining a third quantity according to the register set length and the number of metadata included in the segment data includes: when the segment data read vector instruction is in the unmasked mode, if the product of the number of metadata included in the segment data and the register set length is less than or equal to 4, the third quantity is equal to the number of metadata included in the segment data; otherwise, the third quantity is equal to twice the number of metadata included in the segment data; when the segment data read vector instruction is in the masked mode, if the product of the number of metadata included in the segment data and the register set length is less than or equal to 3, the third quantity is equal to the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register set length is greater than 3 and less than or equal to 6, the third quantity is equal to twice the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register set length is greater than 6, the third quantity is equal to three times the number of metadata included in the segment data. By the above method, the number of data movement microinstructions to be generated in different address calculation modes can be determined.
[0092] In some embodiments of this specification, the data movement micro-instruction includes the position of the first metadata to be copied in the second temporary architecture register and the position of the first metadata in the target register. The data movement micro-instruction may further include the position of the first metadata to be copied in the second temporary architecture register and the position of the first metadata in the target register. Based on the above position parameters, the data movement micro-instruction can read the metadata from the second temporary architecture register and store it in the target register.
[0093] In some embodiments of this specification, the segment data read vector instruction is decoded based on the multiple parameters to obtain multiple micro-instructions, further including: when the segment data read vector instruction is in the mask mode, generating a fourth quantity of vector movement micro-instructions, where the fourth quantity is the product of the register bank length and the number of metadata included in the segment data; the vector movement micro-instruction is used to copy the unmasked metadata in the source register to the target register.
[0094] In this embodiment, since the RISC-V vector instruction supports masked calculation. Therefore, when decoding the segment data read vector instruction, when the segment data read vector instruction is in the mask mode, a fourth quantity of vector movement micro-instructions can also be generated, where the fourth quantity is the product of the register bank length and the number of metadata included in the segment data. The vector movement micro-instruction is used to copy the unmasked metadata in the source register to the target register. That is, the masked segment load instruction generates VMUL×NFIELD VMV micro-instructions (vector movement micro-instructions) to copy the metadata content with mask = 0 in the old target register to the new target register. The Maskless segment load instruction does not need to generate VMV micro-instructions. In the above manner, the masked segment load instruction needs to copy the unmasked metadata from the old target physical register A to the new target physical register mapped by the same architecture register to ensure that the values of these metadata do not change.
[0095] In some embodiments of this specification, after decoding the segment data read vector instruction based on the multiple parameters to obtain multiple microinstructions, the following steps are further included: mapping the first temporary architecture register and the second temporary architecture register to physical registers; allocating physical registers to the target registers corresponding to the integer value read microinstruction, the memory read microinstruction, and the data movement microinstruction; the source register of the vector movement microinstruction uses the physical register initially mapped by the target architecture register, and the target register of the vector movement microinstruction uses the physical register allocated to the data movement microinstruction; distributing the integer value read microinstruction, the memory read microinstruction, the data movement microinstruction, and the vector movement microinstruction; the memory read microinstruction has a data dependency on the integer value read microinstruction; the data movement microinstruction has a data dependency on the memory read microinstruction, and the vector movement microinstruction has no dependency on the integer value read microinstruction, the memory read microinstruction, and the data movement microinstruction.
[0096] In this embodiment, after decoding to obtain multiple microinstructions, register renaming can be performed, that is, mapping an architecture register to a physical register in the kernel to avoid situations such as write-after-read and write-after-write false register dependencies, so as to improve the instruction level parallelism (ILP) and performance of the CPU. For CPU cores that implement register renaming in this way, the masked instruction needs to copy the unmasked metadata from the old target physical register A to the new target physical register B mapped by the same architecture register to ensure that the values of these metadata do not change. Specifically, at this stage, architecture registers, including TAR, will be mapped to physical registers. The target registers of the integer value read microinstruction, the memory read microinstruction, and the data movement microinstruction will be allocated new physical registers. It should be noted that the target register of the VMV microinstruction does not need to be mapped to a new physical register: its source register uses the physical register mapped by the old (i.e., before this segment load instruction) target architecture register, and its target register uses the physical register newly mapped by this instruction VCOMPRESS.
[0097] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. Specifically, reference can be made to the description of the related processing related embodiments above, and details will not be repeated here.
[0098] The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0099] The above method will be described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating this specification and does not constitute an improper limitation of this specification.
[0100] A method for processing RISC-V segment data read vector instructions is proposed in this specific embodiment. This solution uses temporary architecture registers to temporarily store the segment data obtained from memory or cache, and then uses a series of microinstructions to copy the data from the TAR to the target register. Since this solution replaces the memory read microinstructions with register copy microinstructions inside the vector execution unit, both performance and power consumption are improved. The difference between the TAR and normal architecture registers is that the TAR is only used within one instruction to temporarily store and transfer data between its multiple microinstructions, while normal architecture registers need to transfer data between multiple instructions. This solution sets 10 TAR registers, where TAR0 and TAR1 store the integer values required for address generation, and the other 8 TARs (TAR2 to TAR9) support the segment load instruction when LMUL×NFIELD is at most 8.
[0101] Since a RISC-V instruction needs to read four register values (two source registers, mask, and the initial value of the target register), we can reasonably assume that in a CPU core, each RISC-V vector instruction execution unit is allocated 4 vector register file read ports (that is, 4 different register values can be read from the vector register file within one clock cycle).
[0102] In the address calculation of the vector load instruction, both the base (address base value) and the offset (address offset value) of the stride load come from integer registers. High-performance CPU cores generally have separate physical register files for integer and vector registers. When a vector instruction uses an integer register, corresponding I2V microinstructions will be generated to read the corresponding source register value from the integer register file and write it into the TAR target vector register. Other microinstructions in the same instruction obtain the integer value from this TAR. Since the RISC-V vector load instruction uses at most two integer registers, this solution reserves two TAR registers to temporarily store the integer values.
[0103] Figure 4 An example diagram of the RISC-V segment data read vector instruction processing method in this specific embodiment is shown. As Figure 4 shown, it is a schematic diagram of the implementation of a maskless segment strideload with VLEN = 128b, NFIELD = 3, LMUL = 1, and SEW = 16b. In this solution, the instruction is decomposed into 10 (instead of 24 in the prior art solution) VLD data read micro-instructions, and the data in the cache / memory is saved in three TARs. Then, three VCOMPRESS micro-instructions write the metadata from the TARs to three destination registers at the required positions.
[0104] The method in this specific embodiment includes the following steps.
[0105] Decoding: According to the LMUL, NFIELD, and SEW values of a segment load instruction and the vector register width VLEN supported by the CPU core, the instruction is decoded into multiple VLD, VCOMPRESS, and VMV micro-instructions. The cache read width supported by the CPU core is DLEN, and generally DLEN <= VLEN.
[0106] Integer value read micro-instruction I2V: This micro-instruction reads an integer architecture register and writes the result into a vector TAR. Unit-stride and Indexed load instructions only generate one I2V to read the base value and put it into TAR0. Stride load instructions need to generate two I2Vs to read the base and offset values and put them into TAR0 and TAR1.
[0107] Memory read micro-instruction VLD: Similar to the implementation of most CPU cores, this solution requires that one VLD micro-instruction has only one target TAR register. This micro-instruction reads data from memory / cache and writes it into the entire TAR register or a part of the TAR register. In the decoding stage, some of the partially written VLDs need to set the start offset (offset) and data width (size) of the written data in the register.
[0108] For the Unit-stride mode, (VLEN × LMUL × NFIELD / DLEN) VLD micro-instructions are generated. Each micro-instruction reads data with a width of DLEN from memory or cache and writes it into a TAR register. The Masked / maskless segment unit-stride load is the same in this step, and the mask value does not need to be considered.
[0109] For the Stride / Indexed mode, each VLD reads up to min(DLEN, NFIELD × SEW) bits of data from memory or cache. If min(DLEN, NFIELD × SEW) is a power of 2, (VLEN × NFIELD × LMUL) / min(NFIELD × SEW, DLEN) VLD microinstructions are generated. If min(DLEN, NFIELD × SEW) is not a power of 2, (VLEN × NFIELD × LMUL) / min(NFIELD × SEW, DLEN) + LMUL × (NFIELD - 1) VLD microinstructions are generated. The additional LMUL × (NFIELD - 1) microinstructions are due to some segment data (such as [C2, B2, A2] in the figure above) needing to be stored in two registers and thus needing to be split into two VLDs.
[0110] Data movement microinstruction VCOMPRESS: This microinstruction copies metadata from the TAR where all data is already ready to the destination register according to the segment index generated by the decoder.
[0111] For the maskless mode, since there is no need to occupy a register file read port to read the mask, one maskless VCOMPRESS can read 4 TAR registers. For LMUL × NFIELD <= 4, only one VCOMPRESS is required for each destination register. If LMUL × NFIELD > 4, then two VCOMPRESSs are required for each destination register to process the corresponding segment metadata in TAR[0, 1, 2, 3] and TAR[4, 5, 6, 7] respectively.
[0112] For the masked mode, in the masked mode, one source register needs to be reserved for the mask. Therefore, each masked VCOMPRESS can read at most 3 TARs. For LMUL × NFIELD <= 3, only one masked VCOMPRESS is required for each destination register. When 3 < LMUL × NFIELD <= 6, two masked VCOMPRESSs are required for each destination register. When LMUL × NFIELD > 6, three masked VCOMPRESSs are required for each destination register.
[0113] Metadata at position offset in a number of consecutive TARs and metadata at offset+NFIELD, offset+2×NFIELD, …, offset+N×NFIELD will be copied to consecutive positions in the destination register. Therefore, in VCOMPRESS, only the position of the first metadata to be copied in the TAR (start TAR offset) and its position in the destination register (start Dest offset) need to be provided to copy this metadata and multiple subsequent metadata in the same destination register from the given 3 or 4 TAR registers to the destination register.
[0114] In the mask mode, before copying to the destination register, it is necessary to determine whether the mask corresponding to each metadata is 0. If it is 0, the metadata will not be copied into the destination register.
[0115] The masked segment load instruction generates VMUL×NFIELD VMV microinstructions to copy the metadata content with mask=0 in the old destination register to the new destination register. The Maskless segment load instruction does not need to generate VMV microinstructions.
[0116] In the register renaming stage, architectural registers including TAR are mapped to physical registers. The destination registers of I2V / VLD / VCOMPRESS microinstructions are assigned new physical registers. It should be noted that the destination registers of VMV microinstructions do not need to be mapped to new physical registers: their source registers use the old (i.e., before this segment load) mapped physical registers of the destination architectural registers, and their destination registers use the newly mapped physical registers of VCOMPRESS in this instruction.
[0117] Similar to ordinary instructions, when the source registers and other hardware resources (such as execution units and load queues) are ready, microinstructions can be dispatched and executed. Among the microinstructions generated by decoding the same segment load, VLD has a data dependency on I2V, and VCOMPRESS has a data dependency on VLD. While VMV has no data dependency on VCOMPRESS and VLD, so VMV may be executed earlier than VLD / VCOMPRESS of the same instruction.
[0118] During the micro-instruction execution phase, I2V: copies the base / offset value from the integer physical register to TAR0 / TAR1. VLD is the same as the ordinary data load instruction. VLD needs to go through steps such as address generation, address translation (virtual address conversion to physical address), and data cache access. VLD is no different from the ordinary data load operation, and the implementation details are omitted here. For the unmasked mode, VCOMPRESS copies the segment metadata from the specified up to 4 source registers according to the segment index and arranges them in sequence in the target register. For the masked mode, VCOMPRESS copies the segment metadata from the specified up to 3 source registers according to the segment index and arranges them in sequence in the target register. When the mask bit of a certain metadata is 0, this metadata is not copied to the target register. The VMV instruction is only generated in the masked mode. Each VMV micro-instruction reads two source registers: one is the mask, and the other is the old physical register corresponding to the target register. Contrary to the use of the mask by other masked micro-instructions, when the mask corresponding to a metadata is 0, VMV copies this metadata to the target register.
[0119] The retirement of the micro-instructions for segment load is not very different from that of other micro-instructions, and will not be elaborated here.
[0120] Please refer to Figure 5 , which shows the data dependencies between the micro-instructions generated by a segment load instruction, and how this data dependency is established through the TAR register between them. Due to these dependencies, a VLD must be executed after the I2V it depends on has been executed, and a VCOMPRESS must be executed after the VLD it depends on has been executed. And a VMV does not depend on other micro-instructions of the same instruction. Figure 5 TAR[0,1] in is the first temporary architecture register TAR0 and TAR1. TAR[2-9] are the second temporary architecture registers TAR2 to TAR9.
[0121] Since the metadata size in the vector register can vary, ranging from 1 byte to 8 bytes. At the same time, different metadata in the same register may be set by different microinstructions. And before a microinstruction is executed, it must be ensured that every piece of metadata in each source register it uses has been obtained (ready). Therefore, in addition to the data with VLEN bits in each vector register file entry, there are VLEN / 8 valid bits, and each valid bit corresponds to one byte of data in that entry. For a masked segment load instruction, its VCOMPRESS microinstruction will write the masked read metadata into the target physical register and set the corresponding valid bits, while the VMV microinstruction will copy the unmasked metadata from the old physical register corresponding to the target register to the newly mapped physical register of the target register and set the corresponding valid bits. In this way, each piece of metadata in the segment load target register is provided with data by exactly one VCOMPRESS or VMV microinstruction.
[0122] In addition, in the register renaming unit, since 10 new temporary architectural registers (TARs) are added, the corresponding architectural / predicted register mapping table also needs to add 10 entries to include these TARs.
[0123] For the RISC-V vector segment load instruction, this solution proposes a scheme that uses temporary architectural registers to save segment data and the memory access is at the granularity of larger segment data (instead of multiple pieces of metadata within the segment data). This solution reduces the cache / memory access times at the cost of data transfer between vector registers. And the specific implementation steps of this solution are described in detail.
[0124] In this solution, the release of the physical registers mapped by the temporary architectural registers is the same as that of ordinary architectural registers. When an architectural register mapping is written into the architectural mapping table, the physical register corresponding to the replaced entry is released and can be used for a new mapping. Since the temporary architectural registers only transfer data between multiple microinstructions within one instruction, the physical registers they map can be released when the instruction retires. In this way, the utilization rate of physical registers can be improved. However, optimizing specifically for the TAR register mapping will increase the logical complexity of register renaming.
[0125] This solution reserves 8 TARs to temporarily store segment data. A possible solution may only need to reserve fewer TARs, such as only reserving 4. The advantage of this solution is that the register mapping entries for register renaming can be fewer, but the disadvantage is that when LMUL×NFIELD is greater than the number of reserved TARs, the renaming from one TAR to a physical register needs to be performed multiple times, and it may be necessary to generate more VCOMPRESS microinstructions.
[0126] Based on the same inventive concept, an embodiment of this specification also provides a RISC-V segment data read vector instruction processing device as described in the following embodiments. Since the principle of the RISC-V segment data read vector instruction processing device for solving problems is similar to that of the RISC-V segment data read vector instruction processing method, the implementation of the RISC-V segment data read vector instruction processing device can refer to the implementation of the RISC-V segment data read vector instruction processing method, and the repeated parts will not be elaborated. Figure 6 is a structural block diagram of the RISC-V segment data read vector instruction processing device according to an embodiment of this specification, as Figure 6 shown, including: a reading module 601 and a writing module 602, which will be described below.
[0127] The reading module 601 is configured to read segment data from memory or cache in response to a segment data read vector instruction, and temporarily store the read segment data in a temporary architecture register; the segment data includes multiple metadata; the temporary architecture register is an architecture register defined in the implementation of the CPU core, and is only used to temporarily store and transfer data within one instruction.
[0128] The writing module 602 is configured to write the multiple metadata in the segment data in the temporary architecture register into multiple target registers respectively.
[0129] In some embodiments of the present specification, the device further includes a decoding module, and the decoding module is specifically configured to: obtain multiple parameters of a segment data read vector instruction; the multiple parameters include: a vector read address calculation mode, a register bank length, the number of metadata included in the segment data, a metadata width, a vector register width supported by the CPU core, and a cache read width supported by the CPU core; decode the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions; the multiple micro-instructions include: an integer value read micro-instruction, a memory read micro-instruction, and a data movement micro-instruction; correspondingly, the read module is specifically configured to: execute the integer value read micro-instruction to read an address base value or an address base value and an offset value from an integer architecture register and store them in a corresponding first temporary architecture register; execute the memory read micro-instruction to read the address base value or the address base value and the offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read segment data from the memory or the cache based on the address data and write it into a corresponding second temporary architecture register; correspondingly, the write module is specifically configured to: execute the data movement micro-instruction to copy multiple metadata in the segment data in the second temporary architecture register to multiple target registers respectively.
[0130] In some embodiments of the present specification, the decoding module is specifically configured to: determine a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read micro-instructions; determine a second quantity based on the vector read address calculation mode, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core of the segment data read vector instruction to generate a second quantity of memory read micro-instructions; determine a third quantity according to the register bank length and the number of metadata included in the segment data to generate a third quantity of data movement micro-instructions.
[0131] In some embodiments of the present specification, the vector read address calculation mode includes one of the following: a unit stride address calculation mode, a stride address calculation mode, and an index address calculation mode; correspondingly, determining a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read micro-instructions includes: in the case where the vector read address calculation mode is the unit stride address calculation mode or the index address calculation mode, generating one integer value read micro-instruction to read the address base value and store it in one first temporary architecture register; in the case where the vector read address calculation mode is the stride address calculation mode, generating two integer value read micro-instructions to respectively read the address base value and the offset value and store them in two first temporary architecture registers respectively.
[0132] In some embodiments of the present specification, the vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; correspondingly, determining a second quantity based on the vector read address calculation mode of the segment data read vector instruction, the register bank length, the number of metadata included in the segment data, the metadata width, the vector register width supported by the CPU core, and the cache read width supported by the CPU core includes: when the vector read address calculation mode is the unit stride address calculation mode, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the cache read width supported by the CPU core; when the vector read address calculation mode is the stride address calculation mode or the index address calculation mode, determine whether the minimum value among the product of the number of metadata included in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value; otherwise, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value and then plus the product of the register bank length and the number of metadata included in the segment data minus 1.
[0133] In some embodiments of the present specification, determining a third quantity according to the register bank length and the number of metadata included in the segment data includes: when the segment data read vector instruction is in the unmasked mode, if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 4, the third quantity is equal to the number of metadata included in the segment data; otherwise, the third quantity is equal to twice the number of metadata included in the segment data; when the segment data read vector instruction is in the masked mode, if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 3, the third quantity is equal to the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register bank length is greater than 3 and less than or equal to 6, the third quantity is equal to twice the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register bank length is greater than 6, the third quantity is equal to three times the number of metadata included in the segment data.
[0134] In some embodiments of the present specification, the data movement micro-instruction includes the position of the first metadata to be copied in the second temporary architecture register and the position of the first metadata in the target register.
[0135] In some embodiments of the present specification, based on the multiple parameters, the segment data read vector instruction is decoded to obtain multiple micro-instructions, further including: when the segment data read vector instruction is in the mask mode, generating a fourth quantity of vector movement micro-instructions, where the fourth quantity is the product of the register bank length and the number of metadata included in the segment data; the vector movement micro-instruction is used to copy the unmasked calculated metadata in the source register to the target register.
[0136] In some embodiments of the present specification, the device further includes a distribution module, and the distribution module is specifically configured to: after the decoding module decodes the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions, map the first temporary architecture register and the second temporary architecture register to physical registers; allocate physical registers to the target registers corresponding to the integer value read micro-instruction, the memory read micro-instruction, and the data movement micro-instruction; the source register of the vector movement micro-instruction uses the initially mapped physical register of the target architecture register, and the target register of the vector movement micro-instruction uses the physical register allocated to the data movement micro-instruction; distribute the integer value read micro-instruction, the memory read micro-instruction, the data movement micro-instruction, and the vector movement micro-instruction; the memory read micro-instruction has a data dependency on the integer value read micro-instruction; the data movement micro-instruction has a data dependency on the memory read micro-instruction, and the vector movement micro-instruction has no dependency on the integer value read micro-instruction, the memory read micro-instruction, and the data movement micro-instruction.
[0137] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present specification can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present specification are not limited to any specific combination of hardware and software.
[0138] It should be understood that the above description is for illustrative purposes and not for limitation. Upon reading the above description, many embodiments and many applications other than the provided examples will be apparent to those skilled in the art. Therefore, the scope of this specification should not be determined with reference to the above description, but should be determined with reference to the full scope of the foregoing claims and the equivalents thereof.
[0139] The above are only the preferred embodiments of this specification and are not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
Claims
1. A method for processing RISC-V segment data reading vector instructions, characterized in that, Including: In response to a segment data read vector instruction, read segment data from memory or cache and temporarily store the read segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in the implementation of the CPU core and is only used to temporarily store and transfer data within one instruction; Write the plurality of metadata in the segment data in the temporary architecture register into a plurality of target registers respectively; Wherein, the method further includes: Obtain a plurality of parameters of the segment data read vector instruction; the plurality of parameters include: vector read address calculation mode, register group length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core; Decode the segment data read vector instruction based on the plurality of parameters to obtain a plurality of microinstructions; the plurality of microinstructions include: integer value read microinstructions, memory read microinstructions, and data movement microinstructions; Correspondingly, in response to a segment data read vector instruction, read segment data from memory or cache and temporarily store the read segment data in a temporary architecture register, including: Execute the integer value read microinstruction to read an address base value or an address base value and an offset value from an integer architecture register and store them in a corresponding first temporary architecture register; Execute the memory read microinstruction to read an address base value or an address base value and an offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read segment data from memory or cache based on the address data and write it into a corresponding second temporary architecture register; Correspondingly, writing the plurality of metadata in the segment data in the temporary architecture register into a plurality of target registers respectively includes: Execute the data movement microinstruction to copy the plurality of metadata in the segment data in the second temporary architecture register into a plurality of target registers respectively.
2. The RISC-V segment data reading vector instruction processing method according to claim 1, characterized in that, Decoding the segment data read vector instruction based on the plurality of parameters to obtain a plurality of microinstructions includes: Determine a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read microinstructions; Determine a second quantity based on the vector read address calculation mode, register group length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core of the segment data read vector instruction to generate a second quantity of memory read microinstructions; Determine a third quantity according to the register group length and the number of metadata included in the segment data to generate a third quantity of data movement microinstructions.
3. The RISC-V segment data reading vector instruction processing method according to claim 2, wherein The vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; Correspondingly, determining a first quantity according to the vector read address calculation mode of the segment data read vector instruction to generate a first quantity of integer value read microinstructions includes: When the vector read address calculation mode is the unit stride address calculation mode or the index address calculation mode, an integer value read micro-instruction is generated to read the address base value and store it in a first temporary architecture register; When the vector read address calculation mode is the stride address calculation mode, two integer value read micro-instructions are generated to read the address base value and the offset value respectively and store them in two first temporary architecture registers.
4. The RISC-V segment data reading vector instruction processing method according to claim 2, characterized in that, The vector read address calculation mode includes one of the following: unit stride address calculation mode, stride address calculation mode, and index address calculation mode; Correspondingly, determining a second quantity based on the vector read address calculation mode, register bank length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core includes: When the vector read address calculation mode is the unit stride address calculation mode, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the cache read width supported by the CPU core; When the vector read address calculation mode is the stride address calculation mode or the index address calculation mode, determine whether the minimum value among the product of the number of metadata included in the segment data and the metadata width and the cache read width supported by the CPU core is a power of 2. If so, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value; otherwise, the second quantity is the product of the vector register width supported by the CPU core, the register bank length, and the number of metadata included in the segment data divided by the minimum value and then plus the product of the register bank length and the number of metadata included in the segment data minus 1.
5. The RISC-V segment data reading vector instruction processing method according to claim 2, wherein Determining a third quantity according to the register bank length and the number of metadata included in the segment data includes: When the segment data read vector instruction is in the unmasked mode, if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 4, the third quantity is equal to the number of metadata included in the segment data; otherwise, the third quantity is equal to twice the number of metadata included in the segment data; When the segment data read vector instruction is in the masked mode, if the product of the number of metadata included in the segment data and the register bank length is less than or equal to 3, the third quantity is equal to the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register bank length is greater than 3 and less than or equal to 6, the third quantity is equal to twice the number of metadata included in the segment data; if the product of the number of metadata included in the segment data and the register bank length is greater than 6, the third quantity is equal to three times the number of metadata included in the segment data.
6. The RISC-V segment data reading vector instruction processing method according to claim 5, wherein, The data movement micro-instruction includes the position of the first metadata to be copied in the second temporary architecture register and the position of the first metadata in the target register.
7. The RISC-V segment data reading vector instruction processing method according to claim 2, characterized in that Decoding the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions, further including: When the segment data read vector instruction is in the mask mode, generating a fourth quantity of vector movement micro-instructions, where the fourth quantity is the product of the register bank length and the number of metadata included in the segment data; the vector movement micro-instruction is used to copy the unmasked calculated metadata in the source register to the target register.
8. The RISC-V segment data reading vector instruction processing method according to claim 7, wherein After decoding the segment data read vector instruction based on the multiple parameters to obtain multiple micro-instructions, further including: Mapping the first temporary architecture register and the second temporary architecture register to physical registers; Allocating physical registers to the target registers corresponding to the integer value read micro-instruction, the memory read micro-instruction, and the data movement micro-instruction; the source register of the vector movement micro-instruction uses the initially mapped physical register of the target architecture register, and the target register of the vector movement micro-instruction uses the physical register allocated to the data movement micro-instruction; Distributing the integer value read micro-instruction, the memory read micro-instruction, the data movement micro-instruction, and the vector movement micro-instruction; the memory read micro-instruction has a data dependency on the integer value read micro-instruction; the data movement micro-instruction has a data dependency on the memory read micro-instruction, and the vector movement micro-instruction has no dependency on the integer value read micro-instruction, the memory read micro-instruction, and the data movement micro-instruction.
9. A RISC-V segment data reading vector instruction processing device, characterized in that, Including: A reading module, configured to read segment data from a memory or a cache in response to a segment data read vector instruction, and temporarily store the read segment data in a temporary architecture register; the segment data includes a plurality of metadata; the temporary architecture register is an architecture register defined in the CPU core implementation and is only used to temporarily store and transfer data within an instruction; A writing module, configured to write the plurality of metadata in the segment data in the temporary architecture register into a plurality of target registers respectively; Among them, the device further includes a decoding module, and the decoding module is specifically configured to: obtain multiple parameters of the segment data read vector instruction; the multiple parameters include: vector read address calculation mode, register group length, number of metadata included in the segment data, metadata width, vector register width supported by the CPU core, and cache read width supported by the CPU core; decode the segment data read vector instruction based on the multiple parameters to obtain multiple microinstructions; the multiple microinstructions include: integer value read microinstruction, memory read microinstruction, and data movement microinstruction; correspondingly, the reading module is specifically configured to: execute the integer value read microinstruction to read an address base value or an address base value and an offset value from an integer architecture register and store them in a corresponding first temporary architecture register; execute the memory read microinstruction to read an address base value or an address base value and an offset value from the first temporary architecture register, generate address data according to the address base value or the address base value and the offset value, and read segment data from memory or cache based on the address data and write it into a corresponding second temporary architecture register; correspondingly, the writing module is specifically configured to: execute the data movement microinstruction to copy multiple metadata in the segment data in the second temporary architecture register to multiple target registers respectively.
Citation Information
Patent Citations
Segmented data storage method and device, storage medium and electronic device
CN114461599A
Cache data segmentation reading method and device, computer equipment and storage medium
CN117112454A