Instruction processing equipment, system and processing method based on RISC-V architecture

By introducing an extended instruction interface encoding unit into the instruction processing device of the RISC-V architecture, the problem that "V" extended instructions in the prior art cannot meet the needs of special application scenarios is solved, and higher flexibility and computing performance are achieved.

CN120029671AActive Publication Date: 2025-05-23芯来智融半导体科技(上海)股份有限公司

Patent Information

Application Number
CN202510510331.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The "V" extension instructions of the existing RISC-V architecture cannot meet the needs of some special application scenarios, or the efficiency of using basic instructions in some scenarios is low, affecting the computing performance of the processor.

Method used

An instruction processing device based on RISC-V architecture is designed, including a processor core and an instruction co-processing module. The processor core is composed of a decoding module, a distribution module and an operation module. An extended instruction interface encoding unit is set up in the decoding module, which allows users to design instructions of different needs according to different scenarios.

Benefits of technology

Through this design, the device can meet changing business needs, improve processor flexibility, reduce instruction processing delays, ensure effective execution of vector processing operations, and accelerate computing performance by dynamically adjusting processing strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029671A_ABST
    Figure CN120029671A_ABST
Patent Text Reader

Abstract

The invention provides an instruction processing device, system and processing method based on an RISC-V. The device comprises a processor core and an instruction co-processing module, and the processor core comprises a decoding module, a distribution module and an operation module which are connected in sequence; the decoding module comprises an extension instruction interface coding unit; the decoding module is used for integrating the extension interface signals output by the extension instruction interface coding unit into an integration result and sending the integration result to the distribution module; the expansion interface signal is a pre-configured expandable interface signal; the distribution module is used for distributing the instruction to the operation module according to the integration result; an instruction control unit in the operation module is used for receiving the instruction distributed by the distribution module, executing vector processing operation to generate an operation result and sending the operation result to the instruction co-processing module; and the instruction co-processing module is used for interacting with the instruction control unit and executing corresponding processing operation. According to the scheme, constantly changing service requirements are met, and the equipment has higher flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of processor technology, and in particular, to an instruction processing device, system and processing method based on RISC-V architecture. Background Art

[0002] With the rapid development of modern computing systems, RISC-V, as an open source instruction architecture, has modular and extensible features, allowing developers to add different extended instruction sets according to their own needs. Among them, the "V" extension is the officially defined vector extended instruction set in the RISC-V architecture, which provides the RISC-V processor with the basic ability of vector operations. This extended instruction set covers a series of basic vector operation instructions, such as vector loading, storage, arithmetic operations, logical operations, etc., which can be used in standard server and application processor platform profiles. Other platforms, including embedded platforms, can also choose to implement subsets of these extensions to meet different application requirements.

[0003] At present, the official standard "V" extended instructions in the relevant technology cannot meet the needs of some special application scenarios, or the efficiency of using basic instructions in some scenarios is low, which affects the computing performance of the processor. Summary of the invention

[0004] The embodiments of the present application provide an instruction processing device, system and processing method based on the RISC-V architecture.

[0005] In a first aspect of an embodiment of the present application, an instruction processing device based on a RISC-V architecture is provided, the instruction processing device comprising: a processor core and an instruction co-processing module, the processor core comprising: a decoding module, a distribution module and a computing module connected in sequence; The decoding module includes an extended instruction interface encoding unit; the decoding module is used to integrate the extended interface signal output by the extended instruction interface encoding unit to obtain an integration result and send it to the distribution module; the extended interface signal is a pre-configured and extensible interface signal; The distribution module is used to: distribute instructions to the operation module according to the integration result; The operation module includes an instruction control unit, which is used to: receive the instruction distributed by the distribution module, control the operation module to perform vector processing operations to generate operation results and send the operation results to the instruction co-processing module; The instruction co-processing module is used to interact with the instruction control unit and perform corresponding processing operations according to the calculation results.

[0006] A second aspect of the embodiments of the present application provides an instruction processing system, including the instruction processing device provided in the above embodiments.

[0007] A third aspect of the embodiments of the present application provides an instruction processing method based on the RISC-V architecture, the method comprising: The extended interface signal output by the extended instruction interface encoding unit is integrated to obtain an integrated result; the extended interface signal is a pre-configured expandable interface signal; the integrated result includes instructions; Distributing instructions to the computing modules according to the integration result; According to the distributed instructions, execute vector processing operations to generate operation results; A corresponding processing operation is performed according to the calculation result.

[0008] In the embodiment of the present application, an instruction processing device, system and processing method based on RISC-V architecture are provided, and the device includes: a processor core and an instruction co-processing module, and the processor core includes: a decoding module, a distribution module and an operation module connected at one time. The decoding module includes an extended instruction interface encoding unit, and the decoding module is used to integrate the extended interface signal output by the extended instruction interface encoding unit to obtain an integrated result and send it to the distribution module; the extended interface signal is a pre-configured extensible interface signal; the distribution module is used to: distribute the instruction to the operation module according to the integration result; the operation module includes an instruction control unit, and the instruction control unit is used to: receive the instruction distributed by the distribution module, control the operation module to perform vector processing operations to generate operation results and send the operation results to the instruction co-processing module; the instruction co-processing module is used to: interact with the instruction control unit and perform corresponding processing operations according to the operation results. Compared with the prior art, the technical solution in the present application sets up an extended instruction interface encoding unit in the decoding module, so that users can flexibly design instructions with different needs according to different scenarios, meet the ever-changing business needs, and make the device more flexible. In addition, through the orderly collaboration between the decoding module, the distribution module and the operation module, the delay of instruction processing is reduced, and the effective execution of vector processing operations is guaranteed. Through the interaction between the instruction co-processing module and the instruction control unit, the system can flexibly process different instructions. This interactive mechanism can dynamically adjust the processing strategy according to the specific instruction type and business needs, thereby accelerating the computing performance of the processing device. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1A schematic diagram of the structure of a computer device provided in one embodiment of the present application; Figure 2 A schematic diagram of the structure of an instruction processing device provided by an embodiment of the present application; Figure 3 A flowchart of an instruction processing method provided by an embodiment of the present application; Figure 4 A schematic diagram of the structure of an instruction processing device provided for another embodiment of the present application.

[0010] Description of reference numerals: Processor core-10; decoding module-11; distribution module-12; operation module-13; instruction co-processing module-20; extended instruction interface encoding unit-111; instruction control unit-131. DETAILED DESCRIPTION

[0011] In the process of implementing the present application, the inventors discovered that the traditional standard "V" extended instructions cannot meet the needs of some special application scenarios, or the efficiency of using basic instructions in some scenarios is low, thereby affecting the computing performance of the processor.

[0012] In order to make the technical solutions and advantages in the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than an exhaustive list of all the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0013] Based on the above-mentioned defects, the present application provides an instruction processing device based on the RISC-V architecture. Compared with the related art, the technical solution in the present application sets an extended instruction interface encoding unit in the decoding module, so that users can flexibly design instructions with different needs according to different scenarios, meet the ever-changing business needs, and make the device more flexible. In addition, through the orderly collaboration between the decoding module, the distribution module and the operation module, the delay of instruction processing is reduced, and the effective execution of vector processing operations is ensured. Through the interaction between the instruction co-processing module and the instruction control unit, the system can flexibly process different instructions. This interaction mechanism can dynamically adjust the processing strategy according to the specific instruction type and business needs, thereby accelerating the computing performance of the processing device.

[0014] See also Figure 1 , a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 1As shown, the computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium can be, for example, a disk. The non-volatile storage medium stores files (which can be files to be processed or processed files), an operating system and a computer program, etc. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an instruction processing method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0015] See also Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of an instruction processing device based on the RISC-V architecture provided in an embodiment of the present application. Figure 2 As shown, the instruction processing device includes: a processor core 10 and an instruction co-processing module 20, and the processor core 10 includes: a decoding module 11, a distribution module 12 and a calculation module 13 connected in sequence.

[0016] The decoding module 11 includes an extended instruction interface encoding unit 111; the decoding module 11 is used to integrate the extended interface signal output by the extended instruction interface encoding unit to obtain an integrated result and send it to the distribution module; the extended interface signal is a pre-configured extensible interface signal; the distribution module 12 is used to: distribute the instruction to the operation module according to the integration result; the operation module 13 includes an instruction control unit 131, and the instruction control unit 131 is used to: receive the instruction distributed by the distribution module, control the operation module to perform vector processing operations to generate operation results and send the operation results to the instruction co-processing module 20; the instruction co-processing module 20 is used to: interact with the instruction control unit and perform corresponding processing operations according to the operation results.

[0017] It should be noted that the processor core (CORE) is the core part of the processor, which is the key component of the entire processing system and is responsible for coordinating and controlling the execution of instructions. The decoding module (CORE_DECODE) is a module within the processor core and is responsible for decoding input instructions and identifying the type and operation content of the instructions.

[0018] The decoding module may include an extended instruction interface encoding unit (VNICE_DECODE), which is a custom decoding unit. Users can use this unit to perform custom instruction encoding, such as custom encoding according to different application scenarios or different requirements. It is a key part for implementing flexible custom extended instructions.

[0019] The distribution module is used to receive the integration results of the extended interface signals from the decoding module in the processor core. These integration results contain the processed instruction information and are the basis for the distribution module to perform subsequent operations. Then, according to the received integration results, the instructions are analyzed and judged, and the instructions are accurately distributed to the operation module. It plays the role of an instruction "dispatcher", ensuring that each instruction can be correctly sent to the appropriate operation module for processing, ensuring the orderly progress of the instruction processing flow.

[0020] The distribution module accurately distributes instructions, establishes a communication bridge between the decoding module and the operation module, coordinates the work between the two modules, improves the operating efficiency and coordination of the entire instruction processing device, and avoids confusion and errors in the instruction processing process. The distribution module is the key link connecting the decoding module and the operation module in the instruction processing device, and plays an important role in ensuring that instructions can be processed and executed correctly and efficiently.

[0021] The above-mentioned operation module can be a vector processing unit (VPU), which is mainly used to process vector operation instructions and improve data parallel processing capabilities. The operation module provides operation capabilities for the entire device and performs corresponding operation operations in response to operation instructions. The operation operations may include various types of addition, subtraction, multiplication, division, etc.

[0022] The instruction control unit (VNICE_CTRL) is located inside the operation module and is a control module related to the custom extension. It is used to control the operation module to perform corresponding operations according to the instruction encoding information transmitted by the extended instruction interface encoding unit VNICE_DECODE.

[0023] The instruction co-processing module (VNICE_CORE) is the core part of the custom extended instruction processing. It interacts with the instruction control unit VNICE_CTRL and participates in the specific implementation of the instruction function.

[0024] The operation request signal is used to instruct the execution of an operation, and the operation request signal may include operation data and operation mode.

[0025] Specifically, users can customize instruction encoding according to different application scenarios, obtain a variety of extended interface signals, and pre-configure them in the extended instruction interface encoding unit. The decoding module will integrate the output extended interface signals to obtain an integrated result, which includes the processed instruction information, and send the integrated result to the distribution module. After receiving the integration module, the distribution module analyzes and judges the instructions, and accurately distributes the instructions to the operation module. The operation module includes an instruction control unit, and then the instruction control unit controls the computing resources in the operation module to perform vector processing operations on the instructions, thereby obtaining the operation results. The operation module can send the operation results to the instruction co-processing module outside the processor core, and the instruction co-processing module interacts with the instruction control unit in the operation module and performs corresponding processing operations.

[0026] The above calculation results may be expressed in the form of text, in the form of a table, or in the form of an image. This embodiment does not impose any limitation on the form of expression of the calculation results.

[0027] In an embodiment of the present application, an instruction processing device based on a RISC-V architecture is provided, the device comprising: a processor core and an instruction co-processing module, the processor core comprising: a decoding module, a distribution module and an operation module connected once. The decoding module comprises an extended instruction interface encoding unit, the decoding module is used to integrate the extended interface signal output by the extended instruction interface encoding unit to obtain an integrated result and send it to the distribution module; the extended interface signal is a pre-configured extensible interface signal; the distribution module is used to: distribute the instruction to the operation module according to the integration result; the operation module comprises an instruction control unit, the instruction control unit is used to: receive the instruction distributed by the distribution module, control the operation module to perform vector processing operations to generate operation results and send the operation results to the instruction co-processing module; the instruction co-processing module is used to: interact with the instruction control unit, and perform corresponding processing operations according to the operation results. Compared with the prior art, the technical solution in the present application sets up an extended instruction interface encoding unit in the decoding module, so that users can flexibly design instructions with different needs according to different scenarios, meet the ever-changing business needs, and make the device more flexible. In addition, through the orderly collaboration between the decoding module, the distribution module and the operation module, the delay of instruction processing is reduced, and the effective execution of vector processing operations is guaranteed. Through the interaction between the instruction co-processing module and the instruction control unit, the system can flexibly process different instructions. This interactive mechanism can dynamically adjust the processing strategy according to the specific instruction type and business needs, thereby accelerating the computing performance of the processing device.

[0028] In an optional embodiment of the present application, the extended instruction interface encoding unit is used to: Receive vnice command code input from external module; According to the pre-configured instruction encoding rules, the vnice instruction code is analyzed for characteristics and the extended interface signal is output.

[0029] It is understandable that in the instruction processing device based on RISC-V architecture, the vnice instruction code is an instruction set for implementing a specific function. The external module can be, for example, a program memory, an instruction generator, etc. The external module sends the written vnice instruction code to the extended instruction interface encoding unit in the decoding module. Among them, these vnice instruction codes generally exist in binary encoding form, representing different operations and functions. The above-mentioned vnice instruction code can be customized by the user according to different actual application scenarios and different requirements, and is used to implement vector operations, etc. The user can manipulate these control signal combinations to realize various instruction functions through custom instruction encoding.

[0030] The above instruction encoding rules may be predefined instruction encoding specifications, which may include the format of the instruction code, the meaning of each field and the corresponding operation. Through these rules, the system can parse the instruction code and understand its specific function, thereby outputting the extended interface signal.

[0031] Optionally, the above vnice instruction code may be composed of fields such as an operation code (Opcode), an operand address, and a function code. The operation code is used to specify the basic type of instruction, such as addition, multiplication, etc. The operand address indicates the location of the data involved in the operation; the function code further refines the operation of the instruction, such as whether to use a mask, which rounding mode to use, etc.

[0032] In the process of outputting the extended interface signal, each field of the vnice instruction code may be parsed to determine the characteristic information of the vnice instruction, and then the extended interface signal may be output according to the characteristic information of the vnice instruction. These extended interface signals may interact with the distribution module and the operation module in the device, thereby providing necessary control information for the execution of subsequent instructions.

[0033] The characteristics include, but are not limited to, operation type, operand information, and some special functions. Operation types include arithmetic operations, logical operations, data transfer, and other types; operand information may include: determining the number, type (integer, floating point, etc.) and source (register, memory, etc.) of operands involved in the operation; some special functions may include: identifying whether the instruction has special functions, such as whether to use masks, whether to perform saturation operations, etc.

[0034] The extended instruction interface encoding unit in the embodiment of the present application can receive a customized vnice instruction code, enabling developers to create corresponding exclusive instructions based on specific needs, breaking the limitations of traditional instruction sets, greatly improving the flexibility of the instruction system, and can output extended interface signals to provide detailed instruction information for subsequent modules, making it easier for the distribution module to accurately distribute instructions to the corresponding operation modules, reducing the delay in instruction processing, and greatly improving instruction processing efficiency.

[0035] In an optional embodiment of the present application, the extended interface signal includes: At least one first operand interface signal, at least one second operand interface signal, at least one result register interface signal and at least one other interface signal; the first operand interface signal is used to characterize the enable control, source indication and mutual exclusion constraint of the first operand of the vnice instruction; the second operand interface signal is used to characterize the enable control and source limitation of the second operand of the vnice instruction; the result register interface signal is used to characterize the enable control, update type indication, mutual exclusion constraint and read operation indication of the result register of the vnice instruction; the other interface signals are used to indicate the legality judgment, instruction type identification and mask use identification of the vnice instruction.

[0036] It should be noted that the above-mentioned extended interface signal may include an input signal and an output signal. The input signal refers to the signal transmitted from the external module to the extended instruction interface encoding unit, and the output signal refers to the signal output by the extended instruction interface encoding unit (VNICE_DECODE unit). Please refer to the following Table 1, which shows the content of the extended interface signal (VNICE_DECODE interface signal): Table 1

[0037] The signal names in the above table are the names of the VNICE_DECODE interface signals, which refer to the identification of each extended interface signal. The direction indicates the direction of signal transmission. "Input" means that the external signal is transmitted to the VNICE_DECODE unit, such as "dec_vnice_instr" receiving the instruction code; "Output" means that the module outputs signals to the outside, such as various "dec_vnice_ " signal, used to convey the result of instruction-related characteristic judgment. "Bit width" refers to the number of binary bits occupied by the signal, such as "dec_vnice_instr" has a bit width of 32, which can represent a variety of instruction codes; most output signals have a bit width of 1, which is used to represent the judgment result of "yes" or "no". "Description" is used to characterize the functional meaning of each signal, such as judging the requirements of the vnice instruction for operands (op1, op2), the source of operands, the operation on the result register (rd), the legality of the instruction, whether it is a vnice instruction, whether v0 is used as a mask, etc.

[0038] The first operand interface signals refer to signals related to operand op1, including: dec_vnice_op1_en, dec_vnice_op1_scalar_reg_fpu, dec_vnice_op1_scalar_reg_int, and dec_vnice_op1_vector_reg. The second operand interface signals refer to signals related to operand op2, including: dec_vnice_op2_en. The result register interface signals include: dec_vnice_rd_en, dec_vnice_rd_scalar_reg_fpu, dec_vnice_rd_scalar_reg_int, dec_vnice_rd_vector_reg, and dec_vnice_rd_mac. Other interface signals include: dec_vnice_ilgl, dec_vnice_op, and dec_vnice_vm.

[0039] It should be noted that, for the first operand interface signal, dec_vnice_op1_en, bit width 1, is used to indicate whether the vnice instruction requires operand op1. dec_vnice_op1_scalar_reg_fpu is used to indicate whether the operand op1 of the vnice instruction comes from the scalar register of the floating point unit (FPU). dec_vnice_op1_scalar_reg_int is used to indicate whether op1 comes from the integer register (rs1). dec_vnice_op1_vector_reg is used to indicate whether op1 comes from the vector register of the computing module (VPU).

[0040] The above three signals must be mutually exclusive, that is, at a certain moment, operand op1 can only have one source. For example, when the dec_vnice_op1_scalar_reg_fpu signal is 1 (indicating that op1 comes from the scalar register of the FPU), the dec_vnice_op1_scalar_reg_int and dec_vnice_op1_vector_reg signals must be 0, and multiple signals cannot be 1 at the same time to avoid confusing the data source of op1.

[0041] For the second operand interface signal, dec_vnice_op2_en, its bit width is 1, which is used to indicate whether the vnice instruction requires operand op2, and the source of operand op2 is only of one type, which is a vector register. This is different from operand op1 which has multiple possible sources and is relatively clearer and single.

[0042] For the result register rd interface signal, dec_vnice_rd_en, bit width 1, indicates whether the vnice instruction needs to update rd (result register). dec_vnice_rd_scalar_reg_fpu, bit width 1, indicates whether the rd updated by the vnice instruction is an fpu register. dec_vnice_rd_scalar_reg_int, bit width 1, indicates whether the updated rd is a scalar register (scalar register). dec_vnice_rd_vector_reg, bit width 1, indicates whether the updated rd is a vector register (vector register). The user needs to ensure that these three signals related to rd update are mutually exclusive, that is, rd can only update one type of register. dec_vnice_rd_mac, bit width 1, indicates whether the vnice instruction reads rd.

[0043] Similar to the rule for operand op1, these three signals must also be mutually exclusive. Operand rd can only update one type of register after the instruction is executed. For example, when dec_vnice_rd_scalar_reg_fpu is 1, the other two signals must be 0 to ensure the uniqueness of the update target.

[0044] For other interface signals, dec_vnice_ilgl, bit width 1, indicates whether the vnice instruction is an illegal instruction. dec_vnice_op, bit width 1, indicates whether it is a vnice instruction. dec_vnice_vm, bit width 1, indicates whether the instruction uses v0 as a mask.

[0045] In the embodiment of the present application, by setting at least one first operand interface signal, at least one second operand interface signal, at least one result register interface signal and at least one other interface signal, it is possible to meet the special requirements of different application scenarios for instruction processing and improve the flexibility and accuracy of instruction processing. For example, in image processing applications, it may be necessary to use a mask to selectively process certain parts of an image. This function can be achieved by setting the corresponding mask control interface signal. And it is possible to accurately inform the operation module of the specific operation that needs to be performed to ensure that the instruction is executed accurately. For example, when a vector addition operation is required, the corresponding addition operation interface signal will be activated, and after the operation module receives the signal, it will perform the vector addition operation.

[0046] In an optional embodiment of the present application, the instruction co-processing module interacts with the instruction control unit via a co-processing interface signal; The co-processing interface signal includes at least one of the following: a clock signal, a handshake interaction signal, a transmission flag signal, an instruction and configuration signal, an operand signal, and an output result and flag signal; the clock signal is used to characterize the clock reference for the instruction co-processing module to perform operations; the handshake interaction signal is used to characterize the validity of the input data and the readiness state of the module to receive data, as well as the validity of the output data and the readiness state of the downstream module to receive data; the transmission flag signal is used to characterize the flag information of the data transmission; the operand signal is used to characterize the operand information; the operand information includes at least one of: a vector operand and a scalar operand; the output result and flag signal is used to characterize the operation process attribute information and the operation result information.

[0047] It should be noted that the above-mentioned extended interface signal may include an input signal and an output signal. The input signal refers to the signal transmitted from the external module to the instruction co-processing module (VNICE_CORE module), and the output signal refers to the signal output by the instruction co-processing module. Please refer to the following Table 2, which shows the specific content of the co-processing interface signal (VNICE_IF interface signal): Table 2

[0048] The above directions indicate the direction of signal transmission. "Input" means that the external signal is transmitted to the VNICE_DECODE unit, such as "dec_vnice_instr" receiving the instruction code; "Output" refers to the signal output by the VNICE_CORE module to the outside. "Bit width" refers to the number of binary bits occupied by the signal. "Description" is used to characterize the functional meaning of each co-processing interface signal (VNICE_IF interface signal).

[0049] The clock signal is vnice_clk, which is a 1-bit input signal. As a clock signal, it provides a timing reference for the operation of the entire VNICE-CORE module. All logic operations are synchronized based on this clock signal.

[0050] The handshake interaction signals include: i_vnice_valid, i_vnice_ready, vnice_wbck_valid and o_vnice_wbck_ready. Among them, i_vnice_valid is a 1-bit input signal, which is the valid signal of the vnice input handshake. When the signal is high (logic 1), it indicates that the input data and instructions are valid and can be processed. i_vnice_ready is a 1-bit output signal, which is the ready signal of the vnice input handshake. When the signal is high, it indicates that the VNICE module is ready to receive input data and instructions. Only when i_vnice_valid and i_vnice_ready are both high, the input data and instructions can be successfully transmitted and processed. vnice_wbck_valid is a 1-bit output signal, which is the valid signal of the vnice output handshake. When the signal is high, it indicates that the operation result output by the instruction co-processing module is valid. o_vnice_wbck_ready is a 1-bit input signal, which is the ready signal of the vnice output handshake. When this signal is high, it indicates that the receiving end is ready to receive the operation result output by the instruction co-processing module. Similarly, the output result can only be successfully transmitted when o_vnice_wbck_valid and o_vnice_wbck_ready are both high.

[0051] The transmission flag signal includes: i_vnice_beat. The i_vnice_beat is a 2-bit input signal, which is the vnice start and end flag signal. It indicates the transmission stage according to different VLMUL (vector length multiplier) conditions: when VVLMUL>1, there are multiple VLEN (vector length) transmissions: 01 indicates the first VLMUL transmission; 00 indicates the VLMUL transmission in the middle; 10 indicates the last VLMUL transmission; when VLMUL<=1, it is a single VLEN transmission, and 11 indicates only one VLMUL transmission.

[0052] The instructions and configuration signals include: i_vnice_instr, i_vnice_csr_vlmul, i_vnice_csr_vsew, i_vnice_csr_rounding_mode, i_vnice_fpu_rounding_mode. i_vnice_instr is a 32-bit input signal, which is the vnice instruction code and contains the specific instruction information to be executed. i_vnice_csr_vlmul is a 3-bit input signal, which is the vlmul control status register (CSR) signal of vnice, indicating the currently configured VLMUL mode, which is used to determine the length multiplier of the vector operation. i_vnice_csr_vsew is a 3-bit input signal, which is the vsew control status register (CSR) signal of vnice, indicating the currently configured element width, that is, the bit width of each element in the vector. i_vnice_csr_rounding_mode is a 2-bit input signal, which is the rounding control status register (CSR) signal of vnice, indicating the rounding mode used by the current configuration, used for rounding operations of integer operations. i_vnice_fpu_rounding_mode is a 3-bit input signal, which is the floating-point rounding control status register (CSR) signal of vnice, indicating the floating-point rounding type used by the current configuration, used for rounding operations of floating-point operations.

[0053] Operand signals include: i_vnice_vs1, i_vnice_vs2, i_vnice_rs1, i_vnice_vd. i_vnice_vs1 is a VLEN-bit input signal, which is the vs1 operand of vnice and is used as an input operand of vector operations. i_vnice_vs2 is a VLEN-bit input signal, which is the vs2 operand of vnice and is used as another input operand of vector operations. i_vnice_rs1 is a GLEN-bit input signal, which is the rs1 operand of vnice and is usually used as a scalar operand. i_vnice_vd is a VLEN-bit input signal, which is the vd operand of vnice and is used to store the result of vector operations.

[0054] The output result and flag signals include: o_vnice_wbck_wdat, o_vnice_wbck_vxsat, o_vnice_wbck_fflag and o_vnice_active. Among them, o_vnice_wbck_wdat is the VLEN bit output signal, which is the operation result signal output by vnice and contains the final result of the vector operation.

[0055] o_vnice_wbck_vxsat is a 1-bit output signal, which is the operation saturation flag. When the signal is high, it indicates that saturation has occurred during the operation, that is, the operation result exceeds the representable range. o_vnice_wbck_fflag is a 5-bit output signal, which is a floating point exception flag, used to indicate whether an abnormal situation has occurred during the floating point operation, such as overflow, underflow, division by zero, etc. o_vnice_active is a 1-bit output signal, which is used to control the clock. When vnice is in working state, this flag needs to be pulled up to indicate that the module is performing an operation.

[0056] In the embodiment of the present application, by setting the clock signal, a unified time base can be provided for the entire system, so that each module executes the corresponding operation in order according to the clock signal, and by setting the handshake interaction signal, reliable data transmission is ensured, and the processing speed difference of different modules is used. By setting the transmission flag signal, the data transmission website can be guaranteed, and complex transmission modes are supported. Setting the instruction and configuration signal can realize instruction diversity and scalability, setting the operand signal can provide operation data, support multiple operation types, and setting the output result and flag signal can feedback the operation result to indicate the operation status, thereby meeting the special requirements of instruction processing in different application scenarios and improving the flexibility and accuracy of instruction processing.

[0057] In an optional embodiment of the present application, the above-mentioned operand signal includes: a data element status signal; the data element status signal is used to indicate whether the current data element in the vector operand is in an active state; the value of the data element status signal is determined according to different combinations of vector length, vector length multiplier and data element width; the value includes a set of data element status values ​​for each current bit in the vector operand, and the data element status values ​​include 1 and 0.

[0058] When the data element state value is 1, a group of data elements representing the current position is in an active state; when the data element state value is 0, a group of data elements representing the current position is in an inactive state.

[0059] In this embodiment, the data element status signal is the i_vnice_element_active signal, which is used to indicate which data elements in the vector operand are active, that is, which elements will participate in the specific vector operation. The i_vnice_element_active signal is a signal related to sew (element width), which is used to indicate whether the current data element is active. Specifically, its value varies according to different combinations of vector length (VLEN, Vector Length)), vector length multiplier (Vector Length Multiplier, VLMUL) and element width (Vector Element Stride Width, VSEW).

[0060] Among them, VLEN represents the length of the vector. For example, VLEN=128 means that the vector contains a total of 128 bits of data. VLMUL is the vector length multiplier. Here VLMUL=1 means that the vector length is the standard VLEN value. VSEW refers to the width of the vector element, which determines the number of bits occupied by each element in the vector.

[0061] In this embodiment, each bit of the i_vnice_element_active signal corresponds to a group of elements in the vector. When a bit is 1, it indicates that the corresponding data element is in an active state and will participate in subsequent operations; when a bit is 0, it indicates that the corresponding data element is in an inactive state and does not participate in operations. The bit width of the signal is VLEN / 8, because the signal controls the active state of the element in bytes.

[0062] For example, when VSEW = 8, it indicates that the width of each element is 8 bits, and a 128-bit vector contains 128 / 8 = 16 elements in total. i_vnice_element_active = 16'b1111111111111111 means that all 16 elements are active and will participate in vector operations.

[0063] When VSEW = 16, the width of each element is 16 bits, and a 128-bit vector contains 128 / 16=8 elements. Since i_vnice_element_active is controlled in bytes, a 16-bit wide element corresponds to a 2-bit i_vnice_element_active signal. i_vnice_element_active = 16'b0000000011111111 means that the 4 16-bit elements corresponding to the last 8 bytes are active, while the first 4 16-bit elements do not participate in the operation.

[0064] When VSEW = 32, the width of each element is 32 bits, and a 128-bit vector contains 128 / 32 = 4 elements. The 32-bit wide element corresponds to the 4-bit i_vnice_element_active signal. i_vnice_element_active = 16'b0000000000001111 means that the 1 32-bit element corresponding to the last 4 bytes is active, and the first 3 32-bit elements do not participate in the operation.

[0065] When VSEW=64, it indicates that the width of each element is 64 bits, and a 128-bit vector contains 128 / 64= 2 elements. The 64-bit wide element corresponds to the 8-bit i_vnice_element_active signal. i_vnice_element_active = 16'b0000000000000011 means that the 64-bit element corresponding to the last 2 bytes is active, and the first 64-bit element does not participate in the operation.

[0066] In the embodiment of the present application, by setting the i_vnice_element_active signal, the data elements participating in the operation in the vector can be flexibly controlled according to different element widths, thereby improving the efficiency and flexibility of vector operations.

[0067] In an optional embodiment of the present application, the instruction and configuration signal include a vector length multiplier configuration signal, a vector element width configuration signal, a first mode configuration signal, and a second mode configuration signal; The vector length multiplier configuration signal is used to indicate the configuration of the vector length multiplier; the vector element width configuration signal is used to indicate the configuration of the data element width; the first mode configuration signal is used to indicate the configuration of the rounding mode of the conventional operation; the second mode configuration signal is used to indicate the configuration of the rounding mode of the floating-point operation.

[0068] It should be noted that the above-mentioned vector length multiplier configuration signal can be i_vnice_csr_vlmul, which is used to configure the vector length multiplier (LMUL) related information and determine the proportional relationship between the effective length of the vector operand and the standard vector length, which has an important impact on the amount of vector data involved in the operation and the scale of the operation.

[0069] The signal encoding rule of i_vnice_csr_vlmul complies with the "V" extension standard of RISCV, as shown in Table 3 below: Table 3

[0070] Among them, vmul[2:0] in the above table is a 3-bit binary code, which is used to select different vector operation configuration modes. Different codes correspond to different parameter settings. LMUL is used to represent the vector length multiplier, which represents the ratio of the effective length of the vector operand to the standard vector length (VLEN), such as 1 / 8, 1 / 4, etc. #groups refers to the number of groups, indicating the number of vector register groups. VLMAX represents the maximum number of elements in each group. Its value is calculated based on the ratio of VLEN (vector length) and SEW (vector element width). Different configurations have different calculation methods and values. Registersgrouped with register n: describes the other registers grouped with register n, such as reserved means reserved, and also indicates that a single or multiple registers are grouped with n.

[0071] In another exemplary embodiment, the vector element width configuration signal may be i_vnice_csr_vsew, which is used to configure the vector element width (VSEW), and its value will affect the size of the vector element, and further affect the number of elements in the vector and the control of the active element by the i_vnice_element_active signal, etc. Both provide necessary configuration parameters for instruction execution, and assist the i_vnice_instr instruction code in determining the specific operation mode and scale, so they belong to the instruction and configuration signal types.

[0072] The signal encoding rule of i_vnice_csr_vsew complies with the "V" extension standard of RISCV, as shown in Table 4 below: Table 4

[0073] Among them, the above vsew[2:0] is a 3-bit binary code, which is used to select the vector element width, and the specific width configuration is determined by different code values. SEW represents the vector element width, in bits. When vsew[2:0] is 000, SEW is 8 bits; when it is 001, SEW is 16 bits; when it is 010, SEW is 32 bits; when it is 011, SEW is 64 bits. When the highest bit of vsew[2:0] is 1 (i.e. 1XX), the code is reserved and does not correspond to the specific element width setting.

[0074] The first mode configuration signal may be i_vnice_csr_rounding_mode, and the second mode configuration signal may be i_vnice_fpu_rounding_mode. i_vnice_csr_rounding_mode is used to configure the rounding mode during conventional operations, and to determine the processing method for the decimal part during the operation, such as rounding, rounding up, rounding down, etc. Different rounding modes will affect the accuracy of the final operation result. i_vnice_fpu_rounding_mode is mainly used to configure the rounding mode of floating-point operations. In floating-point operations, when the result needs to be converted or truncated, it needs to be processed according to a specific rounding mode. This signal is used to specify this mode to ensure that the accuracy of the floating-point operation result meets the requirements. These two signals are similar to i_vnice_csr_vlmul, i_vnice_csr_vsew, etc. They are all used to assist the execution of instructions and provide specific configuration parameters for operations, so they belong to instructions and configuration signals.

[0075] The signal encoding rule of i_vnice_csr_rounding_mode complies with the "V" extension standard of RISCV, as shown in Table 5 below: Table 5

[0076] It should be noted that the above Table 5 is used to define the rounding modes and rounding increments corresponding to different encodings, as follows: vxrm[1:0] is a 2-bit binary code used to select different rounding modes. Abbreviation is the abbreviation of rounding mode. For example, "rnu", "rne", "rdn", and "rod". Rounding Mode is a detailed description of the rules for each rounding mode. Rounding Increment, r is the rounding increment calculation method for each rounding mode. For example, in the "rnu" mode, the rounding increment is v[d-1]; the "rne" mode has different judgment conditions to determine the rounding increment. This table is mainly used in related operations to select the appropriate rounding mode and calculate the rounding increment according to the vxrm[1:0] code to ensure the precision and accuracy of the operation results.

[0077] The signal encoding rules of the above i_vnice_csr_rounding_mode are in accordance with the "V" extension standard of RISCV, as shown in Table 6 below: Table 6

[0078] Among them, this is an information table about rounding mode, including three columns: Rounding Mode, Mnemonic, and Meaning: Rounding mode is to use binary numbers to represent different rounding options, such as "000", "001", etc. Mnemonics are abbreviated symbols set for each rounding mode, such as "RTZ", "RDN", etc., which are easy to remember and reference. Meaning is to explain the operating rules of each rounding mode in detail. For example, "RTZ" means rounding to zero, "RDN" means rounding down (to negative infinity); "RMM" means rounding to the nearest value, and if the distance is equal, rounding to the direction of the largest absolute value. At the same time, invalid (Invalid) modes and their reserved uses are also indicated. This table provides a reference for selecting the appropriate rounding method in calculations, which helps to accurately process calculation results.

[0079] In the embodiment of the present application, by setting the vector length multiplier configuration signal, the vector element width configuration signal, the first mode configuration signal and the second mode configuration signal, the mode of vector operation can be flexibly configured to adapt to different computing requirements and facilitate the storage and processing of data in vector operations.

[0080] In an optional embodiment of the present application, the operand signal further includes: a first data transmission signal and a second data transmission signal.

[0081] The first data transmission signal is used to transmit the corresponding operand according to the current instruction type; the second data transmission signal is used to write back the original value of the data element in the inactive state in the mask scenario; When the rdmac instruction is executed, the second data transmission signal is used to transfer data to the destination register.

[0082] It should be noted that the first data transmission signal may be i_vnice_rs1, which is an input signal having both integer and floating point transmission functions. According to the current instruction type, if it is an integer operation instruction, it transmits integer operands; if it is a floating point operation instruction, it transmits floating point operands to provide a data source for the operation.

[0083] The second data transfer signal can be i_vnice_vd, which is also an input signal. It is mainly used to write back the original value of the data element (non-active element) that is not in the active state in the mask scenario. The data only comes from the vector register. When the rdmac instruction is executed, this signal is not affected by the vector length (vl), vector mask (vm) and v0, and the data will be transferred to the destination register vd.

[0084] In the embodiment of the present application, by setting the first data transmission signal and the second data transmission signal, the system can selectively process vector elements according to mask rules when performing vector operations, thereby improving the flexibility and accuracy of vector operations, enhancing the versatility and applicability of the system, and meeting different types of computing requirements.

[0085] In an optional embodiment of the present application, the output result and flag signal include: an operation saturation flag signal and a floating point exception flag signal; The operation saturation flag signal is used to determine whether to write back the data element according to the data element status signal; the floating point exception flag signal is used to determine whether to write back the data element according to the data element status signal; If the data element status signal indicates that the corresponding element does not need to be written back, the operation saturation flag signal and the floating point exception flag signal will not be set. The operation saturation flag signal is updated and takes effect each time a vector length multiplication operation is completed; the floating point exception flag signal is updated and takes effect when the last vector length multiplication operation is performed.

[0086] It should be noted that the above operation saturation flag signal can be o_vnice_wbck_vxsat, which is an output signal and represents the operation saturation flag. This flag needs to be combined with the i_vnice_element_active signal to determine whether to write back. If i_vnice_element_active indicates that the corresponding element does not need to be written back, this flag will not be set. In addition, each time a VLMUL (vector length multiplier) operation is completed, this flag will be updated and take effect to indicate whether saturation occurs during the operation.

[0087] The floating point exception flag signal can be o_vnice_wbck_fflag, which is an output signal and is the vnice floating point exception flag. Similar to o_vnice_wbck_vxsat, it also needs to use i_vnice_element_active to determine whether to write back. If the corresponding element does not need to be written back, the flag is not set. The difference is that it takes effect only when the last VLMUL is written back, and is used to identify whether an exception occurs in the floating point operation.

[0088] In the embodiment of the present application, by setting the operation saturation flag signal, it is accurately reflected whether saturation occurs during the operation. In some numerical calculations, saturation occurs when the operation result exceeds the range that the data type can represent. Through this flag signal, the system can be informed of this situation in time, avoid the output of erroneous data caused by the overflow of the result, and provide accurate information for subsequent processing. And setting the floating point exception flag signal can feedback various abnormal conditions that occur in floating point operations, such as division by zero, overflow, underflow, etc. These abnormal conditions may frequently occur in the fields of scientific computing, graphics processing, etc., and timely capture and feedback of these anomalies will help the system to perform corresponding processing and ensure the accuracy and reliability of the operation results.

[0089] On the other hand, the present application embodiment also provides an instruction processing method, see Figure 3 As shown, the instruction processing method includes the following steps 201 to 204: Step 201, integrating the extended interface signal output by the extended instruction interface encoding unit to obtain an integration result; the extended interface signal is a pre-configured extensible interface signal; the integration result includes instructions.

[0090] Step 202: Distribute the instructions to the operation modules according to the integration result.

[0091] Step 203: Execute vector processing operations according to the distributed instructions to generate operation results.

[0092] Step 204: Execute corresponding processing operations according to the calculation results.

[0093] Please note that, see Figure 4 As shown in the figure, taking the decoding module as the CORE-DECODE module, the extended instruction interface encoding unit as the VNICE_DECODE unit, the operation module as the VPU, the distribution module as the DISP module, the instruction control unit as the VNICE-CORE unit, and the instruction co-processing module as the VNICE-CORE module as an example, the VNICE_DECODE unit in the above CORE-DECODE module outputs pre-configured extensible interface signals. These signals contain various information related to the instructions, such as the operation type of the instruction, the relevant information of the operands, the control parameters of special functions, etc. The role of the CORE-DECODE module is to analyze and process these extended interface signals, integrate them together, and form a complete and easy-to-understand integration result. The core part of this integration result is the instruction, which clearly describes the specific operation that the system needs to perform, as well as the various parameters and conditions required to perform the operation.

[0094] The DISP module receives the integration results from the CORE-DECODE module and extracts instructions from them. The DISP module accurately sends instructions to the operation module based on the specific content of the instruction, such as the operation type, target operation module, etc. This step ensures that the instruction can be correctly delivered to the module responsible for executing the operation, avoiding instruction distribution errors or confusion.

[0095] After receiving the instruction sent by the DISP module, the VPU performs the corresponding vector processing operation according to the instruction requirements. This may include basic operations such as vector addition, subtraction, multiplication, and division, and may also involve some special vector operations, such as vector mask operations and vector reduction operations. In the process of executing the operation, the operation module will process the input operands according to the parameters and conditions specified by the instruction, and finally generate the operation result.

[0096] The VNICE-CORE module or other related modules will receive the operation results generated by the operation module. According to the specific situation of the operation results, the system will perform corresponding processing operations. These operations may include storing the results in a specified memory location or register, making conditional judgments based on the results and determining the subsequent execution process, or triggering some exception handling mechanisms (such as overflow, error, etc. in the operation results).

[0097] The above instruction processing flow realizes accurate processing and execution of extended instructions through collaboration between multiple modules, ensuring that the system can complete various vector processing tasks efficiently and reliably.

[0098] Compared with the prior art, the instruction processing method based on the RISC-V architecture in this embodiment is different from the prior art in that an extended instruction interface encoding unit is provided in the decoding module, so that the user can flexibly design instructions with different requirements according to different scenarios, meet the ever-changing business needs, and make the device more flexible. In addition, through the orderly collaboration among the decoding module, the distribution module and the operation module, the delay of instruction processing is reduced, and the effective execution of vector processing operations is ensured. In addition, through the interaction between the instruction co-processing module and the instruction control unit, the system can flexibly process different instructions. This interactive mechanism can dynamically adjust the processing strategy according to the specific instruction type and business needs, thereby accelerating the computing performance of the processing device.

[0099] On the other hand, an embodiment of the present application also provides an instruction processing system, which includes an instruction processing device based on the RISC-V architecture provided in the above embodiment.

[0100] Compared with the prior art, the processing system of the present application, due to the provision of an arbitration module, can generate authorization signals and core identifiers to the corresponding target cores in order of priority, so that each processor core can access the resources of the operation processing unit in sequence to perform operation operations, without the need to configure an independent operation unit for each core, thereby realizing computing resource sharing in multi-core mode, and then sending the operation results to the target core through the first multiplexer, thereby greatly reducing hardware costs.

[0101] It should be understood that, although the various steps in the flow chart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0102] In one embodiment, a computer device is provided, the internal structure diagram of which can be as follows: Figure 1 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an instruction processing method as described above is implemented. It includes: a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, any step in the instruction processing method as described above is implemented.

[0103] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any step in the above instruction processing method can be implemented.

[0104] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0106] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0108] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0109] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. An instruction processing device based on RISC-V architecture, characterized in that: The instruction processing device comprises: a processor core and an instruction co-processing module, wherein the processor core comprises: a decoding module, a distribution module and a calculation module connected in sequence; The decoding module includes an extended instruction interface encoding unit; the decoding module is used to integrate the extended interface signal output by the extended instruction interface encoding unit to obtain an integration result and send it to the distribution module; the extended interface signal is a pre-configured and extensible interface signal; The distribution module is used to: distribute instructions to the operation module according to the integration result; The operation module includes an instruction control unit, which is used to: receive the instruction distributed by the distribution module, control the operation module to perform vector processing operations to generate operation results and send the operation results to the instruction co-processing module; The instruction co-processing module is used to interact with the instruction control unit and perform corresponding processing operations according to the calculation results.

2. The device according to claim 1, characterized in that The extended instruction interface encoding unit is used for: Receive vnice command code input from external module; According to the pre-configured instruction encoding rules, the vnice instruction code is subjected to characteristic analysis, and the extended interface signal is output.

3. The device according to claim 2, characterized in that The extended interface signal includes: at least one first operand interface signal, at least one second operand interface signal, at least one result register interface signal, and at least one other interface signal; The first operand interface signal is used to characterize the enable control, source indication and mutual exclusion constraint of the first operand of the vnice instruction; The second operand interface signal is used to characterize the enable control and source limitation of the second operand of the vnice instruction; The result register interface signal is used to characterize the enable control, update type indication, mutual exclusion constraint and read operation indication of the result register of the vnice instruction; The other interface signals are used to indicate the legality judgment of the vnice instruction, the instruction type identification and the mask use identification.

4. The device according to claim 1, characterized in that The instruction co-processing module interacts with the instruction control unit via a co-processing interface signal; The co-processing interface signal includes at least one of the following: a clock signal, a handshake interaction signal, a transmission flag signal, an instruction and configuration signal, an operand signal, an output result and a flag signal; The clock signal is used to represent the clock reference for the instruction co-processing module to perform operations; The handshake interaction signal is used to characterize the validity of input data and the readiness of the module to receive data, as well as the validity of output data and the readiness of the downstream module to receive data; The transmission flag signal is used to represent flag information of data transmission; The operand signal is used to represent operand information; the operand information includes at least one item: a vector operand and a scalar operand; The output result and the flag signal are used to characterize the operation process attribute information and the operation result information.

5. The device according to claim 4, characterized in that The operand signal includes: a data element status signal; the data element status signal is used to indicate whether the current data element in the vector operand is in an active state; the value of the data element status signal is determined according to different combinations of the vector length, the vector length multiplier and the data element width; the value includes a set of data element status values ​​of each current bit in the vector operand, and the data element status value includes 1 and 0; When the data element state value is 1, a group of data elements representing the current position is in an active state; when the data element state value is 0, a group of data elements representing the current position is in an inactive state.

6. The device according to claim 4, characterized in that The instructions and configuration signals include a vector length multiplier configuration signal, a vector element width configuration signal, a first mode configuration signal, and a second mode configuration signal; The vector length multiplier configuration signal is used to indicate the configuration of the vector length multiplier; The vector element width configuration signal is used to indicate the configuration data element width; The first mode configuration signal is used to indicate the rounding mode of configuring the conventional operation; The second mode configuration signal is used to indicate the configuration of the rounding mode of the floating point operation.

7. The device according to claim 4, characterized in that The operand signal also includes: a first data transmission signal and a second data transmission signal; The first data transmission signal is used to transmit the corresponding operand according to the current instruction type; The second data transmission signal is used to write back the original value of the data element in the inactive state in the mask scenario; When the rdmac instruction is executed, the second data transmission signal is used to transfer data to the destination register.

8. The device according to claim 5, characterized in that The output result and flag signal include: an operation saturation flag signal and a floating point exception flag signal; The operation saturation flag signal is used to determine whether to write back the data element according to the data element status signal; the floating point exception flag signal is used to determine whether to write back the data element according to the data element status signal; Wherein, if the data element status signal indicates that the corresponding element does not need to be written back, the operation saturation flag signal and the floating point exception flag signal will not be set; Each time a vector length multiplication operation is completed, the operation saturation flag signal is updated and becomes effective; When the last vector length multiplication operation is performed, the floating point exception flag signal is updated and becomes effective.

9. A processing system, characterized in that: include: An instruction processing device as claimed in any one of claims 1 to 8.

10. A method for processing an instruction, characterized in that: Applied to the instruction processing device according to any one of claims 1 to 8, the method comprises: Integrating the extended interface signal output by the extended instruction interface encoding unit to obtain an integrated result; the extended interface signal is a pre-configured expandable interface signal; the integrated result includes instructions; Distributing instructions to the computing modules according to the integration result; According to the distributed instructions, execute vector processing operations to generate operation results; A corresponding processing operation is performed according to the calculation result.

Citation Information

Patent Citations

  • Extended instruction interface of superscale RISC-V processor

    CN115309454A

  • Realization method and architecture of RISC-V vector processing unit

    CN115774575A

  • Processor supporting implicit sharing of scalar and vector operation units and application method thereof

    CN118585248A

  • Matrix and vector operation device based on RISC-V extension instruction

    CN119166218A

  • Processor hardware and instructions for vectorized fused and-xor

    US20230305846A1

Cited By

  • Dynamic register grouping method and system based on vector expansion

    CN121255287A

  • Vector extension based register dynamic grouping method and system

    CN121255287B