Processing unit, system-on-chip, computing device and method

By introducing a vector parameter prediction unit in the vector computing system, predicting immediately numeric vector parameters and executing vector operation instructions first, the performance degradation caused by waiting for parameter settings of vector operation instructions is solved, and execution efficiency is improved.

CN113885943BActive Publication Date: 2025-07-04ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010629013.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-02
Publication Date
2025-07-04
Estimated Expiration
2040-11-09

AI Technical Summary

Technical Problem

In the prior art, vector operation instructions need to wait for the vector parameter setting instructions to configure vector parameters before they can be executed, resulting in a degradation in execution performance.

Method used

By introducing a vector parameter prediction unit, the numeric vector parameters are predicted immediately, and without waiting for the vector parameter setting instruction to be executed, the vector operation instruction is first executed, and then the correction is performed according to the actual parameters.

Benefits of technology

It improves the execution performance and efficiency of vector operation instructions, and eliminates the correlation between vector operation instructions and vector parameter setting instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113885943B_ABST
    Figure CN113885943B_ABST
Patent Text Reader

Abstract

The present disclosure provides a processing unit, a system on chip, a computing device and a method. The processing unit includes: an instruction fetch unit configured to sequentially fetch a vector parameter setting instruction and a vector operation instruction; a vector parameter prediction unit configured to predict an immediate-type vector parameter according to the vector parameter setting instruction; an instruction decoding unit configured to separately decode the fetched vector parameter setting instruction and the vector operation instruction; and a vector execution unit configured to execute the decoded vector operation instruction according to the predicted immediate-type vector parameter without waiting for the decoded vector parameter setting instruction to be executed completely. Embodiments of the present disclosure improve the execution performance of vector operation instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of chips, and more particularly, to a processing unit, a system on a chip, a computing device, and a method. Background Art

[0002] Vector operation is an operation that can generate execution results of multiple elements in parallel. That is, for the same type of operations, such as calculating profits for each type of goods, according to requirements such as register capabilities, the unit prices, sales quantities, and profit margins of multiple goods can be fetched at one time and calculated in parallel. These unit prices, sales quantities, and profit margins are the elements of the vector, and the number of elements included in the vector is the number of elements that can be run in parallel at one time. Compared with single-element operations, vector operations greatly improve the operation efficiency.

[0003] When performing the above vector operations, it is necessary to indicate the vector parameters required for the operations, including the element size in the vector and the number of elements for each operation, etc. One existing technology is to encode them in arithmetic operation instructions, such as SIMD instruction sets like arm neon and intel sse. The disadvantage of this is that it occupies the encoding space of the instructions and is not conducive to the reuse of software code segments. Another existing technology is to specify them through other means not encoded in the vector operation instructions, such as vector instruction sets like risc-v vector and ARM SVE. For example, outside the vector operation instructions truly used for vector operations, a vector parameter setting instruction can be set to specify the vector parameters required for vector operations. Subsequently, the vector operation instructions will execute according to the vector parameters specified by the vector parameter setting instruction. It is specified in a separate vector parameter setting instruction outside the vector operation instructions and will not occupy the encoding space of the vector operation instructions. Moreover, since it sets the vector parameters required for many subsequent vector operation instructions through a single separate vector parameter setting instruction, it is conducive to the reuse of software code segments.

[0004] However, since the vector operation instructions need to wait for the vector parameter setting instruction to configure the vector parameters before they can be executed, that is, there is a vector parameter dependency between the vector operation instructions and the vector parameter setting instructions, which will greatly reduce the execution performance of the vector operation instructions. Summary of the Invention

[0005] In view of this, embodiments of the present invention aim to improve the execution performance of vector operation instructions.

[0006] To achieve this purpose, according to one aspect of the present disclosure, there is provided a processing unit, including:

[0007] An instruction fetching unit for sequentially fetching a vector parameter setting instruction and a vector operation instruction;

[0008] A vector parameter prediction unit for predicting an immediate vector parameter according to the vector parameter setting instruction;

[0009] An instruction decoding unit for separately decoding the retrieved vector parameter setting instruction and vector operation instruction;

[0010] A vector execution unit for executing the decoded vector operation instruction according to the predicted immediate vector parameter without waiting for the decoded vector parameter setting instruction to be executed completely.

[0011] Optionally, the vector execution unit includes a vector parameter setting subunit and a vector operation subunit. The vector parameter setting subunit is used to execute the decoded vector parameter setting instruction, and the vector operation subunit is used to execute the decoded vector operation instruction according to the predicted immediate vector parameter.

[0012] Optionally, the vector operation subunit is further configured to: use the non-immediate vector parameter in the previous decoded vector operation instruction of the decoded vector operation instruction received by the vector operation subunit as the non-immediate vector parameter in the received decoded vector operation instruction, and execute the decoded vector operation instruction according to the predicted immediate vector parameter.

[0013] Optionally, the immediate vector parameter includes the element size in the vector, and the non-immediate vector parameter includes the number of elements for a single operation.

[0014] Optionally, the vector parameter setting subunit is provided with a vector parameter register. If the non-immediate vector parameter set after the vector parameter setting subunit executes the decoded vector parameter setting instruction is inconsistent with the non-immediate vector parameter in the vector parameter register, the set non-immediate vector parameter is passed to the vector operation subunit, and the vector operation subunit discards the indication result of the vector operation instruction, and re-executes the decoded vector operation instruction according to the passed non-immediate vector parameter and the predicted immediate vector parameter.

[0015] Optionally, the vector parameter setting subunit also updates the vector parameter register with the vector parameter set after executing the decoded vector parameter setting instruction.

[0016] Optionally, if the non-immediate vector parameter set after the vector parameter setting subunit executes the decoded vector parameter setting instruction is consistent with the non-immediate vector parameter in the vector parameter register, the vector parameter register is updated with the vector parameter set after executing the decoded vector parameter setting instruction.

[0017] Optionally, if the vector operation unit fails to execute, immediate and non-immediate vector parameters in the vector parameter register are called to execute the decoded vector operation instruction.

[0018] Optionally, after the vector parameter setting subunit executes the decoded vector parameter setting instruction, the legality of the set vector parameters is checked, and if the set vector parameters are illegal, the vector parameters are set to be legal according to a predetermined rule.

[0019] Optionally, the instruction fetch unit sequentially fetches vector parameter setting instructions and vector operation instructions from a memory external to the processing unit.

[0020] According to one aspect of the present disclosure, a system-on-chip is provided, including the processing unit as described above.

[0021] According to one aspect of the present disclosure, a computing device is provided, including the processing unit as described above.

[0022] According to one aspect of the present disclosure, a vector operation execution method is provided, including:

[0023] Sequentially obtaining vector parameter setting instructions and vector operation instructions;

[0024] Predicting immediate vector parameters according to the vector parameter setting instructions;

[0025] Decoding the fetched vector parameter setting instructions and vector operation instructions respectively;

[0026] Without waiting for the decoded vector parameter setting instruction to finish execution, executing the decoded vector operation instruction according to the predicted immediate vector parameters.

[0027] Optionally, the executing the decoded vector operation instruction according to the predicted immediate vector parameters includes: using non-immediate vector parameters in the previous decoded vector operation instruction of the decoded vector operation instruction as non-immediate vector parameters in the decoded vector operation instruction, and executing the decoded vector operation instruction according to the predicted immediate vector parameters.

[0028] Optionally, the immediate vector parameters include the element size in the vector, and the non-immediate vector parameters include the number of elements for a single operation.

[0029] Optionally, a vector parameter register is preset. After executing the decoded vector operation instruction according to the predicted immediate vector parameter, the method further includes: if the non-immediate vector parameter set after executing the decoded vector parameter setting instruction is inconsistent with the non-immediate vector parameter in the vector parameter register, discarding the indication result of the vector operation instruction, and re-executing the decoded vector operation instruction according to the set non-immediate vector parameter and the predicted immediate vector parameter.

[0030] Optionally, after discarding the indication result of the vector operation instruction and re-executing the decoded vector operation instruction according to the set non-immediate vector parameter and the predicted immediate vector parameter, the method further includes: updating the vector parameter register with the vector parameter set after executing the decoded vector parameter setting instruction.

[0031] Optionally, after executing the decoded vector operation instruction according to the predicted immediate vector parameter, the method further includes: if the non-immediate vector parameter set after the vector parameter setting subunit executes the decoded vector parameter setting instruction is consistent with the non-immediate vector parameter in the vector parameter register, updating the vector parameter register with the vector parameter set after executing the decoded vector parameter setting instruction.

[0032] Optionally, after executing the decoded vector operation instruction according to the predicted immediate vector parameter, the method further includes: if the vector operation subunit fails to execute, calling the immediate and non-immediate vector parameters in the vector parameter register and executing the decoded vector operation instruction.

[0033] Optionally, after decoding the retrieved vector parameter setting instruction and vector operation instruction, the method further includes: executing the decoded vector parameter setting instruction, checking the legality of the vector parameter set after executing the vector parameter setting instruction, and in the case where the set vector parameter is illegal, setting the vector parameter to be legal according to a predetermined rule.

[0034] Optionally, the sequential acquisition of the vector parameter setting instruction and the vector operation instruction includes: sequentially acquiring the vector parameter setting instruction and the vector operation instruction from a memory outside the processing unit.

[0035] For an immediate vector parameter, it can be directly read out in the vector parameter setting instruction, so it can be predicted and extracted. In this way, the vector operation instruction can be executed first according to the predicted immediate vector parameter without waiting for the vector parameter setting instruction to be executed. That is, it eliminates the correlation between the vector operation instruction and the vector parameter setting instruction and improves the execution performance and efficiency of the vector operation instruction. Description of the Drawings

[0036] Through the description of the embodiments of the present invention with reference to the following drawings, the above and other objects, features and advantages of the present invention will become clearer. In the drawings:

[0037] Figure 1 Showing the architecture diagram of the general terminal to which the processing unit of the embodiment of the present disclosure is applied;

[0038] Figure 2 It is the structural diagram of the processing unit in the general terminal in an embodiment of the present disclosure;

[0039] Figure 3 Showing the architecture diagram of the multi-core high-performance terminal to which the processing unit of the embodiment of the present disclosure is applied;

[0040] Figure 4 Showing the structural diagram of the processing unit in the multi-core high-performance terminal according to an embodiment of the present disclosure;

[0041] Figure 5 Showing the structural diagram of the processing unit in the multi-core high-performance terminal according to an embodiment of the present disclosure;

[0042] Figure 6 Showing the flowchart of the vector operation execution method according to an embodiment of the present disclosure. Detailed Embodiments

[0043] The following describes the present disclosure based on embodiments, but the present disclosure is not limited to these embodiments. In the following detailed description of the present disclosure, some specific details are described in detail. Those skilled in the art can fully understand the present disclosure without the description of these details. In order to avoid obscuring the essence of the present disclosure, well-known methods, processes and procedures are not described in detail. In addition, the drawings are not necessarily drawn to scale.

[0044] The following terms are used in this article.

[0045] Vector operation: A vector operation is an operation that can generate the execution results of multiple elements in parallel. That is, for the same type of operation, such as the operation of calculating the profit for each type of goods, according to requirements such as the register capacity, the unit prices, sales quantities and profit margins of multiple types of goods can be taken out at one time for parallel calculation.

[0046] Element: An operation object for one calculation in the parallel calculation of a vector operation is an element. In the above example of taking out the unit prices, sales quantities and profit margins of multiple types of goods at one time for parallel calculation, these unit prices, sales quantities and profit margins are the elements of the vector, and the number of elements included in the vector is the number of elements that can be run in parallel at one time.

[0047] Vector operation instruction: An instruction used to perform the above-mentioned vector operations. For example, vadd.vvv8, v8, v4 is a vector addition instruction, where v8, v8, v4 are the operands used in this instruction.

[0048] Vector parameter: A vector parameter is a resource configuration parameter used when executing a vector operation instruction, such as the element size in a vector and the number of elements for a single operation. It is not an operand in the vector operation instruction. An operand is the object of the vector operation. A vector parameter is not an object of the operation but reflects the resource allocation (such as the number of bits occupying the register) during the operation, etc.

[0049] Element size in a vector: The element size in a vector refers to how many bits (bit) an element in the vector occupies in the vector register. Assuming the total width of the vector register is 128 bits, determining how many bits an element in the vector occupies decides how many elements the vector register can accommodate at most, that is, how many elements can be parallelly operated simultaneously. For example, if an element in the vector occupies 16 bits, the vector register can accommodate at most 8 elements, and at most 8 elements can be parallelly operated simultaneously.

[0050] Number of elements for a single operation: It indicates the number of elements actually retrieved for parallel calculation. It is less than or equal to the maximum number of elements that the above-mentioned vector register can accommodate. For example, in the case where the vector register can accommodate at most 8 elements, if there are 20 elements to be operated (for example, to calculate the profits of 20 different commodities using profit = unit price of commodity × sales amount of commodity × profit rate of commodity), then in the first batch, 8 elements can be calculated at a time, that is, 8 elements are placed in a vector for vector operation; in the second batch, 8 elements can be calculated at a time, that is, 8 elements are placed in a vector for vector operation; in the third batch, the remaining 4 elements are calculated at a time, that is, the remaining 4 elements are placed in a vector for vector operation. For the first and second batches, the number of elements for a single operation is 8; for the third batch, the number of elements for a single operation is 4.

[0051] Vector parameter setting instruction: An instruction that is separate from the vector operation instruction and is used to set the vector parameters used in the vector operation instruction. Since directly encoding vector parameters into the vector operation instruction is not conducive to the reuse of software code segments, special vector parameter setting instructions are used to uniformly set vector parameters, and the set vector parameters may be reused by multiple subsequent vector operation instructions.

[0052] Immediate vector parameter: A vector parameter directly given in the vector parameter setting instruction without the need to address any register.

[0053] Non-immediate vector parameter: In the vector parameter setting instruction, only the address storing the actual vector parameter is given, and the vector parameter can only be obtained by addressing the address. An important type of non-immediate vector parameter is the register type vector parameter. In this instruction, the register address storing the vector parameter is given, and the vector parameter can be obtained by addressing the corresponding register according to the address.

[0054] General Overview

[0055] In the original processing unit (such as CPU), the instruction execution unit, which is the main body of instruction execution, can only perform one operation on elements for each instruction. The element is, for example, an operand. If another operation is to be performed, it can only be performed with the help of the next instruction, and the operation efficiency is low. However, in practice, similar operations often need to be performed in batches. For example, there is a batch of goods, and the unit price, sales volume and profit margin of each product are known, and the profit of each product is calculated. The profit calculation formula is the same, that is, profit = unit price × sales volume × profit margin. If a single instruction is used to execute an operation, the profit can only be calculated once for each type of goods. For example, if there are 1,000 types of goods, it will be calculated 1,000 times.

[0056] Since the application of vector operations on the chip, this problem has been solved. Vector operations are operations that can generate execution results of multiple elements in parallel. That is, for similar operations, such as the above-mentioned operation of calculating the profit for each type of goods, the unit price, sales quantity and profit rate of multiple goods can be taken out at one time and calculated in parallel according to the requirements of register capacity and other requirements. These unit prices, sales quantities and profit rates are the elements of the vector, and the number of elements contained in the vector is the number of elements that can be run in parallel at a time. For example, if there are 1,000 kinds of goods, and the profit calculation of 4 kinds of goods can be run in parallel each time, it will take 250 times to complete, which greatly improves the calculation efficiency compared to only executing one element at a time.

[0057] When performing the above vector operations, it is necessary to indicate the vector parameters required for the operations. Vector parameters are resource configuration parameters used when executing vector operation instructions, including the element size in the vector and the number of elements in a single operation, etc. One existing technique is to encode them in arithmetic operation instructions, such as SIMD instruction sets like arm neon and intel sse. In this way, in an arithmetic operation instruction, it not only indicates the operands required by the arithmetic operation instruction itself but also contains the vector parameters indicating resource configuration. The disadvantage of this is that it occupies the encoding space of the instruction and is not conducive to the reuse of software code segments. Another existing technique is to specify them through other means not encoded in the vector operation instruction, such as vector instruction sets like risc-v vector and ARM SVE. For example, outside the vector operation instruction for vector operations, a vector parameter setting instruction can be set to specify the vector parameters required for vector operations. Subsequent vector operation instructions will be executed according to the vector parameters specified by this vector parameter setting instruction. It is specified in a separate vector parameter setting instruction outside the vector operation instruction, which does not occupy the encoding space of the vector operation instruction, and since it can set the vector parameters required by many subsequent vector operation instructions through a single separate vector parameter setting instruction, it is conducive to the reuse of software code segments.

[0058] However, since the vector operation instruction needs to wait for the vector parameter setting instruction to configure the vector parameters before it can be executed, that is, there is a dependency on vector parameters between the vector operation instruction and the vector parameter setting instruction, which will greatly reduce the execution performance of the vector operation instruction.

[0059] In order to eliminate the dependency on vector parameters between the vector operation instruction and the vector parameter setting instruction and improve the execution performance of the vector operation instruction, the inventors of the present disclosure thought that for immediate-type vector parameters, they can be directly read out in the vector parameter setting instruction, so they can be predicted and extracted. In this way, the vector operation instruction can be executed first according to the predicted immediate-type vector parameters without waiting for the vector parameter setting instruction to be executed. Since only the immediate-type vector parameters are predicted and not all vector parameters are predicted, the vector parameter setting instruction still needs to be executed. After execution, the actually set vector parameters can be used to correct the previously predicted execution result of the vector operation. If in most cases, the predicted execution result of the vector operation is accurate without being corrected, then this predicted execution will greatly improve the execution efficiency and performance of the vector operation instruction.

[0060] Since the embodiments of the present disclosure can be applied to both general-purpose terminal processing units and multi-core high-performance terminal processing units. The system architectures and internal structures of processing units in two implementation manners will be described in detail below for general-purpose terminal processing units and multi-core high-performance terminal processing units respectively.

[0061] Overview of General Terminal System

[0062] Figure 1 A schematic block diagram of a general terminal to which embodiments of the present disclosure are applied is shown. The general terminal 10 is an example of a "central" system architecture. The general terminal 10 can be built based on various models of processors on the current market and is driven by operating systems such as WINDOWS TM operating system versions, UNIX operating systems, Linux operating systems, etc. In addition, the general terminal 10 can be implemented in hardware and / or software such as a PC, a desktop computer, a laptop, a server, and a mobile communication device.

[0063] As Figure 1 shown, the general terminal 10 of the embodiments of the present disclosure may include one or more processing units 12 and a memory 14.

[0064] The memory 14 in the general terminal 10 may be a main memory (abbreviated as main memory or memory). It is used to store instruction information and / or data information represented by data signals. For example, it stores data provided by the processing unit 12 (such as an operation result), and can also be used to implement data exchange between the processing unit 12 and an external storage device 16 (or referred to as auxiliary memory or external memory).

[0065] In some cases, the processing unit 12 may need to access the memory 14 to obtain data in the memory 14 or modify the data in the memory 14. Since the access speed of the memory 14 is relatively slow, in order to alleviate the speed gap between the processing unit 12 and the memory 14, the general terminal 10 further includes a cache memory 18 coupled to the bus 11. The cache memory 18 is used to cache some program data or message data that may be repeatedly called in the memory 14. The cache memory 18 is implemented by a storage device such as a Static Random Access Memory (abbreviated as SRAM), for example. The cache memory 18 can be a multi-level structure, such as a three-level cache structure having a first-level cache (L1 Cache), a second-level cache (L2 Cache), and a third-level cache (L3 Cache), or can also be a cache structure with three or more levels or other types of cache structures. In some embodiments, a part of the cache memory 18 (such as the first-level cache, or the first-level cache and the second-level cache) can be integrated inside the processing unit 12 or integrated with the processing unit 12 on the same system-on-chip.

[0066] Based on this, the processing unit 12 may include components such as an instruction execution unit 121 and a memory management unit 122. When the instruction execution unit 121 executes some instructions that need to modify the memory, it initiates a write access request, which specifies the write data to be written into the memory and the corresponding physical address; the memory management unit 122 is used to translate the virtual address specified by these instructions into the physical address mapped by the virtual address, and the physical address specified by the write access request may be the same as the physical address specified by the corresponding instruction.

[0067] The information interaction between the memory 14 and the cache memory 18 is usually organized in blocks. In some embodiments, the cache memory 18 and the memory 14 may be divided into data blocks according to the same spatial size, and the data block may be the minimum unit of data exchange between the cache memory 18 and the memory 14 (including one or more data of a preset length). For the sake of concise and clear description, hereinafter, each data block in the cache memory 18 is simply referred to as a cache block (which may be called a cacheline or a cache line), and different cache blocks have different cache block addresses; each data block in the memory 14 is simply referred to as a memory block, and different memory blocks have different memory block addresses. The cache block address includes, for example, a physical address tag for locating the data block.

[0068] Due to space and resource limitations, the cache memory 18 cannot cache all the content in the memory 14, that is, the storage capacity of the cache memory 18 is usually smaller than that of the memory 14, and the cache block addresses provided by the cache memory 18 cannot correspond to all the memory block addresses provided by the memory 14. When the processing unit 12 needs to access the memory, it first accesses the cache memory 18 via the bus 11 to determine whether the content to be accessed has been stored in the cache memory 18. If so, the cache memory 18 is hit, and at this time, the processing unit 12 directly calls the content to be accessed from the cache memory 18; if the content that the processing unit 12 needs to access is not in the cache memory 18, then the cache memory 18 misses, and the processing unit 12 needs to access the memory 14 via the bus 11 to find the corresponding information in the memory 14. Because the access rate of the cache memory 18 is very fast, when the cache memory 18 is hit, the efficiency of the processing unit 12 can be significantly improved, and thus the performance and efficiency of the entire general-purpose terminal 10 are also improved.

[0069] In addition, the general-purpose terminal 10 may further include input / output devices such as a storage device 16, a display device 13, an audio device 14, a mouse / keyboard 15, etc. The storage device 16 is, for example, a device for information access such as a hard disk, an optical disc, and a flash memory that are coupled to the bus 11 through corresponding interfaces. The display device 13 is, for example, coupled to the bus 11 through a corresponding graphics card and is used for displaying according to the display signal provided by the bus 11.

[0070] The general-purpose terminal 10 usually further includes a communication device 17, so it can communicate with a network or other devices in various ways. The communication device 17 may, for example, include one or more communication modules. As an example, the communication device 17 may include a wireless communication module suitable for a specific wireless communication protocol. For example, the communication device 17 may include a WLAN module for implementing Wi-FiTM communication compliant with the 802.11 standard formulated by the Institute of Electrical and Electronics Engineers (IEEE); the communication device 17 may also include a WWAN module for implementing wireless wide-area communication compliant with a cellular or other wireless wide-area protocol; the communication device 17 may further include communication modules using other protocols such as a Bluetooth module, or other custom-type communication modules; the communication device 17 may also be a port for serial data transmission.

[0071] Of course, depending on the motherboard, operating system, and instruction set architecture, the structure of different general-purpose terminals may also vary. For example, many current general-purpose terminals are provided with an input / output control center connected between the bus 11 and each input / output device, and this input / output control center may be integrated within the processing unit 12 or independent of the processing unit 12.

[0072] Processing Unit of General Terminal

[0073] Figure 2 It is a schematic block diagram of the processing unit 12 of the general-purpose terminal in the embodiment of the present disclosure.

[0074] In some embodiments, each processing unit 12 may include one or more processor cores 120 for processing instructions, and the processing and execution of instructions can be controlled by a user (e.g., through an application) and / or the system platform. In some embodiments, each processor core 120 may be used to process a specific instruction set. In some embodiments, the instruction set may support Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or Very Long Instruction Word (VLIW)-based computing. Different processor cores 120 may each process different or the same instruction sets. In some embodiments, the processor core 120 may further include other processing modules, such as a Digital Signal Processor (DSP), etc. As an example, Figure 2 processor cores 1 to m are shown, where m is a non-zero natural number.

[0075] In some embodiments, Figure 1 the shown cache memory 18 may be fully or partially integrated into the processing unit 12. And according to different architectures, the cache memory 18 may be a single or multi-level internal cache memory located inside and / or outside each processor core 101 (such as Figure 2 the three-level cache memories L1 to L3 shown, Figure 2 which are uniformly identified as 18), and may also include an instruction cache for instructions and a data cache for data. In some embodiments, each component in the processing unit 12 may share at least a part of the cache memory. As shown in Figure 2 for example, processor cores 1 to m share the third-level cache memory L3. The processing unit 12 may further include an external cache (not shown), and other cache structures may also serve as the external cache of the processing unit 12.

[0076] In some embodiments, as shown in Figure 2 the processing unit 12 may include a register file 126. The register file 126 may include multiple registers for storing different types of data and / or instructions, and these registers may be of different types. For example, the register file 126 may include: integer registers, floating-point registers, status registers, instruction registers, pointer registers, etc. The registers in the register file 126 may be implemented using general-purpose registers, or may adopt a specific design according to the actual requirements of the processing unit 12.

[0077] The processing unit 12 may include a Memory Management Unit (MMU) 122 for implementing the translation of virtual addresses to physical addresses. A part of the table entries in the page table is cached in the memory management unit 122, and the memory management unit 122 can also obtain the uncached table entries from the memory. One or more memory management units 122 may be provided in each processor core 120, and the memory management units 120 in different processor cores 120 can also be synchronized with the memory management units 120 located in other processing units or processor cores, so that each processing unit or processor core can share a unified virtual storage system.

[0078] The processing unit 12 is used to execute an instruction sequence (i.e., a program). The process of the processing unit 12 executing each instruction includes steps such as fetching an instruction from a memory or cache memory 18 storing the instruction, decoding the fetched instruction, executing the decoded instruction, and saving the instruction execution result, and so on in a loop until all the instructions in the instruction sequence are executed or a halt instruction is encountered. For the execution of vector operations in the embodiments of the present disclosure, this process includes: sequentially fetching vector parameter setting instructions and vector operation instructions (different from the above-mentioned ordinary instructions, vector operations require prior configuration of the resources occupied by vector operations, so vector parameter setting instructions are needed). Among them, the vector parameter setting instructions and vector operation instructions can be fetched from the memory 14 outside the processing unit 12, but in some cases, they may also be fetched from the L1, L2, and L3 cache memories inside the processing unit 12, and with the development in the future, they may also be fetched from other storage units provided inside the processing unit 12; predicting immediate-type vector parameters; decoding the fetched instructions; speculatively executing the decoded vector operation instructions according to the predicted immediate-type vector parameters; executing the decoded vector parameter setting instructions; once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameters used in the speculative execution of the decoded vector operation instruction, re-execute the vector operation instruction according to the execution result of the vector parameter setting instruction. Compared with the existing process of executing vector operations, the embodiments of the present disclosure add a process of predicting immediate-type vector parameters, and modify the existing process of sequentially executing vector parameter setting instructions and vector operation instructions to first speculatively execute the decoded vector operation instructions according to the predicted immediate-type vector parameters, and once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameters used in the speculative execution of the decoded vector operation instruction, correct the execution result of the speculatively executed vector operation instruction. Correspondingly, a vector parameter prediction unit 129 is added to the processing unit 12 for predicting immediate-type vector parameters, and the vector execution unit 1212 is divided into a vector parameter setting sub-unit 12121 and a vector operation sub-unit 12122. Among them, the vector operation sub-unit 12122 is used to speculatively execute the decoded vector operation instructions according to the predicted immediate-type vector parameters, and the vector parameter setting sub-unit 12121 is used to execute the decoded vector parameter setting instructions. Once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameters used in the speculative execution of the decoded vector operation instruction by the vector operation sub-unit 12122, the previous execution result of the vector operation sub-unit 12122 is discarded, and the vector operation sub-unit 12122 is allowed to re-execute the vector operation according to the execution result of the decoded vector parameter setting instruction. The following describes the components of the processing unit 12 in detail.

[0079] The processing unit 12 may include an instruction fetch unit 124, a vector parameter prediction unit 129, an instruction decoding unit 125, an instruction issue unit (not shown), an instruction execution unit 121, an instruction retirement unit (not shown), etc. The instruction execution unit 121 includes an arithmetic operation unit 1211, a vector execution unit 1212, a multiplication and division operation unit 1213, etc. Among them, the vector execution unit 1212 includes a vector parameter setting subunit 12121 and a vector operation subunit 12122.

[0080] The instruction fetch unit 124 serves as the startup engine of the processing unit 12 and is used to transfer instructions from the memory 14 to the instruction register (which can be one of the registers in the register file 26 shown for storing instructions), but it is also possible to transfer instructions from the L1-3 level cache memory 18 inside the processing unit 12 or other storage units that may be set inside the future processing unit 12 to the instruction register, and receive the next instruction fetch address or calculate the next instruction fetch address according to the instruction fetch algorithm. The instruction fetch algorithm includes, for example: incrementing or decrementing the address according to the instruction length. The instructions fetched by the instruction fetch unit 124 may include vector parameter setting instructions, and subsequent vector operation instructions, etc. Figure 2 After the instruction is fetched, if the instruction is a vector parameter setting instruction, at this time, the instruction is not handed over to the instruction decoding unit 125 for decoding first, but the vector parameter prediction unit 129 performs immediate-type vector parameter prediction. Since the immediate-type vector parameter can be directly read from the vector parameter setting instruction, the vector operation subunit 12122 can still execute the vector operation instruction first according to the predicted immediate-type vector parameter under the condition that the vector parameter setting instruction has not been executed by the vector parameter setting subunit 12121.

[0081] Then, the instruction decoding unit 125 decodes the fetched instruction according to a predetermined instruction format to obtain the operand acquisition information required for the fetched instruction, so as to prepare for the operation of the instruction execution unit 121. The operand acquisition information points to, for example, an immediate number, a register, or other software / hardware that can provide source operands.

[0082] The instruction issue unit usually exists in a high-performance processing unit 12 and is located between the instruction decoding unit 125 and the instruction execution unit 121. It is used for instruction scheduling and control to efficiently distribute each instruction to different instruction execution units 121, making it possible to perform parallel operations on multiple instructions. After the instruction is fetched, decoded, and scheduled to the corresponding instruction execution unit 121, the corresponding instruction execution unit 121 starts to execute the instruction, that is, performs the operation indicated by the instruction and realizes the corresponding function.

[0083]

[0084] ​For different types of instructions, different execution units can be correspondingly set in the instruction execution unit 121. The instruction execution unit 121 includes an arithmetic operation unit 1211, a vector execution unit 1212, a memory execution unit 1214, etc., which are respectively responsible for executing different types of instructions. The arithmetic operation unit 1211 is an execution unit that executes only one arithmetic operation each time. The vector execution unit 1212 is a unit that executes vector operation instructions, that is, a unit that executes multiple operations each time. It can parallelly generate execution results of multiple elements. That is, for the same type of operation, such as calculating the profit for each type of goods, according to requirements such as register capabilities, the unit prices, sales quantities, and profit margins of multiple types of goods can be fetched at one time and calculated in parallel. The unit price, sales quantity, and profit margin of each type of goods are one element, and there are multiple elements in the vector. In this way, through vector operations, multiple arithmetic operations can be performed at one time. The memory execution unit 1214 is a unit used to execute memory access instructions. The arithmetic operation unit 1211, the vector execution unit 1212, the memory execution unit 1214, etc. can run in parallel and output corresponding execution results.

[0085] The vector execution unit 1212 includes a vector parameter setting subunit 12121 and a vector operation subunit 12122. The instructions fetched by the instruction fetching unit include vector parameter setting instructions and subsequent vector operation instructions. Among them, the vector parameter setting instruction is an instruction for setting vector parameters for subsequent vector operation instructions so that the vector operation instructions can be executed. The vector parameter setting subunit 12121 is used to execute the decoded vector parameter setting instruction, while the vector operation subunit 12122 is used to execute the decoded vector operation instruction. Different from the prior art, since the vector parameter prediction unit 129 is adopted in the embodiments of the present disclosure, the vector operation subunit 12122 does not need to wait until the vector parameter setting subunit 12121 finishes execution and then start execution using the vector parameter setting result of the vector parameter setting subunit 12121. Instead, it can first execute using the immediate-type vector parameters predicted by the vector parameter prediction unit 129 and adopt post-execution program correction. In this way, the execution efficiency of vector operations is greatly improved.

[0086] The instruction retirement unit (or called the instruction write-back unit) is mainly responsible for writing the execution results generated by the instruction execution unit 121 back to the corresponding storage locations (such as registers inside the processing unit 12) so that subsequent instructions can quickly obtain the corresponding execution results from these storage locations.

[0087] When the instruction execution unit 121 executes a certain type of instruction (such as a memory access instruction), it needs to access the memory 14 to obtain the information stored in the memory 14 or provide the data to be written into the memory 14. This access is performed by the memory execution unit 1214. The memory execution unit is, for example, a Load Store Unit (LSU) and / or other units for memory access.

[0088] After the memory access instruction is fetched by the instruction fetch unit 124, the instruction decoding unit 125 can decode the memory access instruction so that the source operands of the memory access instruction can be obtained. The decoded memory access instruction is provided to the corresponding memory execution unit 1214, and the memory execution unit 1214 can perform corresponding operations on the source operands of the memory access instruction (such as performing operations on the source operands stored in registers) to obtain the address information corresponding to the memory access instruction, and initiate corresponding requests according to this address information, such as an address translation request, a write access request, etc.

[0089] The source operands of the memory access instruction usually include address operands, and the memory execution unit 1214 performs operations on the address operands to obtain the virtual address or physical address corresponding to the memory access instruction. When the memory management unit 122 is disabled, the memory execution unit 1214 can directly obtain the physical address of the memory access instruction through logical operations. When the memory management unit 121 is enabled, the corresponding memory execution unit 1214 initiates an address translation request according to the virtual address corresponding to the memory access instruction. The address translation request includes the virtual address corresponding to the address operand of the memory access instruction; the memory management unit 122 responds to the address translation request and converts the virtual address in the address translation request into a physical address according to the table entry matching the virtual address, so that the memory execution unit 1214 can access the cache memory 18 and / or the memory 14 according to the translated physical address.

[0090] According to different functions, memory access instructions can include load instructions and store instructions. The execution process of a load instruction usually does not modify the information in the memory 14 or the cache memory 18. The memory execution unit 1214 only needs to read the data stored in the memory 14, the cache memory 18, or an external storage device according to the address operand of the load instruction.

[0091] Different from load instructions, the source operands of store instructions include not only address operands but also data information. The execution process of store instructions usually needs to modify the memory 14 and / or the cache memory 18. The data information of the store instruction can point to the data to be written, and the source of the data to be written can be the execution result of instructions such as arithmetic instructions and load instructions, or the data provided by registers or other storage units in the processing unit 12, or an immediate number.

[0092] Overview of Multi-Core High-Performance Terminal

[0093] Figure 3 A schematic block diagram showing a multi-core high-performance terminal in an embodiment of the present disclosure. A multi-core high-performance terminal refers to a terminal or server that divides a large number of processor cores into clusters to improve the processing speed of the terminal and thus enhance the computing performance.

[0094] The multi-core high-performance terminal 10' is an example of a "central" system architecture. The multi-core high-performance terminal 10' can be based on various types of processor cores on the current market and be driven by operating systems such as WINDOWS TM operating system versions, UNIX operating systems, Linux operating systems, etc. In addition, the multi-core high-performance terminal or server 10' can be embodied as a single device, for example, a PC, a desktop computer, a laptop, a server, and a mobile communication device. The multi-core high-performance terminal or server 10' can also be embodied as a cluster composed of multiple devices, for example, a server cluster composed of a large number of servers in the cloud.

[0095] As Figure 3 shown, the multi-core high-performance terminal or server 10' in an embodiment of the present disclosure may include one or more processing units 12'. Each processing unit 12' contains multiple processor clusters 130', and each processor cluster 130' contains multiple processor cores 120'. The reason for dividing into processor clusters 130' is to facilitate the management of a large number of processor cores 120'. Processor cores 120' that perform similar tasks can be grouped into one cluster. In this way, when scheduling the processing unit, it is convenient to decide which processor core 120' to use according to the cluster to which it belongs.

[0096] The multi-core high-performance terminal or server 10' may also include a storage device 16'. The storage device 16' is, for example, a hard disk, an optical disc, and a flash memory, etc. for information access and storage, which are coupled to the system bus 11' through corresponding interfaces. Multiple levels of cache memories 127', 129', 131', etc. are provided inside the processing unit 12' of the multi-core high-performance terminal 10' for storing instruction information and / or data information that may be repeatedly used and represented by data signals. Only instruction information and / or data information that is not frequently used is stored in the storage device 16'.

[0097] In the multi-core high-performance terminal 10', an L3 or last-level cache 18' is also included outside the processing unit 12'. The cache is implemented by a storage device such as a Static Random Access Memory (abbreviated as SRAM). The cache in the multi-core high-performance terminal 10' can be a multi-level structure, for example, a first-level cache ( Figures 4 - 5the L1 cache 127’ therein), the secondary cache ( Figures 2 - 3 the L2 cache 129’ therein), and the tertiary cache ( Figure 4 in the embodiment, the L3 cache 18’ disposed outside the processing unit 12’, as Figure 3 shown, and Figure 5 the L3 cache 131’ therein), and Figure 5 in the embodiment, the last-level cache 18’ disposed outside the processing unit 12’ (as Figure 3 shown).

[0098] The multi-core high-performance terminal 10’ generally further includes a communication device 17’, and thus can communicate with a network or other devices in various ways. The communication device 17’ may include, for example, one or more communication modules. As an example, the communication device 17’ may include a wireless communication module applicable to a specific wireless communication protocol. For example, the communication device 17’ may include a WLAN module for implementing Wi-FiTM communication compliant with the 802.11 standard formulated by the Institute of Electrical and Electronics Engineers (IEEE); the communication device 17’ may include a WWAN module for implementing wireless wide-area communication compliant with a cellular or other wireless wide-area protocol; the communication device 17’ may further include communication modules using other protocols such as a Bluetooth module, or other custom-type communication modules; the communication device 17’ may also be a port for serial data transmission.

[0099] Of course, the structures of different multi-core high-performance terminals 10’ may vary according to different motherboards, operating systems, and instruction set architectures. For example, some multi-core high-performance terminals 10’ may have a display device, input / output devices, and so on.

[0100] Processing Unit of Multi-Core High-Performance Terminal

[0101] Figure 4 is a schematic block diagram of a Figure 3 processing unit according to an embodiment of the present disclosure.

[0102] In this embodiment, the processing unit 12' may include a plurality of processor clusters 130', and each processor cluster 130' may include one or more processor cores 120' for processing instructions. The processing and execution of instructions can be controlled by a user (e.g., through an application) and / or the system platform. In some embodiments, each processor core 120' may be used to process a specific instruction set. In some embodiments, the instruction set may support Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computing based on Very Long Instruction Word (VLIW). Different processor cores 120' may process different or the same instruction sets. As an example, Figure 4 it is shown in Figure 4 that the processing unit 12' includes processor clusters 1 to n, where n is a non-zero natural number; each processor cluster includes processor cores 1 to m, where m is a non-zero natural number.

[0103] In this embodiment, each processor core 120' contains an L1 cache 127' inside for storing instruction information and data information that are frequently used inside the processor core 120'. Additionally, inside each processor cluster 130', there is an L2 cache 1282' shared by multiple processor cores 120'. The processor cores 120' are connected to the shared L2 cache 1282' through a coherence bus 1281' inside the processor cluster 130'. The L2 cache 1282' mainly caches instruction information and data information whose importance is slightly lower than that of the instruction information and data information stored in the L1 cache 127', or instruction information and data information whose access frequency is slightly lower than that of the instruction information and data information stored in the L1 cache 127'. Each processor cluster 130' is connected to the shared system bus 11', and an L3 cache 18' shared by each processor cluster 130' may be connected to the system bus 11'. The L3 cache 18' is used to cache instruction information and data information whose importance is lower than that of the instruction information and data information stored in the L2 cache 1282', or instruction information and data information whose access frequency is lower than that of the instruction information and data information stored in the L2 cache 1282'.

[0104] As Figure 4 shown, each processor core 120' may include an instruction fetch unit 124', a vector parameter prediction unit 129', an instruction decoding unit 125', a register file 126', an instruction execution unit 121', a memory management unit 122', and an L1 cache 127'.

[0105] The register file 126’ may include a plurality of registers for storing different types of data and / or instructions, and these registers may be of different types. For example, the register file 126’ may include: integer registers, floating-point registers, vector registers, status registers, pointer registers, etc. The registers in the register file 126’ may be implemented using general-purpose registers, or may adopt a specific design according to the actual requirements of the processing unit 12’.

[0106] The processing unit 12’ may include a Memory Management Unit (MMU) 122’ for implementing the translation of virtual addresses to physical addresses. The root page table is a table storing the mapping relationship from virtual addresses to physical addresses in the storage device 16’. Generally, the root page table is stored in the storage device 16’. When accessing data in the storage device 16’ each time, first obtain the mapping relationship from virtual addresses to physical addresses in the root page table of the storage device 16’. Then, based on the virtual address to be accessed and this mapping relationship, obtain the physical address in the storage device 16’ to be accessed. Then, access the data on the storage device 16’ according to this physical address. In this way, it is equivalent to accessing the data on the storage device 16’ twice each time, one time for obtaining the mapping relationship and the other time for obtaining the actual data to be accessed. To improve the access speed of some frequently used pages, some frequently used data can be changed from being stored in the storage device 16’ to being stored in the L1 cache 127’, L2 cache 1282’, L3 cache 131’ or the last-level cache 18’. The table entries of the mapping relationship from virtual addresses to physical addresses of the data placed in these caches are stored in the translation lookaside buffer 1221’ set inside the memory management unit 122’. Here, the physical address refers to its physical address in the cache, not the physical address on the storage device 16’. The memory management unit 122’ may store the table entries of the mapping relationship from virtual addresses to physical addresses on the L1 cache 127’, or may simultaneously store the table entries of the mapping relationship from virtual addresses to physical addresses on the L1 cache 127’ and the L2 cache 1282’, or may simultaneously store the table entries of the mapping relationship from virtual addresses to physical addresses on the L1 cache 127’, L2 cache 1282’ and L3 cache 131’, or may simultaneously store the table entries of the mapping relationship from virtual addresses to physical addresses on the L1 cache 127’, L2 cache 1282’, L3 cache 131’ and the last-level cache 18’. This is determined by the number of cache levels set. In addition, not all table entries of the mapping relationship from virtual addresses to physical addresses need to be stored in the root page table. Instead, several hierarchical page tables can be set. For example, look down from the root page table to the first-level page table, then look down from the first-level page table to the second-level page table, and keep looking down until the final page table. In this way, for a relatively wide physical address system, such a management method can compress the memory storage space of the entire page table.

[0107] One or more memory management units 122' can be set in each processor core 120'. The memory management units 120' in different processor cores 120' can also be synchronized with the memory management units 120' located in other processing units or processor cores, so that each processing unit or processor core can share a unified virtual storage system.

[0108] The processing unit 12’ is used to execute an instruction sequence (i.e., a program). The process of the processing unit 12’ executing each instruction includes: fetching an instruction from a storage device 16’ storing the instruction, or an instruction cache 1271’, or an L2 cache 1282’, or other storage units that may be set in the future in the processing unit 12’, decoding the fetched instruction, executing the decoded instruction, saving the instruction execution result, and so on, and repeating this process until all instructions in the instruction sequence are executed or a halt instruction is encountered. For the execution of vector operations in the embodiments of the present disclosure, this process includes: fetching a vector parameter setting instruction and a vector operation instruction from a storage device 16’ storing the instruction, or an instruction cache 1271’, or an L2 cache 1282’, or other storage units that may be set in the future in the processing unit 12’ (different from the above ordinary instructions, vector operations require prior configuration of the resources occupied by vector operations, so vector parameter setting instructions are needed), predicting an immediate-type vector parameter, decoding the fetched instruction, speculatively executing the decoded vector operation instruction according to the predicted immediate-type vector parameter, executing the decoded vector parameter setting instruction, and once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameter used in the speculatively executed decoded vector operation instruction, re-executing the vector operation instruction according to the execution result of the vector parameter setting instruction. Compared with the existing process of executing vector operations, the embodiments of the present disclosure add a process of predicting an immediate-type vector parameter, and modify the existing process of sequentially executing a vector parameter setting instruction and a vector operation instruction to first speculatively execute the decoded vector operation instruction according to the predicted immediate-type vector parameter, and once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameter used in the speculatively executed decoded vector operation instruction, correct the execution result of the speculatively executed vector operation instruction. Correspondingly, a vector parameter prediction unit 129’ is added to the processing unit 12’ for predicting an immediate-type vector parameter, and the vector execution unit 1212’ is divided into a vector parameter setting subunit 12121’ and a vector operation subunit 12122’. Among them, the vector operation subunit 12122’ is used to speculatively execute the decoded vector operation instruction according to the predicted immediate-type vector parameter, and the vector parameter setting subunit 12121’ is used to execute the decoded vector parameter setting instruction. Once the execution result of the decoded vector parameter setting instruction does not match the non-immediate-type vector parameter used in the speculatively executed decoded vector operation instruction by the vector operation subunit 12122’, the previous execution result of the vector operation subunit 12122’ is discarded, and the vector operation subunit 12122’ re-executes the vector operation according to the execution result of the decoded vector parameter setting instruction.

[0109] The L1 cache 127’ includes an instruction cache 1271’ and a data cache 1272’. The instruction cache 1271’ stores the instructions to be executed. The data cache 1272’ stores the operands for the operations to be performed by the instructions, as well as the intermediate results during the execution process, etc.

[0110] The instruction fetch unit 124’ serves as the startup engine of the processor core 120’ and is used to transfer instructions from the storage device 16’ storing the instructions, or the instruction cache 1271’, or the L2 cache 1282’, or other storage units that may be set in the future in the processing unit 12’ into the cache inside the instruction fetch unit 124’ or the instruction decoding unit 125’. Then, for general instructions, it receives the next instruction fetch address or calculates the next instruction fetch address according to the instruction fetch algorithm. The instruction fetch algorithm includes, for example: incrementing or decrementing the address according to the instruction length. If it is the target address of a branch or jump instruction, the instruction fetch unit 124’ needs to predict the jump direction or the target address. Common general techniques related to branch prediction or jump prediction include: using a one-level or two-level branch history table (BHT) to predict the jump direction of branch instructions, using a one-level or multi-level branch target buffer (BTB) to predict the jump target address of branch instructions, using a return address stack (RAS) to predict the jump target address of return instructions, and predicting the target address of indirect jump instructions.

[0111] After the instruction is fetched, if the instruction is a vector parameter setting instruction, this instruction is not handed over to the instruction decoding unit 125’ for decoding immediately, but is subject to immediate-type vector parameter prediction by the vector parameter prediction unit 129’. Since the immediate-type vector parameter can be directly read from the vector parameter setting instruction, the vector operation sub-unit 12122’ can still execute the vector operation instruction according to the predicted immediate-type vector parameter first under the condition that the vector parameter setting instruction has not been executed by the vector parameter setting sub-unit 12121’.

[0112] Next, the instruction decoding unit 125’ decodes the fetched instruction according to a predetermined instruction format to obtain the operand acquisition information required for the fetched instruction, so as to prepare for the operation of the instruction execution unit 121’. The operand acquisition information points to, for example, an immediate number, a register, or other software / hardware that can provide source operands. In Figure 4 it, it can point to the address of the operand in the data cache 1272’.

[0113] Then, the instruction decoding unit 125’ is also used for the scheduling and control of instructions to efficiently allocate each instruction to different instruction execution units 121’, making it possible to perform parallel operations on multiple instructions. After the instructions are fetched, decoded, and scheduled to the corresponding instruction execution units 121’, the corresponding instruction execution units 121’ start to execute the instructions, that is, perform the operations indicated by the instructions and implement the corresponding functions.

[0114] Inside the instruction execution unit 121’, it can be divided into an arithmetic operation unit 1211’, a vector execution unit 1212’, a multiplication and division operation unit 1213’, a storage instruction execution unit 1214’, etc., according to the different specific instructions to be executed. The arithmetic operation unit (ALU) 1211’ is an operation unit for performing integer operations (excluding multiplication and division operations) and logical operations, and it executes only one arithmetic operation at a time. The ALU is the most important component of the computer central processing unit. The vector execution unit 1212’ is an operation unit for performing vector operations, and it can generate execution results of multiple elements in parallel. The multiplication and division operation unit 1213’ is an operation unit for performing multiplication and division operations, and it also executes only one multiplication or division operation at a time. When the arithmetic operation unit 1211’, the vector execution unit 1212’, and the multiplication and division operation unit 1213’ execute instructions, they obtain information according to the operands fetched by the instruction decoding unit 125’ from the corresponding addresses in the register bank 126’, or directly fetch the results that have been generated during the execution of the instruction execution unit 121’, which can be achieved by using the forward technology. The storage instruction execution unit 1213’ is a processing unit for executing storage instructions. The most common processes in high-performance terminals or servers are storage, integer elements, vector operations, multiplication and division operations, logical operations, etc. Therefore, basically the above four execution units can cover the most frequently encountered processes in high-performance terminals. According to needs, the instruction execution unit 121’ may also include a control register, a branch jump processing unit, an encryption and decryption instruction processing unit, etc. (not shown).

[0115] The vector execution unit 1212’ includes a vector parameter setting subunit 12121’ and a vector operation subunit 12122’. The instructions fetched by the instruction fetch unit include vector parameter setting instructions and subsequent vector operation instructions. Among them, the vector parameter setting instruction is an instruction for setting vector parameters for subsequent vector operation instructions so that the vector operation instructions can be executed. The vector parameter setting subunit 12121’ is used to execute the decoded vector parameter setting instruction, while the vector operation subunit 12122’ is used to execute the decoded vector operation instruction. Different from the prior art, since the vector parameter prediction unit 129’ is adopted in the embodiment of the present disclosure, the vector operation subunit 12122 does not need to wait until the vector parameter setting subunit 12121’ finishes execution and then start execution using the vector parameter setting result of the vector parameter setting subunit 12121’. Instead, it can first execute using the immediate-type vector parameters predicted by the vector parameter prediction unit 129’ and perform post-execution program correction. In this way, the execution efficiency of vector operations is greatly improved.

[0116] Optionally, the processor core 120’ may further include an instruction retirement unit 135’, which is mainly responsible for writing the execution results generated by the instruction execution unit 121’ back to the storage locations of the corresponding register file 126’ so that subsequent instructions can quickly obtain the corresponding execution results from these storage locations. If there is this instruction retirement unit, all the paths for writing back to the register file 126’ have to pass through this unit.

[0117] Optionally, the processor core 120’ may further include a debug unit 136’, a trace unit 137’, and an interrupt unit 138’, which are respectively used for debugging, tracing, and interrupt handling of the processor core. They are not essential and can be added according to the design requirements. When these units exist, debug, trace, interrupt, etc. requests are sent to the above-mentioned instruction retirement unit 135’.

[0118] Figure 5 is a schematic block diagram of a processing unit according to another embodiment of the present disclosure. The difference between this embodiment and Figure 3 the embodiment of Figure 4 is that Figure 5In the processor core 120' in the embodiment, there are two levels of cache, namely, L1 cache 127' and L2 cache 1282'. Among them, the importance level or access frequency of the instruction information and data information cached in L2 cache 1282' is slightly lower than that of the instruction information and data information cached in L1 cache 127'. Multiple processor cores 120' are connected to the shared L3 cache 131' through the coherence bus 1281' inside the processor cluster 130'. The importance level or access frequency of the instruction information and data information cached in L3 cache 131' is slightly lower than that of the instruction information and data information cached in L2 cache 1282'. Each processor cluster 130' is connected to the shared system bus 11', and the last-level cache 18' shared by each processor cluster 130' can be connected to the system bus 11'. The last-level cache 18' is used to cache some instruction information and data information with a lower importance level or a lower access frequency than the instruction information and data information stored in L3 cache 131. In Figure 5 In the architecture of, since an additional level of cache is set in the processor core 120', it is suitable for scenarios where the processing load of the processor core is greater. In the case of a greater processing load, the processing efficiency of the processing unit is improved through the two-level cache structure inside the processor core 120'.

[0119] Detailed Implementation Process of Embodiments of the Present Disclosure

[0120] Vector operations cannot be executed relying solely on one vector operation instruction, and vector parameter setting instructions are also required. A vector parameter setting instruction is an instruction that is separate from the vector operation instruction and is used to set the vector parameters used in the vector operation instruction.

[0121] First, the instruction fetch units 124, 124' sequentially fetch the vector parameter setting instruction and the vector operation instruction from the memory 14 outside the processing units 12, 12' from the storage device 16' storing the instructions, or the instruction cache 1271', or the L2 cache 1282', or other storage units that may be set in the future in the processing unit 12'. For example, the vector parameter setting instruction is vsetvli t0,x0,e32, and the vector operation instruction is vadd.vv v8,v8,v4. Among them, vsetvli and vadd are instruction names, representing the vector parameter setting instruction and the vector addition instruction respectively, and v8, v8, v4 are operands, that is, the objects of the operation.

[0122] Vector parameters are resource configuration parameters used when executing vector operation instructions (such as the above vadd.vv v8,v8,v4), such as the element size in a vector and the number of elements in a single operation. It is not an operand as described above. An operand is the object of a vector operation. Vector parameters are parameters that reflect the resource allocation (such as the number of bits occupied in a register) during the operation, such as t0, x0, e32 as described above (the meanings of these parameters will be detailed below). Vector parameters are divided into immediate vector parameters and non-immediate vector parameters.

[0123] An immediate vector parameter refers to a vector parameter directly given in a vector parameter setting instruction without the need to address any register. For example, e32 in vsetvli t0,x0,e32 as described above, which indicates that the element size in the vector is 32 bits. This information is directly written in the instruction without the need to address other registers. A non-immediate vector parameter refers to a vector parameter whose address storing the actual vector parameter is only given in a vector parameter setting instruction, and the vector parameter can only be obtained by addressing this address. For example, t0 and x0 in vsetvli t0,x0,e32 as described above, which respectively represent the number of elements in a single operation and the maximum number of elements that can be stored in a vector register. The maximum number of elements x0 that can be stored in a vector register can actually be obtained by dividing the bit width of the vector register by the element size in the vector. If the bit width of the vector register is 128 bits and the element size in the vector is 16 bits, then at most 8 elements can be placed in the vector register simultaneously. The maximum number of elements that can be executed in a single operation is 8. In the case where the above vector register can accommodate at most 8 elements simultaneously, if there are 20 elements to be operated (for example, the profits of 20 different commodities need to be calculated), then in the first batch, 8 elements can be calculated in a single operation (the unit price, sales amount, and profit rate of each commodity are one element), that is, 8 elements are placed in a vector for vector operation; in the second batch, 8 elements can be calculated in a single operation, that is, 8 elements are placed in a vector for vector operation; in the third batch, the remaining 4 elements are calculated in a single operation, that is, the remaining 4 elements are placed in a vector for vector operation. For the first and second batches, the number of elements in a single operation is 8; for the third batch, the number of elements in a single operation is 4. In the above case, for the three batches, x0 is equal to 8; for the first two batches, t0 = 8; for the last batch, t0 = 4. For t0 and x0, their values are not directly written in the vector parameter setting instruction, and only their register identifiers t0 and x0 are given in the instruction. Therefore, to find their true values, it is necessary to address the register according to this identifier. Therefore, they are non-immediate vector parameters.

[0124] Although in the above example, the immediate vector parameter includes the element size in the vector, and the non-immediate vector parameter includes the number of elements in a single operation. However, it can also be set conversely, that is, setting the number of elements in a single operation as the immediate vector parameter and setting the element size in the vector as the non-immediate vector parameter.

[0125] The vector register is one of the registers in the register files 126, 126', and is dedicated to storing the vector in vector operations.

[0126] The execution of the vector operation instruction vadd.vv v8, v8, v4 is restricted by the above vector parameter setting instruction vsetvli t0, x0, e32. The execution of the vector operation instruction vadd.vv v8, v8, v4 should be performed according to the element size in the vector and the number of elements in a single operation specified by the vector parameter setting instruction vsetvli t0, x0, e32. Therefore, in the prior art, it is necessary to wait until the vector parameter setting instruction vsetvli t0, x0, e32 is executed before the vector operation instruction vadd.vv v8, v8, v4 can be executed.

[0127] In the embodiments of the present disclosure, considering that the immediate vector parameter in the vector parameter setting instruction can be directly read out from the vector parameter setting instruction, the vector parameter prediction units 129, 129' can predict the immediate vector parameter according to the vector parameter setting instruction, and then let the vector operation sub-units 12122, 12122' execute the vector operation instruction according to the predicted immediate vector parameter without waiting for the vector parameter setting sub-units 12121, 12121' to execute the vector parameter setting instruction.

[0128] The method for the vector parameter prediction units 129, 129' to predict the immediate vector parameter according to the vector parameter setting instruction can be to directly read it out from the vector parameter setting instruction, such as e32 in the above vector parameter setting instruction vsetvli t0, x0, e32.

[0129] Then, the instruction decoding units 125, 125' respectively decode the fetched vector parameter setting instruction and vector operation instruction. The decoding process has been introduced in detail in the previous Figure 2 、 4 description of -5, so it will not be elaborated here.

[0130] Next, the vector execution units 1212 and 1212' execute the decoded vector operation instructions according to the predicted immediate vector parameters without waiting for the decoded vector parameter setting instruction to be executed completely. The vector execution units 1212 and 1212' include vector parameter setting subunits 12121 and 12121' and vector operation subunits 12122 and 12122', which execute the above-mentioned decoded vector parameter setting instructions and vector operation instructions respectively. When the vector operation subunits 12122 and 12122' execute the decoded vector operation instructions, since the vector parameter setting subunits 12121 and 12121' have not finished execution and the vector parameter setting results are not available, the decoded vector operation instructions they execute are based on the predicted immediate vector parameters. Of course, to execute the decoded vector operation instructions, not only immediate vector parameters but also non-immediate vector parameters are required. In the embodiments of the present disclosure, for non-immediate vector parameters, the vector operation subunits 12122 and 12122' can use the non-immediate vector parameters in the previous decoded vector operation instruction of the decoded vector operation instruction as the non-immediate vector parameters in the currently received decoded vector operation instruction. The reason for using the non-immediate vector parameters in the previous decoded vector operation instruction in history as the non-immediate vector parameters in the currently decoded vector operation instruction is that vector parameters are relatively continuous in the context. Among multiple consecutive vector operation instructions, the possibility that the vector parameters remain unchanged is much greater than the possibility of change. Even if the vector parameters change, they usually remain unchanged in several subsequent vector operation instructions. Therefore, such processing greatly improves the vector operation efficiency and makes the cost for error correction not too large.

[0131] Still taking the vector parameter setting instruction vsetvli t0,x0,e32 and the vector operation instruction vadd.vv v8,v8,v4 received in the above-mentioned order as an example. Assume that the vector register bit width is 128. Before the vector operation subunits 12122 and 12122' receive the decoded vector parameter setting instruction vsetvli t0,x0,e32, the number of elements in a single operation (non-immediate vector parameter) of the previous decoded vector parameter setting instruction received is 8. After receiving the vector parameter setting instruction vsetvli t0,x0,e32, the vector parameter prediction units 129 and 129' predict that the element size in the vector is 32 bits. At this time, the vector operation subunits 12122 and 12122' will use the element size of 32 bits in the vector and the number of elements in a single operation of 8 as vector parameters and perform the vector operation of vadd.vv v8,v8,v4 under these vector parameters. Since the non-immediate vector parameter used at this time, that is, the number of elements in a single operation of 8, is historical and inaccurate and may not conform to the actual situation, correction may be required.

[0132] In one embodiment, a vector parameter register (not shown) is provided inside the vector parameter setting subunits 12121 and 12121', or a vector parameter register is provided in the register banks 126 and 126' for storing the vector parameters set as the execution results of the vector parameter setting instructions of the vector parameter setting subunits 12121 and 12121' historically. In one embodiment, this vector parameter register stores only the vector parameters obtained from one vector parameter setting instruction, that is, the vector parameters obtained from the most recent vector parameter setting instruction. Once a new vector parameter setting instruction is executed, its result is used to update this vector parameter register. In another embodiment, the vector parameters obtained from the most recent N vector parameter setting instructions can be saved, and the first-in-first-out method is followed in the vector parameter register. That is, once a new vector parameter setting instruction is executed, its result is used to overwrite the vector parameters obtained from the vector parameter setting instruction that entered the register earliest.

[0133] If the non-immediate vector parameters set after the vector parameter setting subunits 12121 and 12121' execute the decoded vector parameter setting instructions are inconsistent with the non-immediate vector parameters in the vector parameter register, this means that it is incorrect to continue using the historical non-immediate vector parameters and correction is needed. At this time, the vector parameter setting subunits 12121 and 12121' can pass the set non-immediate vector parameters to the vector operation subunits 12122 and 12122'. The vector operation subunits 12122 and 12122' discard the execution result of this vector operation instruction (that is, the result obtained by performing vector operations using the historical non-immediate vector parameters), and re-execute the decoded vector operation instruction according to the passed non-immediate vector parameters and the predicted immediate vector parameters, and use the obtained vector operation result to overwrite the operation result obtained by performing vector operations using the historical non-immediate vector parameters, thereby correcting the error caused by performing vector operations using the historical non-immediate vector parameters.

[0134] Taking the vector operation vadd.vv v8,v8,v4 as an example, where the element size of the vector predicted by the above vector operation units 12122 and 12122' and the vector parameter prediction units 129 and 129' is 32 bits, and the number of elements in the single operation in the history is 8. When the vector parameter setting units 12121 and 12121' execute the vector parameter setting instruction vsetvli t0,x0,e32, it is found that the number of elements in the single operation should be set to 4, which is inconsistent with the number of elements in the single operation stored in the vector parameter register, which is 8. At this time, the vector operation units 12122 and 12122' are notified to discard the speculative vector operation result for vadd.vv v8,v8,v4, and the vector operation units 12122 and 12122' are required to perform the vector operation again according to the element size 32 in the vector and the number of elements in the single operation 4, and replace the previous speculative vector operation result with the vector operation result.

[0135] On the other hand, in one embodiment, the vector parameter setting units 12121 and 12121' also update the vector parameter register with the vector parameters set after executing the decoded vector parameter setting instruction. In this way, the vector parameter register can always reflect the vector parameters set by the most recent vector parameter setting instruction, and can form a correct basis for determining whether the non-immediate vector parameters set by the current vector parameter setting instruction are consistent with the non-immediate vector parameters set by the previous vector parameter setting instruction in history.

[0136] In the above example, when the vector parameter setting units 12121 and 12121' execute the vector parameter setting instruction vsetvli t0,x0,e32, it is found that the element size in the vector is 32 and the number of elements in the single operation is 4, and the element size 32 in the vector and the number of elements in the single operation 4 are updated to the vector parameter register.

[0137] If the non-immediate vector parameters set by the vector parameter setting units 12121 and 12121' after executing the decoded vector parameter setting instruction are consistent with the non-immediate vector parameters in the vector parameter register, this means that it is correct to follow the non-immediate vector parameters in history and no correction is required. However, at this time, in order to make the vector parameter register always reflect the vector parameters set by the most recent vector parameter setting instruction, the vector parameter register still needs to be updated with the vector parameters set after executing the decoded vector parameter setting instruction, so as to lay a correct foundation for subsequent comparisons.

[0138] For example, when the vector parameter setting subunits 12121 and 12121' execute the vector parameter setting instruction vsetvli t0,x0,e32, it is found that the element size in the vector is 32 and the number of elements for a single operation is 8, which is consistent with the number of elements for a single operation, i.e., 8, stored in the vector parameter register. At this time, the vector operation subunits 12122 and 12122' are not notified to re - operate, but the element size 32 and the number of elements for a single operation 8 in the vector still need to be updated to the vector parameter register.

[0139] As long as the vector parameter setting subunits 12121 and 12121' finish executing the decoded vector parameter setting instruction, regardless of whether the non - immediate vector parameters set after execution are consistent with the non - immediate vector parameters in the vector parameter register, the vector parameter register needs to be updated. Besides laying a correct foundation for subsequent comparisons, it also provides a remedial measure for the speculative execution failure of the vector operation subunits 12122 and 12122'. This execution failure does not refer to the error of the non - immediate vector parameters in history, but rather the failure caused by hardware failures, network communication failures, etc. during the process of vector operation using the non - immediate vector parameters in history, and the correct execution result cannot be obtained. At this time, the vector parameter setting subunits 12121 and 12121' usually have completed setting the vector parameters and stored the set vector parameters in the vector parameter register. As long as the immediate and non - immediate vector parameters in the vector parameter register are called, the vector operation subunits 12122 and 12122' can re - execute the decoded vector operation instruction.

[0140] In addition, the vector parameter setting subunits 12122 and 12122' can check the legality of the set vector parameters after executing the decoded vector parameter setting instruction. Legality means whether the vector parameters conform to the actual hardware environment and can be implemented in the hardware environment. For example, if the element size in the vector is 256, which has exceeded the bit width of the vector register, i.e., 128, and cannot be executed in practice, it is considered illegal.

[0141] One way to check the legality of vector parameters is to determine whether the vector parameters conform to a predetermined legality standard. For example, the legality standard can be several valid values or a range of values. Only the vector parameters among these valid values or the vector parameters falling within this range of values are considered legal. For example, if the predetermined legality standard is that the element size in the vector is 32 or 64, then if the set element size in the vector is 128, it is illegal.

[0142] In one embodiment, when the set vector parameter is illegal, the vector parameter setting subunits 12122 and 12122' can set the vector parameter to be legal according to a predetermined rule. When the legality standard is a number of valid values, the predetermined rule can be: setting the vector parameter to the value among the number of valid values that is closest to the vector parameter. For example, if the predetermined legality standard is that the element size in the vector is 32 or 64, and if the element size in the set vector is 128, which is closer to 64, then the element size in the vector is set to 64. When the legality standard is a range of values, the predetermined rule can be: setting the vector parameter to the endpoint value in the range of values that is closest to the vector parameter. For example, if the predetermined legality standard is that the element size in the vector is 32 - 64, and if the element size in the set vector is 128, which is closer to the endpoint value 64, then the element size in the vector is set to 64.

[0143] This application also discloses a system on chip, including the processing units 12 and 12' as Figures 1 - 5 shown.

[0144] Vector Operation Execution Method According to Embodiments of the Present Disclosure

[0145] As Figure 6 shown, according to an embodiment of the present disclosure, there is also provided a method for executing vector operations, including:

[0146] Step 610: Sequentially obtain a vector parameter setting instruction and a vector operation instruction;

[0147] Step 620: Predict an immediate-type vector parameter according to the vector parameter setting instruction;

[0148] Step 630: Decode the retrieved vector parameter setting instruction and vector operation instruction respectively;

[0149] Step 640: Without waiting for the decoded vector parameter setting instruction to be executed, execute the decoded vector operation instruction according to the predicted immediate-type vector parameter.

[0150] The implementation details of the above method embodiments can refer to the detailed description of the previous device embodiments. Its implementation is similar to that of the device embodiments, only the description angle is different. For the sake of saving space, it will not be elaborated here.

[0151] It should be understood that the above are only the preferred embodiments of the present disclosure and are not used to limit the present disclosure. For those skilled in the art, there are many variations in the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure should be included within the protection scope of the present disclosure.

[0152] It should be understood that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for method embodiments, since they are basically similar to the methods described in the device and system embodiments, the description is relatively simple, and reference can be made to the relevant parts in other embodiments for the relevant content.

[0153] It should be understood that the above has described specific embodiments of this specification. Other embodiments are within the scope of the claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0154] It should be understood that an element described herein in the singular form or shown as only one in the figures does not represent limiting the quantity of that element to one. Additionally, a module or element described or shown herein as separate may be combined into a single module or element, and a module or element described or shown herein as a single one may be split into multiple modules or elements.

[0155] It should also be understood that the terms and expressions used herein are only for description, and one or more embodiments of this specification should not be limited to these terms and expressions. Using these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Correspondingly, the claims should be regarded as covering all such equivalents.

Claims

1. A processor, comprising: an instruction fetch unit configured to sequentially fetch vector parameter setting instructions and vector operation instructions, wherein a vector operation refers to an operation capable of concurrently generating execution results of multiple elements; a vector parameter prediction unit configured to predict immediate-type vector parameters according to the vector parameter setting instructions; an instruction decoding unit configured to separately decode the fetched vector parameter setting instructions and vector operation instructions; a vector execution unit configured to execute the decoded vector operation instructions according to the predicted immediate-type vector parameters without waiting for the decoded vector parameter setting instructions to be executed completely; wherein, the vector execution unit includes a vector operation sub-unit, and the vector operation sub-unit is configured to: use a non-immediate-type vector parameter in a previous decoded vector operation instruction of the decoded vector operation instruction received by the vector operation sub-unit as the non-immediate-type vector parameter in the received decoded vector operation instruction, and execute the decoded vector operation instruction according to the predicted immediate-type vector parameters.

2. The processor according to claim 1, wherein, The vector execution unit further includes a vector parameter setting sub-unit configured to execute the decoded vector parameter setting instructions.

3. The processor according to claim 1, wherein, The immediate-type vector parameters include element sizes in a vector, and the non-immediate-type vector parameters include the number of elements for a single operation.

4. The processor according to claim 2, wherein, The vector parameter setting sub-unit is provided with a vector parameter register, wherein if the non-immediate-type vector parameter set after the vector parameter setting sub-unit executes the decoded vector parameter setting instructions is inconsistent with the non-immediate-type vector parameter in the vector parameter register, the set non-immediate-type vector parameter is passed to the vector operation sub-unit, and the vector operation sub-unit discards the execution result of the vector operation instruction, and re-executes the decoded vector operation instruction according to the passed non-immediate-type vector parameter and the predicted immediate-type vector parameters.

5. The processor according to claim 4, wherein, The vector parameter setting sub-unit also updates the vector parameter register with the vector parameters set after executing the decoded vector parameter setting instructions.

6. The processor according to claim 4, wherein, If the non-immediate-type vector parameter set after the vector parameter setting sub-unit executes the decoded vector parameter setting instructions is consistent with the non-immediate-type vector parameter in the vector parameter register, the vector parameter register is updated with the vector parameters set after executing the decoded vector parameter setting instructions.

7. The processor according to claim 4, wherein, If the vector operation sub-unit fails to execute, the immediate-type and non-immediate-type vector parameters in the vector parameter register are called to execute the decoded vector operation instructions.

8. The processor according to claim 2, wherein After the vector parameter setting sub-unit executes the decoded vector parameter setting instructions, it checks the legality of the set vector parameters, and in the case where the set vector parameters are illegal, sets the vector parameters to be legal according to a predetermined rule.

9. The processor according to claim 1, wherein, The instruction fetch unit sequentially fetches vector parameter setting instructions and vector operation instructions from a memory external to the processor.

10. A system-on-chip, comprising the processor according to any one of claims 1-9.

11. A computing device, comprising a processor as described in any one of claims 1-9.

12. A method for performing vector operations, comprising: sequentially obtaining a vector parameter setting instruction and a vector operation instruction, wherein the vector operation refers to an operation that can generate execution results of multiple elements in parallel; predicting an immediate-type vector parameter according to the vector parameter setting instruction; decoding the retrieved vector parameter setting instruction and vector operation instruction respectively; without waiting for the decoded vector parameter setting instruction to be executed, using the non-immediate-type vector parameter in the previous decoded vector operation instruction of the decoded vector operation instruction as the non-immediate-type vector parameter in the decoded vector operation instruction, and executing the decoded vector operation instruction according to the predicted immediate-type vector parameter.

13. The method according to claim 12, wherein, The immediate-type vector parameter includes the element size in the vector, and the non-immediate-type vector parameter includes the number of elements for a single operation.

14. The method according to claim 12, wherein, A vector parameter register is pre-set. After executing the decoded vector operation instruction according to the predicted immediate-type vector parameter, the method further includes: if the non-immediate-type vector parameter set after executing the decoded vector parameter setting instruction is inconsistent with the non-immediate-type vector parameter in the vector parameter register, discarding the execution result of the vector operation instruction, and re-executing the decoded vector operation instruction according to the set non-immediate-type vector parameter and the predicted immediate-type vector parameter.

15. The method according to claim 14, wherein, After discarding the indication result of the vector operation instruction and re-executing the decoded vector operation instruction according to the set non-immediate-type vector parameter and the predicted immediate-type vector parameter, the method further includes: updating the vector parameter register with the vector parameter set after executing the decoded vector parameter setting instruction.

16. The method according to claim 14, wherein, After executing the decoded vector operation instruction according to the predicted immediate-type vector parameter, the method further includes: if the non-immediate-type vector parameter set after the vector parameter setting sub-unit executes the decoded vector parameter setting instruction is consistent with the non-immediate-type vector parameter in the vector parameter register, updating the vector parameter register with the vector parameter set after executing the decoded vector parameter setting instruction.

17. The method according to claim 14, wherein, After executing the decoded vector operation instruction according to the predicted immediate-type vector parameter, the method further includes: if the vector operation sub-unit fails to execute, invoking the immediate-type and non-immediate-type vector parameters in the vector parameter register and executing the decoded vector operation instruction.

18. The method according to claim 12, wherein After decoding the retrieved vector parameter setting instruction and vector operation instruction, the method further includes: executing the decoded vector parameter setting instruction, checking the legality of the vector parameter set after executing the vector parameter setting instruction, and in the case where the set vector parameter is illegal, setting the vector parameter to be legal according to a predetermined rule.

19. The method according to claim 12, wherein, The sequentially obtaining a vector parameter setting instruction and a vector operation instruction includes: sequentially obtaining a vector parameter setting instruction and a vector operation instruction from a memory external to the processor.

Citation Information

Patent Citations

  • A four-stage assembly line RISC-V processor with a rapid data bypass structure

    CN109918130A

  • Systems, methods, and apparatuses for heterogeneous computing

    CN110121698A