Vector processor, high performance processor, and electronic device

By designing the interactive mechanism of vector program control unit, functional unit and matrix register stack of vector processors, the efficiency problem of vector processor processing large amounts of data in scientific computing, graphics processing and artificial intelligence technologies is solved, and efficient vector data processing is achieved.

CN120540709AActive Publication Date: 2025-08-26SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Patent Information

Application Number
CN202510574548.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-26
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The need for existing vector processors to be difficult to efficiently process large amounts of data in scientific computing, graphics processing and artificial intelligence technologies is not met.

Method used

A vector processor is designed, including a vector program control unit, multiple functional units, a matrix register stack and a scalar register. The vector program control unit fetches fingers and transmits instructions, and interacts with the scalar registers. The functional unit performs functional processing, the matrix register stack receives read and write requests and returns data, and interacts with the functional unit to configure the registers of the vector program control unit.

Benefits of technology

It realizes efficient processing of vector data, and improves the computing power and processing efficiency of vector processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540709A_ABST
    Figure CN120540709A_ABST
Patent Text Reader

Abstract

The invention provides a vector processor, a high-performance processor and electronic equipment. The vector processor comprises a vector program control unit, a plurality of functional units, a matrix register file and a scalar register, wherein the vector program control unit is used for fetching instructions and transmitting the instructions; the vector program control unit interacts with the scalar register; the function unit is used for performing function processing according to the instruction; the matrix register file is used for returning data after receiving the read-write request; rearranging the data and then returning the data; performing read-write interaction with the functional unit; and configuring a configuration register of the vector program control unit through data in the matrix register file. The vector processor provided by the invention can efficiently process the vector data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a vector processor, a high-performance processor, and an electronic device. Background Art

[0002] Vector processors can execute various types of instructions, including arithmetic operations, logical operations, data transfer, etc. Vector processors have super computing power and can handle various data types and tasks.

[0003] With the development of scientific computing, graphics processing, and artificial intelligence technologies, the rapid processing of large amounts of data through vector processors has become a common demand, which places higher demands on vector processors.

[0004] Therefore, a vector processor is needed that can process vector data efficiently. Summary of the Invention

[0005] In order to solve one of the above technical defects, the present application provides a vector processor, a high-performance processor and an electronic device.

[0006] In a first aspect of the present application, a vector processor is provided, the vector processor comprising: a vector program control unit, a plurality of function units, a matrix register file and a scalar register;

[0007] Among them, the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register;

[0008] Functional unit, used for performing functional processing according to instructions;

[0009] The matrix register stack is used to receive read and write requests and return data; rearrange the data and return it; perform read and write interactions with the functional units; and configure the configuration registers of the vector program control unit through the data in the matrix register stack.

[0010] Optionally, the vector program control unit is configured to, after fetching the instruction, determine whether to execute it, and transmit the instruction to the functional unit based on the determination result;

[0011] Vector program control unit, also used to control instruction jumps;

[0012] The vector program control unit has scalar computing capabilities.

[0013] Optionally, the functional unit includes: one or more vector operation units, one or more vector interleaving units, and one or more vector access units;

[0014] Any vector operation unit, used to perform vector operations according to instructions;

[0015] Any vector interleaving unit, used to perform data interleaving and logic processing according to instructions;

[0016] Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.

[0017] Optionally, any vector operation unit includes: a floating-point multiplication-addition operator unit, a floating-point multiplication-accumulation operator unit, a floating-point arithmetic operator unit, a tensor multiplication subunit and an intermediate result register;

[0018] The floating-point multiplication-addition operator unit, the floating-point multiplication-accumulation operator unit, the floating-point arithmetic operator unit, and the tensor multiplication subunit share the intermediate result register;

[0019] The floating-point multiplication-addition operator unit and the floating-point arithmetic operator unit share a single launch slot;

[0020] The floating-point multiply-accumulate operator and the tensor multiplication subunit share one emit slot.

[0021] Optionally, the vector processor further includes: a private vector register of the vector interleaving unit and a private vector register of the vector access unit;

[0022] Wherein, the private vector register of the vector interleaving unit corresponds to the vector interleaving unit one-to-one;

[0023] The private vector registers of the vector access unit are shared by multiple vector access units.

[0024] Optionally, the matrix register stack is also used to write data of other operation processors; and receive status information sent by other operation processors to query whether the data has been written.

[0025] Optionally, the vector program control unit is further configured to receive a start command sent by other operation processors to start the vector processor; and return an indication signal indicating whether the vector processor has ended to the other operation processors.

[0026] Optionally, a read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and other arithmetic processors;

[0027] The vector program control unit and other operation processors both perform read operations on the read FIFO unit and write operations on the write FIFO unit;

[0028] Other arithmetic processors perform read operations or write operations on the scalar registers.

[0029] A second aspect of the present application provides a high-performance processor, comprising: the vector processor and the scalar processor as described in the first aspect;

[0030] The matrix register stack is used to write data to the scalar processor; it receives status information from the scalar processor to inquire whether the data has been written;

[0031] The vector program control unit is used to receive a start command sent by the scalar processor, start the vector processor, and return an indication signal to the scalar processor indicating whether the vector processor has ended;

[0032] A read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and the scalar processor;

[0033] The vector program control unit and the scalar processor both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit;

[0034] The scalar processor reads or writes scalar registers.

[0035] In a third aspect of the present application, an electronic device is provided, comprising: a high-performance processor as described in the second aspect; or comprising one or more processor clusters, wherein each processor cluster includes multiple high-performance processors as described in the second aspect.

[0036] The present application provides a vector processor, a high-performance processor, and an electronic device. The vector processor includes: a vector program control unit, multiple functional units, a matrix register stack, and a scalar register. The vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register; the functional unit is used to perform functional processing according to the instruction; the matrix register stack is used to receive read and write requests and return data; rearrange the data before returning it; perform read and write interactions with the functional units; and the configuration registers of the vector program control unit are configured using data in the matrix register stack. The vector processor provided by the present application can efficiently process vector data. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0038] Figure 1 A schematic diagram of the architecture of a vector processor provided in an embodiment of the present application;

[0039] Figure 2 A schematic diagram of the structure of a vector operation unit provided in an embodiment of the present application;

[0040] Figure 3 A schematic diagram of the architecture of another vector processor provided in an embodiment of the present application;

[0041] Figure 4 A schematic diagram of the structure of a high-performance processor provided in an embodiment of the present application;

[0042] Figure 5 A schematic diagram of the structure of a scalar processor provided in an embodiment of the present application;

[0043] Figure 6 A schematic structural diagram of a synchronization unit of a scalar processor provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.

[0045] During the development of this application, the inventors discovered that with the advancement of scientific computing, graphics processing, and artificial intelligence technologies, the rapid processing of large amounts of data using vector processors has become a common need, placing high demands on vector processors. Therefore, a vector processor is needed that can efficiently process vector data.

[0046] To address the above issues, embodiments of the present application provide a vector processor, a high-performance processor, and an electronic device. The vector processor includes: a vector program control unit, multiple functional units, a matrix register stack, and a scalar register. The vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register; the functional unit is used to perform functional processing according to the instruction; the matrix register stack is used to receive read and write requests and return data; rearrange the data and return it; perform read and write interactions with the functional units; and the configuration registers of the vector program control unit are configured using data in the matrix register stack. The vector processor provided by the present application can efficiently process vector data.

[0047] See also Figure 1 This embodiment provides a vector processor, which includes: a vector program control unit, multiple function units, a register file and a scalar register.

[0048] In addition, the vector processor also includes: a private vector register of a vector interleaving unit and a private vector register of a vector access unit.

[0049] 1. Vector program control unit

[0050] Vector program control unit, used for instruction fetching and instruction issuance.

[0051] That is, the vector program control unit is used to fetch instructions, determine whether to execute them, and send the instructions to the functional units based on the determination result.

[0052] The vector program control unit is also used to control instruction jumps.

[0053] The vector program control unit has scalar computing capabilities.

[0054] The vector program control unit interacts with the scalar registers.

[0055] In specific implementation, the vector program control unit is an instruction fetch and instruction issuance unit. It takes instructions from the cache according to the PC value, and after determining whether to execute them, it issues the instructions to each functional unit according to the wait value (configured by the wait instruction). At the same time, it controls the jump of instructions and has some scalar computing capabilities.

[0056] In addition, the vector program control unit is further configured to receive a start command from other operation processors to start the vector processor and return an indication signal to the other operation processors indicating whether the vector processor has finished operation.

[0057] Taking the scalar processor as an example, the vector program control unit receives the start command issued by the synchronization unit of the scalar processor, starts the vector processor execution, and also returns an indication signal indicating whether the synchronization unit vector processor execution is completed.

[0058] 2. Functional Unit

[0059] Functional unit, used to perform functional processing according to instructions.

[0060] For example, the functional unit receives an instruction from the vector program control unit, processes data accordingly according to the instruction, and outputs the processing result according to the address specified in the instruction.

[0061] The functional units include: one or more vector operation units, one or more vector interleaving units, and one or more vector access units.

[0062] 1) Vector operation unit

[0063] Any vector arithmetic unit, used to perform vector operations according to instructions.

[0064] like Figure 2 As shown, any vector operation unit includes: a floating-point multiplication-addition operator unit, a floating-point multiplication-accumulation operator unit, a floating-point arithmetic operator unit, a tensor multiplication subunit and an intermediate result register.

[0065] The floating-point multiplication and addition unit and the floating-point arithmetic unit share a single issue slot, so a maximum of eight vector unit instructions can be issued per cycle.

[0066] The floating-point multiply-accumulate operator and the tensor multiplication subunit share one emit slot.

[0067] The floating-point multiply-add operator unit is a functional unit that executes floating-point multiply-add operator unit related instructions. For example, floating-point multiply-add operator unit related instructions include integer and floating-point vector multiplication and accumulation, multiplication, addition, tensor calculation, etc.

[0068] 1 vector operation unit has independent intermediate result registers.

[0069] One floating-point multiply-add operator unit, one floating-point multiply-accumulate operator unit, one tensor multiplication subunit, and one floating-point arithmetic operator unit share intermediate result registers.

[0070] (1) Floating-point multiplication and addition subunits and floating-point multiplication and accumulation subunits, which can perform integer and floating-point vector multiplication, multiplication and accumulation, etc. Supported types include but are not limited to int32, fp32, and fp64.

[0071] (2) The floating-point arithmetic subunit can perform integer and floating-point vector arithmetic operations, such as comparison, addition, subtraction, bitwise operations, etc. Supported types include but are not limited to int8, uint8, int16, uint16, int32, uint32, bool, fp16, bf16, fp32, tf32, and fp64.

[0072] (3) The tensor multiplication subunit can perform tensor multiplication, multiply-accumulate and other operations. Supported types include but are not limited to int8, bf16, fp16, and tf32.

[0073] 2) Vector interleaving unit

[0074] Any vector interleaving unit is used to perform data interleaving and logic processing according to instructions.

[0075] The vector interleaver unit is the control and data processing unit within the vector processor. It is responsible for interleaving data and supports logical and some fixed-point and floating-point calculations. It also supports a wide range of customized instructions, including table lookup, lateral calculations, sparse matrix calculations, precision conversion, and FIFO (First Input First Output). It also executes instructions such as data broadcasting, extraction, and internal interleaving.

[0076] Each vector interleaving unit has a set of private vector registers, so the private vector registers of the vector interleaving units correspond one to one with the vector interleaving units.

[0077] 3) Vector access unit

[0078] Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.

[0079] The vector access unit is a memory access unit within a vector processor, primarily responsible for reading / writing instructions and various scalar calculations.

[0080] The read instruction / write instruction supports multiple memory access modes, such as row mode, column mode / discrete mode / extended mode / accumulated mode.

[0081] It also supports multiple parameter configurations, with a maximum read / write instruction data width of 1024 bits. It can also perform address calculation, load / store and other instructions.

[0082] All vector access units share a set of private vector registers, so the private vector registers of a vector access unit are shared by multiple vector access units.

[0083] 3. Register file stack

[0084] The register file receives read and write requests and returns data. It rearranges the data before returning it. It interacts with the functional units for read and write operations. The vector program control unit's configuration registers are configured using data in the register file.

[0085] The register file is a general-purpose vector register stack and is the main storage unit within the vector processor. It is responsible for receiving read and write requests and returning data. In some functions, it can rearrange the data before returning it to the request module.

[0086] The register file stack performs read and write interactions with the functional units in the vector processor (such as the floating-point multiplication-addition operator unit, the floating-point arithmetic operator unit, the floating-point multiplication-accumulation operator unit, and the tensor multiplication subunit), and supports the use of data in the register file stack to configure the instruction fetch unit configuration register.

[0087] The register file stack is also used to write data to other processors and receive status information from other processors to check whether the data has been written.

[0088] Taking a scalar processor as an example, the synchronization unit of the scalar processor can write data into the register file stack, and the register file stack can also receive status information from the synchronization unit of the scalar processor to inquire whether the data has been written.

[0089] The depth of the register file is configurable.

[0090] Figure 3 A schematic diagram of a vector processor is shown, in which the functional units include four vector operation units, four vector interleaving units, and four vector access units.

[0091] The vector processor provided in this embodiment supports a VLIW (Very Long Instruction Word) instruction set. Each VLIW may be composed of one or more instructions, and each instruction corresponds to a functional unit.

[0092] In addition, a read FIFO unit and a write FIFO unit are provided between the vector processor and other operation processors.

[0093] The vector program control unit and other operation processors both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit.

[0094] Other arithmetic processors perform read operations or write operations on the scalar registers.

[0095] Taking a scalar processor as an example, a read FIFO and a write FIFO unit for transmitting data are provided between the scalar processor and the vector processor. The scalar processor and the vector program control unit can perform read operations or write operations on the read and write FIFO.

[0096] At the same time, the synchronization unit of the scalar processor can read or write the scalar registers of the vector processor.

[0097] This embodiment provides a vector processor, comprising: a vector program control unit, multiple functional units, a matrix register stack, and scalar registers; wherein the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar registers; the functional units are used for performing functional processing according to instructions; the matrix register stack is used for receiving read and write requests and returning data; rearranging the data before returning it; performing read and write interactions with the functional units; and configuring the configuration registers of the vector program control unit using data in the matrix register stack. The vector processor provided by this embodiment can efficiently process vector data.

[0098] Based on the same inventive concept of a vector processor, this embodiment provides a high-performance processor, which includes a scalar processor and a vector processor.

[0099] The connection relationship between the scalar processor and the vector processor can be shown as follows Figure 4 shown.

[0100] The scalar and vector processors share data storage. The vector processors can only access the data storage and are executed only by the scalar processors.

[0101] A connection is established between the scalar processor and the vector processor. For example, the scalar processor and the vector processor are connected via a dedicated instruction channel.

[0102] In addition, the high-performance processor may also include two registers, one register corresponding to the scalar processor and the other register corresponding to the vector processor. The vector processor can read and write its corresponding register, and the scalar processor can read and write its corresponding register as well as the vector processor's corresponding register.

[0103] The scalar processor can read and write the registers of the vector processor.

[0104] Scalar processors establish connections to global memory.

[0105] (1) Scalar processor

[0106] See also Figure 5 The scalar processor includes: an instruction fetch unit, a register renaming unit, an operation reservation stack unit, a storage reservation stack unit, a scalar operation unit, a memory access unit, a program control unit, a synchronization unit, a pipeline control unit, a register file unit, and a special vector register file unit.

[0107] In addition, the scalar processor may also include one or more other units, such as one or more other functional modules, one or more instruction caches, one or more data storages, one or more special vector registers, one or more status flag registers, etc.

[0108] 1. Instruction fetch unit

[0109] The instruction fetch unit is used to fetch and dispatch instructions.

[0110] Specifically, the instruction fetch unit generates an instruction fetch request address, outputs the fetch request address to the instruction cache for instruction fetching, receives instructions from the instruction cache, and stores them in the data store. Each cycle, it sequentially reads qualified instructions from the data store, decodes and performs relevant checks on the read instructions, and then dispatches the checked instructions sequentially.

[0111] For example, the instruction fetch unit generates an instruction fetch request address and outputs it to the instruction cache for instruction fetch, and receives instructions from the instruction cache and stores them in the data storage. In each cycle, it sequentially selects one or more instructions from the qualified instructions for decoding and related checks, and dispatches the qualified instructions sequentially, dispatching at most one program control unit instruction and one synchronization unit instruction at a time. In addition, one or more scalar operation unit instructions and one or more memory access unit instructions can also be dispatched each time.

[0112] 2. Register renaming unit

[0113] The register renaming unit is used to receive instructions dispatched by the instruction fetch unit and rename registers.

[0114] Specifically, the register renaming unit receives and stores instructions dispatched by the instruction fetch unit, renames special vector registers, conditionally decodes instructions, and generates pipeline stall signals. It also receives and writes data from one or more of the scalar arithmetic unit, memory access unit, program control unit, synchronization unit, special vector registers, condition registers, and flag registers. It also sends instructions to one or more of the arithmetic hold stack unit, memory hold stack unit, program control unit, and synchronization unit.

[0115] For example, the register renaming unit is used in a scalar processor to receive instructions dispatched by the instruction fetch unit and perform register and special vector register renaming, instruction conditional decoding, and generate pipeline congestion signals. At the same time, it receives data from the execution unit (such as the scalar operation unit, memory access unit, program control unit, synchronization unit) to write back registers, special vector registers, condition registers, and status flag registers and writes them back to the corresponding registers.

[0116] The scalar processor write-back supports out-of-order write-back, has high execution efficiency, and distributes instructions to the operation retention stack unit, the storage retention stack unit, the program control unit, or the synchronization unit.

[0117] The register renaming unit bandwidth may be 6 bits, wherein multiple (eg, 4) input instructions may be valid at the same time.

[0118] There can be multiple condition registers, which are located in the register renaming unit.

[0119] The instructions of the scalar operation unit and memory access unit support the operations of reading and writing condition registers.

[0120] The synchronous unit's instructions support the operation of reading the condition register.

[0121] The jump and function call instructions of the program control unit support the operation of reading the condition register.

[0122] When an instruction enters the condition register, if there is an unexecuted instruction in the condition register, the pipeline is blocked.

[0123] That is, the condition register is not renamed, and when a read-write dependency occurs, a dispatch block is triggered to wait. The read and write rules for the condition register are as follows:

[0124] ●Read the rules:

[0125] (1) All instructions of the scalar arithmetic unit, memory access unit, and synchronization unit support conditional execution and require reading the value of the condition register.

[0126] (2) The scalar arithmetic unit also supports read conditional register instruction operations.

[0127] (3) The jump and function call instructions of the program control unit support read condition register operations.

[0128] ●Write rules:

[0129] (1) The scalar arithmetic unit supports write conditional register instructions.

[0130] (2) Scalar operation unit logic instructions and comparison instructions support the option of writing condition registers.

[0131] When the previously issued instruction to write the condition register has not yet been completed, and an instruction to read or write the same condition register enters, the pipeline is blocked and a conditional execution blocking signal is generated, waiting for the previous condition register to be written.

[0132] In addition, the register renaming unit includes: one or more physical registers and one or more logical registers.

[0133] Wherein, any physical register is one of the following: a scalar physical register, a vector physical register, a condition register, and a flag register.

[0134] Any logical register is one of the following: a scalar logical register, a vector logical register.

[0135] For example, the register renaming unit includes one or more physical registers, such as a plurality of 512-bit wide special vector registers, a plurality of condition registers, and a status flag register.

[0136] Among them, special vector registers are renamed, while condition registers and status flag registers are not renamed.

[0137] There are multiple logical registers, such as multiple logical registers including scalar logical registers and multiple vector logical registers.

[0138] In addition, the mapping relationship between logical registers and physical registers is maintained by the register mapping table. The mapping relationship between vector logical registers and vector physical registers is maintained by the special vector register mapping table.

[0139] 1) Register Mapping Table

[0140] Initially, the physical registers mapped to the entries corresponding to all logical register indices in the register mapping table are all 0. When an instruction is executed or an interrupt occurs, the logical registers allocated to the relevant physical registers are determined, and the mappings of the entries corresponding to the allocated logical register indices in the register mapping table are updated to the identifiers of the relevant physical registers.

[0141] For example, a register map table with a depth of 32 bits and a width of 6 bits stores the mapping between all logical registers and all physical registers. Initially, the mapping in the register map table is invalid, and all entries for the mapped physical registers are set to zero. When a physical register is assigned to a logical register, the entry corresponding to the logical register index in the register map table is changed to the ID of the physical register.

[0142] It should be noted that the register mapping table is updated only when the instruction is actually executed. If the conditional execution instruction is not executed, the register mapping table will not be updated. In addition, the register mapping table will not be updated when a jump occurs. However, when an interrupt occurs, the interrupt return address must update the register mapping table to ensure that the interrupt can return normally.

[0143] 2) Special vector register mapping table

[0144] Initially, the mapping vector physical registers of all entries corresponding to the vector logical register indexes in the special vector register mapping table are all 0. When an instruction is executed, the vector logical register allocated to the relevant vector physical register is determined, and the mapping of the entry corresponding to the allocated vector logical register index in the special vector register mapping table is updated to the identifier of the relevant vector physical register.

[0145] For example, the special vector register mapping table has a depth of 4 and a width of 3 bits, storing the mapping relationship between all vector logical registers and all vector physical registers. Initially, the special vector register mapping table is invalid, and all entries for the mapped vector physical registers are all zeros. When a vector physical register is assigned to a vector logical register, the entry corresponding to the vector logical register index in the special vector register mapping table is changed to the ID of the vector physical register.

[0146] It should be noted that the special vector register mapping table is updated only when the instruction is actually executed. If the conditional execution instruction is not executed, the special vector register mapping table will not be updated. In addition, the special vector register mapping table will not be updated when a jump occurs.

[0147] 3. Operation reservation stack unit

[0148] The operation reserve stack unit is the emission queue of the scalar operation unit.

[0149] The operation reserve stack unit receives instructions, dispatch and renaming information from the register renaming unit and pushes them into the queue. It then pops ready instructions to the scalar operation unit for execution.

[0150] The operation reserve stack unit is also used to decode input instructions and store instruction type information.

[0151] In other words, the Arithmetic Hold Stack unit acts as the issue queue for the scalar arithmetic unit. It receives instructions and associated dispatch and renaming information from the register renaming unit, pushes them into the queue, and then pops ready instructions onto the scalar arithmetic unit for execution. The Arithmetic Hold Stack unit decodes the incoming instructions and stores the instruction type information.

[0152] In specific implementation, the depth of the operation reserve stack unit can be flexibly adjusted, for example, the depth of the operation reserve stack unit is 8. Multiple scalar operation units share one operation reserve stack unit.

[0153] The rules for issuing and receiving instructions for the operation reserve stack unit are as follows:

[0154] (1) The output of the register renaming unit enters the operation preservation stack unit.

[0155] (2) When there is any idle scalar operation unit, it will fetch instructions and operands from the operation reservation stack unit for execution.

[0156] (3) The principle of executing instructions from the operation reservation stack unit is to execute the executable instructions that can be sent from the operation reservation stack unit in a forward-to-back order.

[0157] (4) Whether the transmission can be made is determined by whether the values ​​of all source registers or special vector registers or condition registers and status flag registers are ready.

[0158] (5) If there are multiple instructions that can be sent, the oldest instruction will be sent first according to the instruction order.

[0159] (6) If any scalar operation unit is blocked, it can no longer receive new instructions.

[0160] (7) If the instruction previously sent to any scalar operation unit is a division instruction, a new division instruction can only be sent to it after the division result calculation is completed and the calculation completion En signal is returned.

[0161] 4. Storage Retention Stack Unit

[0162] The storage reservation stack unit is the transmit queue of the memory access unit.

[0163] The storage reservation stack unit is used to receive instructions and register renaming information from the register renaming unit and push them into the queue.

[0164] The storage reservation stack unit is also used to send a read request to the register renaming unit when the instruction address register is ready, and save the read address operand.

[0165] The register renaming unit is also used to calculate the address, decode the address, and save the decoding information after the instruction obtains the address.

[0166] The register renaming unit is also used to detect when the source register of any instruction is ready and the address decoding is completed, and then send it to the memory access unit for execution.

[0167] In specific implementation, the depth of the storage reservation stack unit can be flexibly adjusted, such as the depth of the storage reservation stack unit is 16. Multiple memory access units share one storage reservation stack unit. The storage reservation stack unit is the transmission queue of the memory access unit. The storage reservation stack unit receives instructions and register renaming information from the register renaming unit and pushes them into the queue. When the instruction address register in the storage reservation stack unit is ready, a read request is sent to the register renaming unit and the read address operand is saved in the queue. After the instruction in the storage reservation stack unit obtains the address, the address can be calculated and decoded, and the generated decoding information is saved in the queue. When the source register of an instruction (such as a write instruction) in the storage reservation stack unit is ready and the address decoding is completed, it can be transmitted to the memory access unit for execution. Before transmission, it must undergo a series of checks, such as address type check, address comparison check, and address forward check.

[0168] The rules for sending and receiving instructions to the storage reserve stack unit are as follows:

[0169] (1) The output of the register renaming unit enters the storage reservation stack unit.

[0170] (2) When the source operand for calculating the address is ready, the memory access address is calculated and saved in the storage reservation stack unit.

[0171] (3) Instructions with unrelated addresses: They can be sent out of order. The out-of-order rules are: read instructions after read instructions, write instructions after read instructions, and read instructions after write instructions. They can all be sent out of order. Write instructions after write instructions need to maintain order (they cannot be sent to different memory access units at the same time). Even if the addresses are unrelated, write instructions after write instructions still need to maintain order.

[0172] (4) Address-related instructions: read instructions followed by write instructions, write instructions followed by read instructions, write instructions followed by write instructions, and read instructions followed by read instructions all need to be performed in order.

[0173] (5) When the addresses are unrelated but all instructions that have not been sent successfully (i.e., instructions on the way that have not been sent to the destination, including those at the memory access unit level and the memory access unit output level) are located in the same storage space, they can be sent out of order to the same memory access unit, but they cannot be sent to two or more memory access units.

[0174] (6) Only one memory access instruction located in the same storage space but with unrelated addresses can be sent at the same time, and two or more memory access units cannot be sent at the same time.

[0175] (7) Address correlation judgment principle: Whether the addresses are related is irrelevant if they are located in different storage spaces. If they are located in the same storage space, whether the addresses are related is determined based on the data granularity.

[0176] 5. Scalar arithmetic unit

[0177] In a specific implementation, there may be one or more scalar operation units.

[0178] For example, the scalar processor includes two scalar arithmetic units, namely scalar arithmetic unit 0 and scalar arithmetic unit 1.

[0179] The scalar operation unit is used to receive instructions and data sent by the operation reservation stack unit, perform operations on the data based on the instructions, and write the operation results back to the register renaming unit.

[0180] The scalar arithmetic unit is the computing unit of the scalar processor, which can perform various types of fixed-point and floating-point operations, such as addition, subtraction, multiplication, division, logical operations, comparison operations, shifts, etc. It receives instructions and data sent by the operation reserve stack unit, performs operations, and writes the results back to the register file unit of the register renaming unit or the special vector register file unit.

[0181] Several instruction examples are provided below as examples. In specific implementation, they are not limited to the following instructions, nor are they limited to including all instructions.

[0182] Instructions with execution level one include: fixed-point addition and subtraction, logical instructions, shift instructions, fixed-point and floating-point comparison instructions, read and write Flag instructions, fixed-point maximum and minimum instructions, ABS instructions, bit reversal instructions, selection instructions, special vector register distribution instructions, read special vector register instructions, Byte reversal instructions, Merge instructions, immediate value assignment instructions, FirstOne instructions, CRC instructions, floating-point classification instructions, floating-point partial domain extraction, and Rounding instructions.

[0183] Instructions with execution level three include: fixed-point multiplication instructions, fixed-floating-point conversion instructions, bit filtering instructions, Count instructions, and floating-point addition and subtraction instructions.

[0184] Instructions that support Bypass include: selection instructions, fixed-point addition and subtraction instructions, shift instructions, immediate value assignment instructions, ABS instructions, logical instructions, comparison instructions, and maximum and minimum instructions.

[0185] The execution cycle of the division instruction is uncertain and is related to the data of the divisor and the dividend. When the instruction is executed, a DivEn instruction will be generated to indicate that the instruction is executed and the result is output to the register stack. No new division instructions can be input during the execution of the division instruction, but other scalar calculation unit instructions can be input. After the division is executed, the output result is reused with the output port of the first-stage pipeline. When the output port of the first-stage pipeline is not used by other scalar calculation unit instructions, the division outputs its result and outputs the DivEn identifier at the same time. The DivEn identifier is output to the operation retention stack unit, indicating that the Div instruction can continue to be output to the current scalar calculation unit.

[0186] 6. Memory access unit

[0187] In a specific implementation, there may be one or more memory access units.

[0188] For example, a scalar processor includes two memory access units, namely memory access unit 0 and memory access unit 1.

[0189] The memory access unit is used to receive instructions, data and register information sent by the storage reservation stack unit, and read and write the data based on the instructions and register information.

[0190] The memory access unit is a functional module that executes memory access-related instructions for scalar processors. The memory access unit receives instructions and data, as well as register-related information, from the storage retention stack unit. It executes instructions accordingly and interacts with other units to read and write data. Read instructions and atomic write instructions require writing data back to the register renaming unit. These instructions include register-level read and write instructions, including 8-bit, 16-bit, 32-bit, 64-bit, or other bit granularities, as well as vector read and write instructions, including 128-bit, 256-bit, 512-bit, or other bit granularities. Different instructions have different processing cycles.

[0191] In addition, the memory access unit is responsible for providing the relevant instruction quantity information required by FENCE. The memory access unit interacts with the storage reservation stack unit to complete the data storage configuration.

[0192] 7. Program control unit

[0193] In a specific implementation, there is only one program control unit.

[0194] The program control unit is configured to receive instructions and data from the register renaming unit, process the data based on the instructions, and output a processing result.

[0195] The program control unit (PCU) executes instructions related to the scalar processor's program execution sequence. The PCU receives instructions and data from the register renaming unit (RRU), processes the data accordingly, and outputs the results to other modules in the scalar processor. Different instructions are processed in different time periods.

[0196] The program control unit is responsible for controlling the direction of program execution (such as stop, interrupt, jump, function call), involving the execution of related instructions and the reading and writing control of configuration information; the program control unit is responsible for the configuration and prefetch operations of the instruction cache, as well as FENCE operations; the program control unit is responsible for the reading, writing and control of the counter, as well as the reading and writing of some other control information.

[0197] 8. Synchronization unit

[0198] In a specific implementation, there is only one synchronization unit.

[0199] Synchronization unit, used to synchronize the scalar processor and the vector processor.

[0200] like Figure 6 As shown, a communication connection is established between the synchronization unit and the pipeline control unit, the register renaming unit, the program control unit, and the vector processor.

[0201] The instructions of the synchronization unit come from the register renaming unit, and the reading and writing of the data of the synchronization unit interact with the register renaming unit.

[0202] The synchronization unit is used to receive the pause signal sent by the pipeline control unit and send the execution level pause signal generated when communicating with the vector processor to the pipeline control unit so as to generate the execution pause signal of the scalar processor.

[0203] The synchronization unit is used to generate instructions and transmit them to the program control unit.

[0204] That is to say, the synchronization unit is a unit that synchronizes the scalar processor and the vector processor. It receives instructions and data sent by the register renaming unit, reads data from the vector processor and writes it back to the register stack, reads data from the register stack unit or the special vector register stack unit and sends it to the functional module of the vector processor. It is responsible for the startup and status query of the vector processor, such as querying the reading and writing of the read and write FIFO (FirstInput FirstOutput) in the vector program control unit of the vector processor, the configuration of the register file stack, the reading or writing of the scalar register, the register file stack status query, the reading FIFO depth, the reading of the startup vector processor instruction counter, etc., and providing the program control unit with the synchronization unit instruction information.

[0205] The synchronization unit interacts with the pipeline control unit, register renaming unit, and program control unit within the scalar processor, as well as with the external vector processor, scalar processor, and vector processor transfer queue module. Synchronization unit instructions originate from the register renaming unit, and data reading and writing must interact with the register renaming unit. The synchronization unit receives a stall signal from the pipeline control unit and, when communicating with the vector processor, generates its own execute-level stall signal, which is sent to the pipeline control unit to generate the ExeStall signal for the entire scalar processor. The synchronization unit generates the instruction to be executed in the next cycle and transmits it to the program control unit for use by the program control unit's counter instruction. Interactions with the vector processor include, but are not limited to, configuring the register file with special vector registers or registers, reading and writing scalar registers, and querying the write status of the register file. Interactions with the scalar processor and vector processor transfer queue module include, but are not limited to, starting the vector processor, querying vector processor status, reading and writing data in the vector processor's instruction fetch unit FIFO, reading the FIFO depth, and reading the start vector processor instruction counter.

[0206] Therefore, in a specific implementation, the synchronization unit may have the following functions (it should be noted that the following functions are only examples, and other functions may be provided. This embodiment and subsequent embodiments do not limit the specific functions of the synchronization unit):

[0207] The start vector processor function is used to start the vector processor, including immediate start and register start, such as the pipeline waits until the start is successful, or writes the result of the start success or failure back to the destination register.

[0208] Query the vector processor execution status function, support option B.

[0209] Read and write FIFO function, the FIFO is located in the instruction fetch unit of the vector processor, such as the FIFO bit width 32 bits, the read and write FIFO such as the read and write FIFO waits until success, or the read and write FIFO success or failure result is written back to the register.

[0210] Write register file stack functions, including special vector register writes or register writes.

[0211] Read and write scalar register functions, including immediate index or register index read and write.

[0212] Query the register file stack write back status function, such as waiting until all writes to the register file stack are completed, or returning the result of whether the write to the register file stack is completed to the register.

[0213] When the related operations are not completed, the synchronization unit will generate its own blocking signal, blocking and waiting, and the signal will be sent to the pipeline control unit to generate a pipeline blocking signal.

[0214] A FIFO (such as a 32-bit deep FIFO) can also be added between the scalar processor and the vector processor to store the request to start the vector processor, and move the read and write FIFO previously located in the vector processor to the scalar processor and vector processor transmission queue module. The scalar processor and vector processor transmission queue module unit implements the startup of the vector processor, queries the execution status of the vector processor, reads and writes FIFO functions, reads FIFO depth functions, and reads the startup vector processor instruction counter function. The condition for the successful startup of the vector processor is that the startup vector processor FIFO is not full, and the query of the vector processor execution status is passed. The condition for the vector processor status to be stopped is that the vector processor execution is completed and the startup vector processor FIFO is empty.

[0215] 9. Assembly line control unit

[0216] The pipeline control unit is used to generate a pipeline pause signal and / or generate a start and stop signal for the scalar processor.

[0217] The pipeline control unit is the pipeline control unit of the scalar processor, which is connected to each unit inside the scalar processor and is responsible for generating pipeline blocking signals, such as blocking in normal working mode and blocking in debug mode.

[0218] The pipeline control unit also communicates with the communication and synchronization unit to generate signals for starting and stopping the scalar processor.

[0219] Furthermore, in practical applications, scalar processors can also perform conditional execution decoding. For example, when performing conditional execution decoding, a scalar processor determines the execution condition of the instruction's preset bits. If the condition is met, a valid instruction is output; otherwise, a null instruction is output. A null instruction refers to an empty or invalid instruction.

[0220] If there are read or write operations on the condition register, the pipeline is blocked and the read operation can only be performed after the write operation is completed. There is no bypass when reading or writing the condition register.

[0221] Taking two condition registers, namely condition register 0 and condition register 1, and the preset bits being [29:28] as an example, when the scalar processor performs conditional execution decoding, the scalar processor judges the execution condition of the input instruction based on the [29:28] bits of the instruction set encoding. If the condition is met, a valid instruction is output, otherwise an empty instruction is output.

[0222] Among them, [29:28] bits are 00, indicating that condition register 0 is 1 and execution is executed, [29:28] bits are 01, indicating that condition register 1 is 1 and execution is executed, [29:28] bits are 10, indicating that condition register 0 is executed, and [29:28] bits are 11, indicating unconditional execution. If the condition is not met, the instruction is invalid and a null instruction is output.

[0223] If the condition register is read or written, the pipeline is blocked and the read operation is performed after the condition register is written. There is no bypass when reading or writing the condition register.

[0224] (2) Vector Processor

[0225] The vector processor can be Figure 1 or Figure 3 The vector processor shown.

[0226] Specifically, the vector processor includes: a vector program control unit, multiple function units, a matrix register file and a scalar register;

[0227] Among them, the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register;

[0228] Functional unit, used for performing functional processing according to instructions;

[0229] The matrix register stack is used to receive read and write requests and return data; rearrange the data and return it; perform read and write interactions with the functional units; and configure the configuration registers of the vector program control unit through the data in the matrix register stack.

[0230] Optionally, the vector program control unit is configured to, after fetching the instruction, determine whether to execute it, and transmit the instruction to the functional unit based on the determination result;

[0231] Vector program control unit, also used to control instruction jumps;

[0232] The vector program control unit has scalar computing capabilities.

[0233] Optionally, the functional unit includes: one or more vector operation units, one or more vector interleaving units, and one or more vector access units;

[0234] Any vector operation unit, used to perform vector operations according to instructions;

[0235] Any vector interleaving unit, used to perform data interleaving and logic processing according to instructions;

[0236] Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.

[0237] Optionally, any vector operation unit includes: a floating-point multiplication-addition operator unit, a floating-point multiplication-accumulation operator unit, a floating-point arithmetic operator unit, a tensor multiplication subunit and an intermediate result register;

[0238] The floating-point multiplication-addition operator unit, the floating-point multiplication-accumulation operator unit, the floating-point arithmetic operator unit, and the tensor multiplication subunit share the intermediate result register;

[0239] The floating-point multiplication-addition operator unit and the floating-point arithmetic operator unit share a single launch slot;

[0240] The floating-point multiply-accumulate operator and the tensor multiplication subunit share one emit slot.

[0241] Optionally, the vector processor further includes: a private vector register of the vector interleaving unit and a private vector register of the vector access unit;

[0242] Wherein, the private vector register of the vector interleaving unit corresponds to the vector interleaving unit one-to-one;

[0243] The private vector registers of the vector access unit are shared by multiple vector access units.

[0244] Optionally, the matrix register stack is also used to write data of other operation processors; and receive status information sent by other operation processors to query whether the data has been written.

[0245] Optionally, the vector program control unit is further configured to receive a start command sent by other operation processors to start the vector processor; and return an indication signal indicating whether the vector processor has ended to the other operation processors.

[0246] Optionally, a read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and other arithmetic processors;

[0247] The vector program control unit and other operation processors both perform read operations on the read FIFO unit and write operations on the write FIFO unit;

[0248] Other arithmetic processors perform read operations or write operations on the vector registers.

[0249] That is to say, in the high-performance processor provided by this embodiment,

[0250] The matrix register file is used to write data to the scalar processor and receive status information from the scalar processor to check whether the data has been written.

[0251] The vector program control unit is configured to receive a start command from the scalar processor, start the vector processor, and return an indication signal to the scalar processor indicating whether the vector processor has finished executing.

[0252] A read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and the scalar processor.

[0253] The vector program control unit and the scalar processor both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit.

[0254] The scalar processor reads or writes scalar registers.

[0255] This embodiment provides a high-performance processor. The vector processor in the high-performance processor includes: a vector program control unit, multiple functional units, a matrix register stack, and scalar registers. The vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar registers; the functional units are used to perform functional processing according to instructions; the matrix register stack receives read and write requests and returns data; rearranges the data before returning it; and interacts with the functional units for read and write operations. The configuration registers of the vector program control unit are configured using data in the matrix register stack. This vector arithmetic processor can efficiently process vector data.

[0256] Based on the same inventive concept of the vector processor, this embodiment provides an electronic device, which includes a high-performance processor, or the electronic device includes one or more processor clusters, wherein each processor cluster includes multiple high-performance processors.

[0257] Among them, high-performance processors can be Figure 4 As shown, the implementation details of the high performance processor can also be as shown in Figure 4 The embodiments shown in the drawings have been described and will not be described in detail here.

[0258] For example, the high-performance processor includes a scalar processor and a vector processor.

[0259] The vector processor can be Figure 1 or Figure 3 The vector processor shown.

[0260] Specifically, the vector processor includes: a vector program control unit, multiple function units, a matrix register file and a scalar register;

[0261] Among them, the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register;

[0262] Functional unit, used for performing functional processing according to instructions;

[0263] The matrix register stack is used to receive read and write requests and return data; rearrange the data and return it; perform read and write interactions with the functional units; and configure the configuration registers of the vector program control unit through the data in the matrix register stack.

[0264] Optionally, the vector program control unit is configured to, after fetching the instruction, determine whether to execute it, and transmit the instruction to the functional unit based on the determination result;

[0265] Vector program control unit, also used to control instruction jumps;

[0266] The vector program control unit has scalar computing capabilities.

[0267] Optionally, the functional unit includes: one or more vector operation units, one or more vector interleaving units, and one or more vector access units;

[0268] Any vector operation unit, used to perform vector operations according to instructions;

[0269] Any vector interleaving unit, used to perform data interleaving and logic processing according to instructions;

[0270] Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.

[0271] Optionally, any vector operation unit includes: a floating-point multiplication-addition operator unit, a floating-point multiplication-accumulation operator unit, a floating-point arithmetic operator unit, a tensor multiplication subunit and an intermediate result register;

[0272] The floating-point multiplication-addition operator unit, the floating-point multiplication-accumulation operator unit, the floating-point arithmetic operator unit, and the tensor multiplication subunit share the intermediate result register;

[0273] The floating-point multiplication-addition operator unit and the floating-point arithmetic operator unit share a single launch slot;

[0274] The floating-point multiply-accumulate operator and the tensor multiplication subunit share one emit slot.

[0275] Optionally, the vector processor further includes: a private vector register of the vector interleaving unit and a private vector register of the vector access unit;

[0276] Wherein, the private vector register of the vector interleaving unit corresponds to the vector interleaving unit one-to-one;

[0277] The private vector registers of the vector access unit are shared by multiple vector access units.

[0278] Optionally, the matrix register stack is also used to write data of other operation processors; and receive status information sent by other operation processors to query whether the data has been written.

[0279] Optionally, the vector program control unit is further configured to receive a start command sent by other operation processors to start the vector processor; and return an indication signal indicating whether the vector processor has ended to the other operation processors.

[0280] Optionally, a read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and other arithmetic processors;

[0281] The vector program control unit and other operation processors both perform read operations on the read FIFO unit and write operations on the write FIFO unit;

[0282] Other arithmetic processors perform read operations or write operations on the scalar registers.

[0283] This embodiment provides an electronic device, wherein a vector processor in the electronic device includes: a vector program control unit, multiple functional units, a matrix register stack, and scalar registers; wherein the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar registers; the functional units are used for performing functional processing according to instructions; the matrix register stack is used for receiving read and write requests and returning data; rearranging the data before returning it; performing read and write interactions with the functional units; and configuring the configuration registers of the vector program control unit using data in the matrix register stack. This vector arithmetic processor can efficiently process vector data.

[0284] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.

[0285] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0286] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0287] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0288] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0289] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A vector processor, characterized in that: The vector processor comprises: a vector program control unit, a plurality of function units, a matrix register file and a scalar register; Wherein, the vector program control unit is used for instruction fetching and instruction issuance; the vector program control unit interacts with the scalar register; The functional unit is used to perform functional processing according to the instruction; The matrix register stack is used to return data after receiving read and write requests; rearrange the data and return it; perform read and write interactions with the functional unit; and configure the configuration register of the vector program control unit through the data in the matrix register stack.

2. The vector processor according to claim 1, wherein: The vector program control unit is used to determine whether to execute the instruction after fetching it, and transmit the instruction to the functional unit based on the determination result; The vector program control unit is also used to control the jump of instructions; The vector program control unit has scalar computing capability.

3. The vector processor according to claim 1, wherein: The functional units include: one or more vector operation units, one or more vector interleaving units, and one or more vector access units; Any vector operation unit, used to perform vector operations according to instructions; Any vector interleaving unit, used to perform data interleaving and logic processing according to instructions; Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.

4. The vector processor according to claim 3, wherein: Any vector operation unit, including: floating-point multiply-add operator unit, floating-point multiply-accumulate operator unit, floating-point arithmetic operator unit, tensor multiplication subunit and intermediate result register; The floating-point multiplication-addition operator unit, the floating-point multiplication-accumulation operator unit, the floating-point arithmetic operator unit, and the tensor multiplication subunit share the intermediate result register; The floating-point multiplication-addition operator unit and the floating-point arithmetic operator unit share a launch slot; The floating-point multiplication-accumulation operator unit and the tensor multiplication subunit share one emission slot.

5. The vector processor according to claim 1, wherein: The vector processor further comprises: a private vector register of a vector interleaving unit and a private vector register of a vector access unit; Wherein, the private vector register of the vector interleaving unit corresponds to the vector interleaving unit one-to-one; The private vector registers of the vector access unit are shared by multiple vector access units.

6. The vector processor according to claim 1, wherein: The matrix register stack is also used to write data from other operation processors and receive status information from other operation processors to inquire whether data has been written.

7. The vector processor according to claim 1, wherein: The vector program control unit is further configured to receive a start command sent by other operation processors to start the vector processor; and return an indication signal indicating whether the vector processor has ended to the other operation processors.

8. The vector processor according to claim 1, wherein: A read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and other operation processors; The vector program control unit and other operation processors both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit; Other arithmetic processors perform read operations or write operations on the scalar registers.

9. A high-performance processor, characterized in that: include: The vector processor and scalar processor according to any one of claims 1 to 8; Matrix register file, used to write data to the scalar processor; Receive status information from the scalar processor to inquire whether the data has been written; A vector program control unit, configured to receive a start command sent by a scalar processor, start the vector processor, and return an indication signal to the scalar processor indicating whether the vector processor has ended; A read first-in first-out FIFO unit and a write FIFO unit are provided between the vector processor and the scalar processor; The vector program control unit and the scalar processor both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit; The scalar processor reads or writes scalar registers.

10. An electronic device, characterized in that: include: The high-performance processor according to claim 9; or, comprising one or more processor clusters, wherein each processor cluster comprises a plurality of high-performance processors according to claim 9.

Citation Information

Patent Citations

  • Parallel vector processing engine structure

    CN101833441A

  • Single instruction multiple data (SIMD) vector processor supporting fast Fourier transform (FFT) acceleration

    CN102495721A

  • Communication processor

    CN110096307A

  • Vector processor with vector first and multiple lane configuration

    CN113366462A

  • Single-instruction-multiple-data processing with combined scalar / vector operations

    CN1188275A

Cited By

  • Scalar processor, high-performance processor, and electronic device

    CN120540705A

  • Scalar processors, high-performance processors, and electronic devices

    CN120540705B