High-performance processor, processor cluster and electronic equipment
By introducing the design of sharing local memory with scalar processors and vector processors into the processor, the problem of how to coordinate the use of heterogeneous cores in heterogeneous multi-core architecture is solved, and more efficient computing task adaptability and performance is achieved.
Patent Information
- Application Number
- CN202510572336.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-30
AI Technical Summary
In a heterogeneous multi-core architecture, how to coordinately use multiple heterogeneous cores to meet different computing needs, especially when different types of cores in the processor have different architectures, clock frequency and function consumption, how to better adapt to different types of tasks.
It provides a high-performance processor, including a scalar processor and a vector processor, which share local memory. The vector processor only accesses local memory and is executed only by the scalar processor. The scalar processor establishes a connection with the global memory. The scalar processor obtains instructions and parameters from the global memory, and calls the vector processor to perform tasks when the execution conditions are met.
It realizes the efficient and coordinated use of multiple heterogeneous cores in a heterogeneous multi-core architecture to meet different computing needs and improves the adaptability and computing efficiency of the processor.
Smart Images

Figure CN120540708A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a high-performance processor, a processor cluster, and an electronic device. Background Art
[0002] In heterogeneous multi-core architectures, the processing cores can be different, with different architectures, clock frequencies, and power consumption. The goal is to make the processor better suited to different types of tasks by combining different types of cores.
[0003] In a heterogeneous multi-core architecture, it is very important to know how to process heterogeneous multi-core data and then coordinately use multiple heterogeneous cores to meet different computing needs. Summary of the Invention
[0004] In order to solve one of the above technical deficiencies, the present application provides a high-performance processor, a processor cluster and an electronic device.
[0005] In a first aspect of the present application, a high-performance processor is provided, the high-performance processor comprising: a scalar processor, a vector processor, and a local memory;
[0006] The scalar processor and the vector processor share local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor;
[0007] Establish connections between scalar processors and vector processors;
[0008] The scalar processor establishes a connection with the global memory;
[0009] The scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters.
[0010] Optionally, the scalar processor is an out-of-order multi-issue scalar processor;
[0011] The vector processor is a very long instruction word vector processor.
[0012] Optionally, the high-performance processor further includes:
[0013] Instruction first-in-first-out FIFO memory;
[0014] Among them, the depth of the instruction FIFO memory is 32 bits.
[0015] Optionally, the high-performance processor further comprises: an instruction FIFO memory storage and a data FIFO memory storage;
[0016] The depth of the instruction FIFO memory and the data FIFO memory are both 32 bits.
[0017] Optionally, the execution condition is that the vector processor is not performing task processing.
[0018] Optionally, the execution condition is that the instruction FIFO memory is not full;
[0019] A scalar processor for storing the storage addresses of instructions and parameters into an instruction FIFO memory;
[0020] The vector processor is used to read instructions and data from the instruction FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0021] Optionally, the execution condition is that both the instruction FIFO memory and the data FIFO memory are not full;
[0022] A scalar processor, configured to store the memory address of the instruction into the instruction FIFO memory and store the memory address of the parameter into the data FIFO memory;
[0023] The vector processor is configured to read instructions from the instruction FIFO memory and data from the data FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0024] Optionally, the task includes an identifier of the thread to which the task belongs.
[0025] A second aspect of the present application provides a processor cluster comprising: a plurality of high-performance processors as described in the first aspect.
[0026] In a third aspect of the present application, an electronic device is provided, comprising: the high-performance processor described in the first aspect, or the high-performance processor described in the second aspect.
[0027] The present application provides a high-performance processor, a processor cluster, and an electronic device. The high-performance processor includes: a scalar processor, a vector processor, and a local memory; the scalar processor and the vector processor share the local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; a connection is established between the scalar processor and the vector processor; the scalar processor is connected to the global memory; the scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute a task based on the parameters. The high-performance processor provided by the present application, after the scalar processor obtains instructions and parameters from the global memory, calls the vector processor to execute the task based on the parameters when it determines that the execution conditions are met, thereby coordinately using multiple heterogeneous cores to meet different computing needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0029] Figure 1 A schematic diagram of the structure of a high-performance processor provided in an embodiment of the present application;
[0030] Figure 2 A schematic diagram of the structure of another high-performance processor provided in an embodiment of the present application;
[0031] Figure 3 A schematic diagram of a high-performance processor according to an embodiment of the present application implementing an execution condition in which a vector processor does not process a task;
[0032] Figure 4 A schematic diagram of a high-performance processor according to an embodiment of the present application implementing an execution condition in which an instruction FIFO memory is not full;
[0033] Figure 5 A schematic diagram of the structure of a scalar processor provided in an embodiment of the present application;
[0034] Figure 6 A schematic structural diagram of a synchronization unit of a scalar processor provided in an embodiment of the present application;
[0035] Figure 7 A schematic diagram of the architecture of a vector processor provided in an embodiment of the present application;
[0036] Figure 8 A schematic diagram of the structure of a vector operation unit provided in an embodiment of the present application;
[0037] Figure 9 A schematic diagram of the architecture of another vector processor provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the technical solutions and advantages of the embodiments of the present application more clearly understood, the exemplary embodiments of the present application are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, and are not an exhaustive list of all the embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless they conflict.
[0039] In the process of implementing this application, the inventors discovered that in a heterogeneous multi-core architecture, the processing cores can be different, and they may have different architectures, clock frequencies, and power consumption. The goal of this approach is to combine different types of cores together to make the processor better adapted to different types of tasks. In a heterogeneous multi-core architecture, it is very important to know how to perform heterogeneous multi-core processing and then coordinately use multiple heterogeneous cores to meet different computing needs.
[0040] In response to the above problems, an embodiment of the present application provides a high-performance processor, a processor cluster, and an electronic device. The high-performance processor includes: a scalar processor, a vector processor, and a local memory; the scalar processor and the vector processor share the local memory; and the vector processor only accesses the local memory and is only called and executed by the scalar processor; a connection is established between the scalar processor and the vector processor; the scalar processor establishes a connection with the global memory; the scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters. In the high-performance processor provided in the present application, after the scalar processor obtains instructions and parameters from the global memory, when it determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters, thereby coordinately using multiple heterogeneous cores to meet different computing needs.
[0041] This embodiment provides a high-performance processor, which is a heterogeneous asynchronous processor composed of a scalar processor and a vector processor.
[0042] See also Figure 1 , the high-performance processor includes: a scalar processor, a vector processor, and a local memory.
[0043] The scalar processor and the vector processor share local memory. The vector processor can only access local memory and can only be called and executed by the scalar processor.
[0044] A connection is established between the scalar processor and the vector processor. For example, the scalar processor and the vector processor are connected via a dedicated instruction channel.
[0045] like Figure 2 As shown, the high-performance processor may also include two registers, one register corresponding to the scalar processor and the other register corresponding to the vector processor. The vector processor can read and write its corresponding register, and the scalar processor can read and write its corresponding register as well as the vector processor's corresponding register.
[0046] The scalar processor can read and write the registers of the vector processor.
[0047] Scalar processors establish connections to global memory.
[0048] (1) Scalar processor
[0049] The scalar processor is an out-of-order multi-issue scalar processor.
[0050] The scalar processor primarily handles data movement, task management, synchronization, and switching, and is responsible for invoking the vector processor to perform accelerated operations. For example, it retrieves instructions and parameters from global memory. Once the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters.
[0051] In specific implementation, the scalar processor can be Figure 5 As shown, the scalar processor includes: an instruction fetch unit, a register renaming unit, an operation reservation stack unit, a storage reservation stack unit, a scalar operation unit, a memory access unit, a program control unit, a synchronization unit, a pipeline control unit, a register file unit, and a special vector register file unit.
[0052] In addition, the scalar processor may also include one or more other units, such as one or more other functional modules, one or more instruction caches, one or more data storages, one or more special vector registers, one or more status flag registers, etc.
[0053] 1. Instruction fetch unit
[0054] The instruction fetch unit is used to fetch and dispatch instructions.
[0055] Specifically, the instruction fetch unit generates an instruction fetch request address, outputs the fetch request address to the instruction cache for instruction fetching, receives instructions from the instruction cache, and stores them in the data store. Each cycle, it sequentially reads qualified instructions from the data store, decodes and performs relevant checks on the read instructions, and then dispatches the checked instructions sequentially.
[0056] For example, the instruction fetch unit generates an instruction fetch request address and outputs it to the instruction cache for instruction fetch, and receives instructions from the instruction cache and stores them in the data storage. In each cycle, it sequentially selects one or more instructions from the qualified instructions for decoding and related checks, and dispatches the qualified instructions sequentially, dispatching at most one program control unit instruction and one synchronization unit instruction at a time. In addition, one or more scalar operation unit instructions and one or more memory access unit instructions can also be dispatched each time.
[0057] 2. Register renaming unit
[0058] The register renaming unit is used to receive instructions dispatched by the instruction fetch unit and rename registers.
[0059] Specifically, the register renaming unit receives and stores instructions dispatched by the instruction fetch unit, renames special vector registers, conditionally decodes instructions, and generates pipeline stall signals. It also receives and writes data from one or more of the scalar arithmetic unit, memory access unit, program control unit, synchronization unit, special vector registers, condition registers, and flag registers. It also sends instructions to one or more of the arithmetic hold stack unit, memory hold stack unit, program control unit, and synchronization unit.
[0060] For example, the register renaming unit is used in a scalar processor to receive instructions dispatched by the instruction fetch unit and perform register and special vector register renaming, instruction conditional decoding, and generate pipeline congestion signals. At the same time, it receives data from the execution unit (such as the scalar operation unit, memory access unit, program control unit, synchronization unit) to write back registers, special vector registers, condition registers, and status flag registers and writes them back to the corresponding registers.
[0061] The scalar processor write-back supports out-of-order write-back, has high execution efficiency, and distributes instructions to the operation retention stack unit, the storage retention stack unit, the program control unit, or the synchronization unit.
[0062] The register renaming unit bandwidth may be 6 bits, wherein multiple (eg, 4) input instructions may be valid at the same time.
[0063] There can be multiple condition registers, which are located in the register renaming unit.
[0064] The instructions of the scalar operation unit and memory access unit support the operations of reading and writing condition registers.
[0065] The synchronous unit's instructions support the operation of reading the condition register.
[0066] The jump and function call instructions of the program control unit support the operation of reading the condition register.
[0067] When an instruction enters the condition register, if there is an unexecuted instruction in the condition register, the pipeline is blocked.
[0068] That is, the condition register is not renamed, and when a read-write dependency occurs, a dispatch block is triggered to wait. The read and write rules for the condition register are as follows:
[0069] ●Read the rules:
[0070] (1) All instructions of the scalar arithmetic unit, memory access unit, and synchronization unit support conditional execution and require reading the value of the condition register.
[0071] (2) The scalar arithmetic unit also supports read conditional register instruction operations.
[0072] (3) The jump and function call instructions of the program control unit support read condition register operations.
[0073] ●Write rules:
[0074] (1) The scalar arithmetic unit supports write conditional register instructions.
[0075] (2) Scalar operation unit logic instructions and comparison instructions support the option of writing condition registers.
[0076] When the previously issued instruction to write the condition register has not yet been completed, and an instruction to read or write the same condition register enters, the pipeline is blocked and a conditional execution blocking signal is generated, waiting for the previous condition register to be written.
[0077] In addition, the register renaming unit includes: one or more physical registers and one or more logical registers.
[0078] Wherein, any physical register is one of the following: a scalar physical register, a vector physical register, a condition register, and a flag register.
[0079] Any logical register is one of the following: a scalar logical register, a vector logical register.
[0080] For example, the register renaming unit includes one or more physical registers, such as a plurality of 512-bit wide special vector registers, a plurality of condition registers, and a status flag register.
[0081] Among them, special vector registers are renamed, while condition registers and status flag registers are not renamed.
[0082] There are multiple logical registers, such as multiple logical registers including scalar logical registers and multiple vector logical registers.
[0083] In addition, the mapping relationship between logical registers and physical registers is maintained by the register mapping table. The mapping relationship between vector logical registers and vector physical registers is maintained by the special vector register mapping table.
[0084] 1) Register Mapping Table
[0085] Initially, the physical registers mapped to the entries corresponding to all logical register indices in the register mapping table are all 0. When an instruction is executed or an interrupt occurs, the logical registers allocated to the relevant physical registers are determined, and the mappings of the entries corresponding to the allocated logical register indices in the register mapping table are updated to the identifiers of the relevant physical registers.
[0086] For example, a register map table with a depth of 32 bits and a width of 6 bits stores the mapping between all logical registers and all physical registers. Initially, the mapping in the register map table is invalid, and all entries for the mapped physical registers are set to zero. When a physical register is assigned to a logical register, the entry corresponding to the logical register index in the register map table is changed to the ID of the physical register.
[0087] It should be noted that the register mapping table is updated only when the instruction is actually executed. If the conditional execution instruction is not executed, the register mapping table will not be updated. In addition, the register mapping table will not be updated when a jump occurs. However, when an interrupt occurs, the interrupt return address must update the register mapping table to ensure that the interrupt can return normally.
[0088] 2) Special vector register mapping table
[0089] Initially, the mapping vector physical registers of all entries corresponding to the vector logical register indexes in the special vector register mapping table are all 0. When an instruction is executed, the vector logical register allocated to the relevant vector physical register is determined, and the mapping of the entry corresponding to the allocated vector logical register index in the special vector register mapping table is updated to the identifier of the relevant vector physical register.
[0090] For example, the special vector register mapping table has a depth of 4 and a width of 3 bits, storing the mapping relationship between all vector logical registers and all vector physical registers. Initially, the special vector register mapping table is invalid, and all entries for the mapped vector physical registers are all zeros. When a vector physical register is assigned to a vector logical register, the entry corresponding to the vector logical register index in the special vector register mapping table is changed to the ID of the vector physical register.
[0091] It should be noted that the special vector register mapping table is updated only when the instruction is actually executed. If the conditional execution instruction is not executed, the special vector register mapping table will not be updated. In addition, the special vector register mapping table will not be updated when a jump occurs.
[0092] 3. Operation reservation stack unit
[0093] The operation reserve stack unit is the emission queue of the scalar operation unit.
[0094] The operation reserve stack unit receives instructions, dispatch and renaming information from the register renaming unit and pushes them into the queue. It then pops ready instructions to the scalar operation unit for execution.
[0095] The operation reserve stack unit is also used to decode input instructions and store instruction type information.
[0096] In other words, the Arithmetic Hold Stack unit acts as the issue queue for the scalar arithmetic unit. It receives instructions and associated dispatch and renaming information from the register renaming unit, pushes them into the queue, and then pops ready instructions onto the scalar arithmetic unit for execution. The Arithmetic Hold Stack unit decodes the incoming instructions and stores the instruction type information.
[0097] In specific implementation, the depth of the operation reserve stack unit can be flexibly adjusted, for example, the depth of the operation reserve stack unit is 8. Multiple scalar operation units share one operation reserve stack unit.
[0098] The rules for issuing and receiving instructions for the operation reserve stack unit are as follows:
[0099] (1) The output of the register renaming unit enters the operation preservation stack unit.
[0100] (2) When there is any idle scalar operation unit, it will fetch instructions and operands from the operation reservation stack unit for execution.
[0101] (3) The principle of executing instructions from the operation reservation stack unit is to execute the executable instructions that can be sent from the operation reservation stack unit in a forward-to-back order.
[0102] (4) Whether the transmission can be made is determined by whether the values of all source registers or special vector registers or condition registers and status flag registers are ready.
[0103] (5) If there are multiple instructions that can be sent, the oldest instruction will be sent first according to the instruction order.
[0104] (6) If any scalar operation unit is blocked, it can no longer receive new instructions.
[0105] (7) If the instruction previously sent to any scalar operation unit is a division instruction, a new division instruction can only be sent to it after the division result calculation is completed and the calculation completion En signal is returned.
[0106] 4. Storage Retention Stack Unit
[0107] The storage reservation stack unit is the transmit queue of the memory access unit.
[0108] The storage reservation stack unit is used to receive instructions and register renaming information from the register renaming unit and push them into the queue.
[0109] The storage reservation stack unit is also used to send a read request to the register renaming unit when the instruction address register is ready, and save the read address operand.
[0110] The register renaming unit is also used to calculate the address, decode the address, and save the decoding information after the instruction obtains the address.
[0111] The register renaming unit is also used to detect when the source register of any instruction is ready and the address decoding is completed, and then send it to the memory access unit for execution.
[0112] In specific implementation, the depth of the storage reservation stack unit can be flexibly adjusted, such as the depth of the storage reservation stack unit is 16. Multiple memory access units share one storage reservation stack unit. The storage reservation stack unit is the transmission queue of the memory access unit. The storage reservation stack unit receives instructions and register renaming information from the register renaming unit and pushes them into the queue. When the instruction address register in the storage reservation stack unit is ready, a read request is sent to the register renaming unit and the read address operand is saved in the queue. After the instruction in the storage reservation stack unit obtains the address, the address can be calculated and decoded, and the generated decoding information is saved in the queue. When the source register of an instruction (such as a write instruction) in the storage reservation stack unit is ready and the address decoding is completed, it can be transmitted to the memory access unit for execution. Before transmission, it must undergo a series of checks, such as address type check, address comparison check, and address forward check.
[0113] The rules for sending and receiving instructions to the storage reserve stack unit are as follows:
[0114] (1) The output of the register renaming unit enters the storage reservation stack unit.
[0115] (2) When the source operand for calculating the address is ready, the memory access address is calculated and saved in the storage reservation stack unit.
[0116] (3) Instructions with unrelated addresses: They can be sent out of order. The out-of-order rules are: read instructions after read instructions, write instructions after read instructions, and read instructions after write instructions. They can all be sent out of order. Write instructions after write instructions need to maintain order (they cannot be sent to different memory access units at the same time). Even if the addresses are unrelated, write instructions after write instructions still need to maintain order.
[0117] (4) Address-related instructions: read instructions followed by write instructions, write instructions followed by read instructions, write instructions followed by write instructions, and read instructions followed by read instructions all need to be performed in order.
[0118] (5) When the addresses are unrelated but all instructions that have not been sent successfully (i.e., instructions on the way that have not been sent to the destination, including those at the memory access unit level and the memory access unit output level) are located in the same storage space, they can be sent out of order to the same memory access unit, but they cannot be sent to two or more memory access units.
[0119] (6) Only one memory access instruction located in the same storage space but with unrelated addresses can be sent at the same time, and two or more memory access units cannot be sent at the same time.
[0120] (7) Address correlation judgment principle: Whether the addresses are related is irrelevant if they are located in different storage spaces. If they are located in the same storage space, whether the addresses are related is determined based on the data granularity.
[0121] 5. Scalar arithmetic unit
[0122] In a specific implementation, there may be one or more scalar operation units.
[0123] For example, the scalar processor includes two scalar arithmetic units, namely scalar arithmetic unit 0 and scalar arithmetic unit 1.
[0124] The scalar operation unit is used to receive instructions and data sent by the operation reservation stack unit, perform operations on the data based on the instructions, and write the operation results back to the register renaming unit.
[0125] The scalar arithmetic unit is the computing unit of the scalar processor, which can perform various types of fixed-point and floating-point operations, such as addition, subtraction, multiplication, division, logical operations, comparison operations, shifts, etc. It receives instructions and data sent by the operation reserve stack unit, performs operations, and writes the results back to the register file unit of the register renaming unit or the special vector register file unit.
[0126] Several instruction examples are provided below as examples. In specific implementation, they are not limited to the following instructions, nor are they limited to including all instructions.
[0127] Instructions with execution level one include: fixed-point addition and subtraction, logical instructions, shift instructions, fixed-point and floating-point comparison instructions, read and write Flag instructions, fixed-point maximum and minimum instructions, ABS instructions, bit reversal instructions, selection instructions, special vector register distribution instructions, read special vector register instructions, Byte reversal instructions, Merge instructions, immediate value assignment instructions, FirstOne instructions, CRC instructions, floating-point classification instructions, floating-point partial domain extraction, and Rounding instructions.
[0128] Instructions with execution level three include: fixed-point multiplication instructions, fixed-floating-point conversion instructions, bit filtering instructions, Count instructions, and floating-point addition and subtraction instructions.
[0129] Instructions that support Bypass include: selection instructions, fixed-point addition and subtraction instructions, shift instructions, immediate value assignment instructions, ABS instructions, logical instructions, comparison instructions, and maximum and minimum instructions.
[0130] The execution cycle of the division instruction is uncertain and is related to the data of the divisor and the dividend. When the instruction is executed, a DivEn instruction will be generated to indicate that the instruction is executed and the result is output to the register stack. No new division instructions can be input during the execution of the division instruction, but other scalar calculation unit instructions can be input. After the division is executed, the output result is reused with the output port of the first-stage pipeline. When the output port of the first-stage pipeline is not used by other scalar calculation unit instructions, the division outputs its result and outputs the DivEn identifier at the same time. The DivEn identifier is output to the operation retention stack unit, indicating that the Div instruction can continue to be output to the current scalar calculation unit.
[0131] 6. Memory access unit
[0132] In a specific implementation, there may be one or more memory access units.
[0133] For example, a scalar processor includes two memory access units, namely memory access unit 0 and memory access unit 1.
[0134] The memory access unit is used to receive instructions, data and register information sent by the storage reservation stack unit, and read and write the data based on the instructions and register information.
[0135] The memory access unit is a functional module that executes memory access-related instructions for scalar processors. The memory access unit receives instructions and data, as well as register-related information, from the storage retention stack unit. It executes instructions accordingly and interacts with other units to read and write data. Read instructions and atomic write instructions require writing data back to the register renaming unit. These instructions include register-level read and write instructions, including 8-bit, 16-bit, 32-bit, 64-bit, or other bit granularities, as well as vector read and write instructions, including 128-bit, 256-bit, 512-bit, or other bit granularities. Different instructions have different processing cycles.
[0136] In addition, the memory access unit is responsible for providing the relevant instruction quantity information required by FENCE. The memory access unit interacts with the storage reservation stack unit to complete the data storage configuration.
[0137] 7. Program control unit
[0138] In a specific implementation, there is only one program control unit.
[0139] The program control unit is configured to receive instructions and data from the register renaming unit, process the data based on the instructions, and output a processing result.
[0140] The program control unit (PCU) executes instructions related to the scalar processor's program execution sequence. The PCU receives instructions and data from the register renaming unit (RRU), processes the data accordingly, and outputs the results to other modules in the scalar processor. Different instructions are processed in different time periods.
[0141] The program control unit is responsible for controlling the direction of program execution (such as stop, interrupt, jump, function call), involving the execution of related instructions and the reading and writing control of configuration information; the program control unit is responsible for the configuration and prefetch operations of the instruction cache, as well as FENCE operations; the program control unit is responsible for the reading, writing and control of the counter, as well as the reading and writing of some other control information.
[0142] 8. Synchronization unit
[0143] In a specific implementation, there is only one synchronization unit.
[0144] Synchronization unit, used to synchronize the scalar processor and the vector processor.
[0145] like Figure 6 As shown, a communication connection is established between the synchronization unit and the pipeline control unit, the register renaming unit, the program control unit, and the vector processor.
[0146] The instructions of the synchronization unit come from the register renaming unit, and the reading and writing of the data of the synchronization unit interact with the register renaming unit.
[0147] The synchronization unit is used to receive the pause signal sent by the pipeline control unit and send the execution level pause signal generated when communicating with the vector processor to the pipeline control unit so as to generate the execution pause signal of the scalar processor.
[0148] The synchronization unit is used to generate instructions and transmit them to the program control unit.
[0149] That is to say, the synchronization unit is a unit that synchronizes the scalar processor and the vector processor. It receives instructions and data sent by the register renaming unit, reads data from the vector processor and writes it back to the register stack, reads data from the register stack unit or the special vector register stack unit and sends it to the functional module of the vector processor. It is responsible for the startup and status query of the vector processor, such as querying the reading and writing of the read and write FIFO (FirstInput FirstOutput) in the vector program control unit of the vector processor, the configuration of the register file stack, the reading or writing of the scalar register, the register file stack status query, the reading FIFO depth, the reading of the startup vector processor instruction counter, etc., and providing the program control unit with the synchronization unit instruction information.
[0150] The synchronization unit interacts with the pipeline control unit, register renaming unit, and program control unit within the scalar processor, as well as with the external vector processor, scalar processor, and vector processor transfer queue module. Synchronization unit instructions originate from the register renaming unit, and data reading and writing must interact with the register renaming unit. The synchronization unit receives a stall signal from the pipeline control unit and, when communicating with the vector processor, generates its own execute-level stall signal, which is sent to the pipeline control unit to generate the ExeStall signal for the entire scalar processor. The synchronization unit generates the instruction to be executed in the next cycle and transmits it to the program control unit for use by the program control unit's counter instruction. Interactions with the vector processor include, but are not limited to, configuring the register file with special vector registers or registers, reading and writing scalar registers, and querying the write status of the register file. Interactions with the scalar processor and vector processor transfer queue module include, but are not limited to, starting the vector processor, querying vector processor status, reading and writing data in the vector processor's instruction fetch unit FIFO, reading the FIFO depth, and reading the start vector processor instruction counter.
[0151] Therefore, in a specific implementation, the synchronization unit may have the following functions (it should be noted that the following functions are only examples, and other functions may be provided. This embodiment and subsequent embodiments do not limit the specific functions of the synchronization unit):
[0152] The start vector processor function is used to start the vector processor, including immediate start and register start, such as the pipeline waits until the start is successful, or writes the result of the start success or failure back to the destination register.
[0153] Query the vector processor execution status function, support option B.
[0154] Read and write FIFO function, the FIFO is located in the instruction fetch unit of the vector processor, such as the FIFO bit width 32 bits, the read and write FIFO such as the read and write FIFO waits until success, or the read and write FIFO success or failure result is written back to the register.
[0155] Write register file stack functions, including special vector register writes or register writes.
[0156] Read and write scalar register functions, including immediate index or register index read and write.
[0157] Query the register file stack write back status function, such as waiting until all writes to the register file stack are completed, or returning the result of whether the write to the register file stack is completed to the register.
[0158] When the related operations are not completed, the synchronization unit will generate its own blocking signal, blocking and waiting, and the signal will be sent to the pipeline control unit to generate a pipeline blocking signal.
[0159] A FIFO (such as a 32-bit deep FIFO) can also be added between the scalar processor and the vector processor to store the request to start the vector processor, and move the read and write FIFO previously located in the vector processor to the scalar processor and vector processor transmission queue module. The scalar processor and vector processor transmission queue module unit implements the startup of the vector processor, queries the execution status of the vector processor, reads and writes FIFO functions, reads FIFO depth functions, and reads the startup vector processor instruction counter function. The condition for the successful startup of the vector processor is that the startup vector processor FIFO is not full, and the query of the vector processor execution status is passed. The condition for the vector processor status to be stopped is that the vector processor execution is completed and the startup vector processor FIFO is empty.
[0160] 9. Assembly line control unit
[0161] The pipeline control unit is used to generate a pipeline pause signal and / or generate a start and stop signal for the scalar processor.
[0162] The pipeline control unit is the pipeline control unit of the scalar processor, which is connected to each unit inside the scalar processor and is responsible for generating pipeline blocking signals, such as blocking in normal working mode and blocking in debug mode.
[0163] The pipeline control unit also communicates with the communication and synchronization unit to generate signals for starting and stopping the scalar processor.
[0164] Furthermore, in practical applications, scalar processors can also perform conditional execution decoding. For example, when performing conditional execution decoding, a scalar processor determines the execution condition of the instruction's preset bits. If the condition is met, a valid instruction is output; otherwise, a null instruction is output. A null instruction refers to an empty or invalid instruction.
[0165] If there are read or write operations on the condition register, the pipeline is blocked and the read operation can only be performed after the write operation is completed. There is no bypass when reading or writing the condition register.
[0166] Taking two condition registers, namely condition register 0 and condition register 1, and the preset bits being [29:28] as an example, when the scalar processor performs conditional execution decoding, the scalar processor judges the execution condition of the input instruction based on the [29:28] bits of the instruction set encoding. If the condition is met, a valid instruction is output, otherwise an empty instruction is output.
[0167] Among them, [29:28] bits are 00, indicating that condition register 0 is 1 and execution is executed, [29:28] bits are 01, indicating that condition register 1 is 1 and execution is executed, [29:28] bits are 10, indicating that condition register 0 is executed, and [29:28] bits are 11, indicating unconditional execution. If the condition is not met, the instruction is invalid and a null instruction is output.
[0168] If the condition register is read or written, the pipeline is blocked and the read operation is performed after the condition register is written. There is no bypass when reading or writing the condition register.
[0169] (2) Vector Processor
[0170] The vector processor is a Very Long Instruction Word (VLIW) vector processor.
[0171] The vector processor can be unbalanced clustered and can execute more than 20 instructions concurrently, such as multiple vector / matrix operation acceleration instructions, loop acceleration methods, etc.
[0172] In specific implementation, the vector processor can be Figure 7 As shown, the vector processor includes: a vector program control unit, multiple function units, a register file and a scalar register.
[0173] In addition, the vector processor also includes: a private vector register of a vector interleaving unit and a private vector register of a vector access unit.
[0174] 1. Vector program control unit
[0175] Vector program control unit, used for instruction fetching and instruction issuance.
[0176] That is, the vector program control unit is used to fetch instructions, determine whether to execute them, and send the instructions to the functional units based on the determination result.
[0177] The vector program control unit is also used to control instruction jumps.
[0178] The vector program control unit has scalar computing capabilities.
[0179] The vector program control unit interacts with the scalar registers.
[0180] In specific implementation, the vector program control unit is an instruction fetch and instruction issuance unit. It takes instructions from the cache according to the PC value, and after determining whether to execute them, it issues the instructions to each functional unit according to the wait value (configured by the wait instruction). At the same time, it controls the jump of instructions and has some scalar computing capabilities.
[0181] In addition, the vector program control unit is further configured to receive a start command from other operation processors to start the vector processor and return an indication signal to the other operation processors indicating whether the vector processor has finished operation.
[0182] Taking the scalar processor as an example, the vector program control unit receives the start command issued by the synchronization unit of the scalar processor, starts the vector processor execution, and also returns an indication signal indicating whether the synchronization unit vector processor execution is completed.
[0183] 2. Functional Unit
[0184] Functional unit, used to perform functional processing according to instructions.
[0185] For example, the functional unit receives an instruction from the vector program control unit, processes data accordingly according to the instruction, and outputs the processing result according to the address specified in the instruction.
[0186] The functional units include: one or more vector operation units, one or more vector interleaving units, and one or more vector access units.
[0187] 1) Vector operation unit
[0188] Any vector arithmetic unit, used to perform vector operations according to instructions.
[0189] like Figure 8 As shown, any vector operation unit includes: a floating-point multiplication-addition operator unit, a floating-point multiplication-accumulation operator unit, a floating-point arithmetic operator unit, a tensor multiplication subunit and an intermediate result register.
[0190] The floating-point multiplication and addition unit and the floating-point arithmetic unit share a single issue slot, so a maximum of eight vector unit instructions can be issued per cycle.
[0191] The floating-point multiply-accumulate operator and the tensor multiplication subunit share one emit slot.
[0192] The floating-point multiply-add operator unit is a functional unit that executes floating-point multiply-add operator unit related instructions. For example, floating-point multiply-add operator unit related instructions include integer and floating-point vector multiplication and accumulation, multiplication, addition, tensor calculation, etc.
[0193] 1 vector operation unit has independent intermediate result registers.
[0194] One floating-point multiplication-addition operator unit, one floating-point multiplication-accumulation operator unit, one tensor multiplication subunit, and one floating-point arithmetic operator unit share intermediate result registers.
[0195] (1) Floating-point multiplication-addition operator unit and floating-point multiplication-accumulation operator unit, which can perform integer and floating-point vector multiplication, multiplication-accumulation and other operations. Supported types include but are not limited to int32, fp32, and fp64.
[0196] (2) The floating-point arithmetic subunit can perform integer and floating-point vector arithmetic operations, such as comparison, addition, subtraction, bitwise operations, etc. Supported types include but are not limited to int8, uint8, int16, uint16, int32, uint32, bool, fp16, bf16, fp32, tf32, and fp64.
[0197] (3) The tensor multiplication subunit can perform tensor multiplication, multiply-accumulate and other operations. Supported types include but are not limited to int8, bf16, fp16, and tf32.
[0198] 2) Vector interleaving unit
[0199] Any vector interleaving unit is used to perform data interleaving and logic processing according to instructions.
[0200] The vector interleaver unit is the control and data processing unit within the vector processor. It is responsible for interleaving data and supports logical and some fixed-point and floating-point calculations. It also supports a wide range of customized instructions, including table lookup, lateral calculations, sparse matrix calculations, precision conversion, and FIFO (First Input First Output). It also executes instructions such as data broadcasting, extraction, and internal interleaving.
[0201] Each vector interleaving unit has a set of private vector registers, so the private vector registers of the vector interleaving units correspond one to one with the vector interleaving units.
[0202] 3) Vector access unit
[0203] Any vector access unit, used to perform multi-mode memory access, address calculation and scalar calculation according to the instruction.
[0204] The vector access unit is a memory access unit within a vector processor, primarily responsible for reading / writing instructions and various scalar calculations.
[0205] The read instruction / write instruction supports multiple memory access modes, such as row mode, column mode / discrete mode / extended mode / accumulated mode.
[0206] It also supports multiple parameter configurations, with a maximum read / write instruction data width of 1024 bits. It can also perform address calculation, load / store and other instructions.
[0207] All vector access units share a set of private vector registers, so the private vector registers of a vector access unit are shared by multiple vector access units.
[0208] 3. Register file stack
[0209] The register file receives read and write requests and returns data. It rearranges the data before returning it. It interacts with the functional units for read and write operations. The vector program control unit's configuration registers are configured using data in the register file.
[0210] The register file is a general-purpose vector register stack and is the main storage unit within the vector processor. It is responsible for receiving read and write requests and returning data. In some functions, it can rearrange the data before returning it to the request module.
[0211] The register file stack performs read and write interactions with the functional units in the vector processor (such as the floating-point multiplication-addition operator unit, the floating-point arithmetic operator unit, the floating-point multiplication-accumulation operator unit, and the tensor multiplication subunit), and supports the use of data in the register file stack to configure the instruction fetch unit configuration register.
[0212] The register file stack is also used to write data to other processors and receive status information from other processors to check whether the data has been written.
[0213] Taking a scalar processor as an example, the synchronization unit of the scalar processor can write data into the register file stack, and the register file stack can also receive status information from the synchronization unit of the scalar processor to inquire whether the data has been written.
[0214] The depth of the register file is configurable.
[0215] Figure 9 A schematic diagram of a vector processor is shown, in which the functional units include four vector operation units, four vector interleaving units, and four vector access units.
[0216] The vector processor provided in this embodiment supports a VLIW (Very Long Instruction Word) instruction set. Each VLIW may be composed of one or more instructions, and each instruction corresponds to a functional unit.
[0217] In addition, a read FIFO unit and a write FIFO unit are provided between the vector processor and other operation processors.
[0218] The vector program control unit and other operation processors both perform a read operation on the read FIFO unit and a write operation on the write FIFO unit.
[0219] Other arithmetic processors perform read operations or write operations on the vector registers.
[0220] Taking a scalar processor as an example, a read FIFO and a write FIFO unit for transmitting data are provided between the scalar processor and the vector processor. The scalar processor and the vector program control unit can perform read operations or write operations on the read and write FIFO.
[0221] At the same time, the synchronization unit of the scalar processor can read or write the scalar registers of the vector processor.
[0222] In addition, the high-performance processor may further include an instruction FIFO (First Input First Output) memory. Alternatively, the high-performance processor may further include an instruction FIFO memory and a data FIFO memory.
[0223] Among them, the depth of the instruction FIFO memory is 32 bits, and the depth of the data FIFO memory is 32 bits.
[0224] In a specific implementation, there may be various execution conditions, for example: the execution condition is that the vector processor is not processing a task. Alternatively, the execution condition is that the instruction FIFO memory is not full (for example, the depth of the instruction FIFO memory is 32 bits, in which case the instruction FIFO memory constitutes an asynchronous queue). Alternatively, the execution condition is that both the instruction FIFO memory and the data FIFO memory are not full (for example, the depth of the instruction FIFO memory and the data FIFO memory is 32 bits, in which case the instruction FIFO memory constitutes an asynchronous queue for instructions, and the data FIFO memory constitutes an asynchronous queue for parameters).
[0225] Depending on the execution conditions, the specific implementation schemes of high-performance processors are also different, which are explained below:
[0226] ●The execution condition is that the vector processor is not processing the task
[0227] Take two tasks (task0 and task1) to be executed as an example, see Figure 3 ( Figure 3 The white boxes in the middle are executed by the scalar processor, and the gray boxes are executed by the vector processor. The processing process of the high-performance processor for this execution condition is as follows:
[0228] The scalar processor gets task0 and its parameters from global memory.
[0229] The scalar processor determines whether the vector processor is currently processing a task. If the vector processor is not currently processing a task, the scalar processor calls the vector processor to execute task0 based on the parameters of task0.
[0230] The vector processor will then execute task0 based on the parameters of task0. At the same time, the scalar processor can again obtain task1 and the parameters of task1 from the global memory to prepare for the execution of task1.
[0231] The scalar processor synchronizes with the vector processor to confirm whether the vector processor is processing task0. If the vector processor has not completed task0 and is currently processing a task, the scalar processor waits until the vector processor completes processing task0.
[0232] The scalar processor again determines that the vector processor is not processing the task, and then the scalar processor calls the vector processor to execute task1 based on the parameters of task1.
[0233] In this heterogeneous multi-core processing process, the scalar processor needs to be synchronized with the vector processor. The scalar processor can call the vector processor to process the next task only after the vector processor has completed processing the previous task.
[0234] ●The execution condition is that the instruction FIFO memory is not full
[0235] Under these execution conditions, the scalar processor in a high-performance processor calls the vector processor to execute a task based on parameters. The scalar processor stores the storage address of the instruction and parameters in the instruction FIFO memory. The vector processor reads the instruction and parameters from the instruction FIFO memory when not processing a task and executes the instruction based on the parameters.
[0236] Take two tasks (task0 and task1) to be executed as an example, see Figure 4 ( Figure 4 The white boxes in the middle are executed by the scalar processor, and the gray boxes are executed by the vector processor). The implementation process of this execution condition in the high-performance processor is based on the asynchronous queue instruction FIFO memory.
[0237] For example, the scalar processor obtains task0 and its parameters from the global memory.
[0238] The scalar processor determines whether the instruction FIFO is full. If not, it packages task0 and its parameters into task packet 0 (e.g., it packages the first address of task0 and its parameters into task packet 0) and stores the task packet in the instruction FIFO. The vector processor, when not processing a task, reads task packet 0 from the instruction FIFO and executes task0 based on its parameters.
[0239] During the execution of the task by the vector processor, the scalar processor can re-obtain task1 and the parameters of task1 from the global memory. Task1 and the parameters of task1 are obtained from the global memory to prepare for the execution of task1. The scalar processor determines whether the instruction FIFO memory is full. If not, task1 and the parameters of task1 are packaged into task package 1, and the task package is stored in the instruction FIFO memory. In other words, the process of the scalar processor reading tasks and parameters and packaging them into task packages has nothing to do with whether the vector processor executes tasks. As long as the instruction FIFO memory is not full (that is, the number of task packages stored in the instruction FIFO memory does not reach 32), the scalar processor can repeatedly execute the steps of obtaining instructions and parameters from the global memory. After the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the steps of executing tasks based on the parameters, continuously obtaining tasks and parameters, and packaging them into task packages and storing them in the instruction FIFO memory.
[0240] However, whether the vector processor processes new tasks has nothing to do with the scalar processor. As long as the vector processor completes processing a task (at this time, the vector processor is not processing tasks) and the instruction FIFO memory is not empty, it can obtain a task packet from the instruction FIFO memory and then execute the task packet. If the vector processor completes processing a task (at this time, the vector processor is not processing tasks) but the instruction FIFO memory is empty, the task execution will be stopped.
[0241] In this heterogeneous multi-core processing process, the scalar processor and the vector processor do not need to be synchronized, and the task reading and dispatching of the scalar processor are independent of the task execution of the vector processor.
[0242] ● The execution condition is that both the instruction FIFO memory and the data FIFO memory are not full
[0243] Under these execution conditions, the scalar processor in a high-performance processor calls the vector processor to execute a task based on parameters. The scalar processor stores the instruction's address in the instruction FIFO and the parameter's address in the data FIFO. When not processing a task, the vector processor reads the instruction from the instruction FIFO and the parameters from the data FIFO, then executes the instruction based on the parameters.
[0244] Taking the case of two tasks (task0 and task1) that need to be executed as an example, for the heterogeneous multi-core processing process under this execution condition, an asynchronous queue instruction FIFO memory and an asynchronous queue data FIFO memory are constructed. The depth of the instruction FIFO memory and the data FIFO memory are both 32 bits, that is, the instruction FIFO memory can store 32 task packets simultaneously, and the data FIFO memory can also store 32 task packets simultaneously. At the same time, the instructions and parameters stored in the instruction FIFO memory and the data FIFO memory are corresponding (that is, if instruction 2 is stored in the second position of the instruction FIFO memory, the parameters of instruction 2 are also stored in the second position of the data FIFO memory). This ensures that the parameters read by the vector processor from the instruction FIFO memory and the data FIFO memory are the parameters of the read instructions.
[0245] First, the scalar processor gets task0 and its parameters from global memory.
[0246] The scalar processor determines whether both the instruction FIFO memory and the data FIFO memory are not full (since the instruction FIFO memory and the data FIFO memory correspond to each other during storage, if the instruction FIFO memory is not full, the data FIFO memory is not full either; if the instruction FIFO memory is full, the data FIFO memory is also full). If both the instruction FIFO memory and the data FIFO memory are not full, the scalar processor stores instruction 0 in the instruction FIFO memory and stores the storage address of the parameters of instruction 0 in the data FIFO memory (e.g., the scalar processor applies for a space on the local memory, stores the parameters of instruction 0 in the space, and stores the first address of the space in the data FIFO memory). When not processing a task, the vector processor reads task0 from the instruction FIFO memory, reads the parameters of task0 from the data FIFO memory, and executes task0 based on the parameters of task0.
[0247] While the vector processor is executing a task, the scalar processor can retrieve task1 and its parameters from the global memory again. Task1 and its parameters are retrieved from the global memory to prepare for the execution of task1. The scalar processor determines whether both the instruction FIFO memory and the data FIFO memory are not full. If not, the scalar processor stores instruction 1 in the instruction FIFO memory and the storage address of instruction 1's parameters in the data FIFO memory (e.g., the scalar processor allocates a space in the local memory, stores instruction 1's parameters in the space, and stores the first address of the space in the data FIFO memory). In other words, the process of the scalar processor reading tasks and parameters and storing them in the instruction FIFO memory and data FIFO memory respectively is independent of whether the vector processor is executing a task. As long as both the instruction FIFO memory and the data FIFO memory are not full (i.e., the number of task packets stored in the instruction FIFO memory does not reach 32, and the number of task packets stored in the data FIFO memory does not reach 32), the scalar processor can repeatedly retrieve instructions and parameters from the global memory. After the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the steps of executing the task based on the parameters, continuously obtains tasks and parameters, and stores the tasks and parameters in the instruction FIFO memory and the data FIFO memory respectively.
[0248] However, whether the vector processor processes new tasks has nothing to do with the scalar processor. As long as the vector processor completes processing a task (at this time, the vector processor is not processing tasks) and the instruction FIFO memory and data FIFO memory are not empty, the task and parameters can be obtained from the instruction FIFO memory and data FIFO memory respectively, and then the task is executed based on the parameters. If the vector processor completes processing a task (at this time, the vector processor is not processing tasks), but the instruction FIFO memory and data FIFO memory are empty, the task execution will be stopped.
[0249] In this heterogeneous multi-core data processing process, the scalar processor and the vector processor do not need to be synchronized. The task reading and dispatching of the scalar processor are independent of the task execution of the vector processor. At the same time, when the vector processor takes a task from the instruction FIFO memory, it also obtains the parameter address from the data FIFO memory, and then obtains the list of task actual parameters. In this way, the task is called and the actual parameters are obtained.
[0250] Furthermore, in its implementation, the scalar processor includes the thread identifier of each task it dispatches with the task. This identifier indicates which thread the task belongs to. This identifier allows the processor to determine whether each thread's task has completed execution, transforming synchronization into asynchrony, with the instruction FIFO memory serving as the asynchronous execution mechanism.
[0251] This embodiment provides a high-performance processor, which includes: a scalar processor, a vector processor, and a local memory; the scalar processor and the vector processor share the local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; a connection is established between the scalar processor and the vector processor; the scalar processor is connected to the global memory; the scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute a task based on the parameters. In the high-performance processor provided by this embodiment, after the scalar processor obtains instructions and parameters from the global memory and determines that the execution conditions are met, it calls the vector processor to execute a task based on the parameters, thereby coordinately using multiple heterogeneous cores to meet different computing needs.
[0252] Based on the same inventive concept of a high-performance processor, this embodiment provides a processor cluster, which includes multiple processors such as Figure 1 or Figure 2 High-performance processor shown.
[0253] For example, the high-performance processor includes: a scalar processor, a vector processor, and a local memory;
[0254] The scalar processor and the vector processor share local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor;
[0255] Establish connections between scalar processors and vector processors;
[0256] The scalar processor establishes a connection with the global memory;
[0257] The scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters.
[0258] Wherein, the scalar processor is an out-of-order multi-issue scalar processor;
[0259] The vector processor is a very long instruction word vector processor.
[0260] Among them, high-performance processors also include:
[0261] Instruction first-in-first-out FIFO memory;
[0262] Among them, the depth of the instruction FIFO memory is 32 bits.
[0263] Wherein, the high performance processor further comprises: an instruction FIFO memory storage and a data FIFO memory storage;
[0264] The depth of the instruction FIFO memory and the data FIFO memory are both 32 bits.
[0265] The execution condition is that the vector processor is not processing a task.
[0266] Among them, the execution condition is that the instruction FIFO memory is not full;
[0267] A scalar processor for storing the storage addresses of instructions and parameters into an instruction FIFO memory;
[0268] The vector processor is used to read instructions and data from the instruction FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0269] The execution condition is that both the instruction FIFO memory and the data FIFO memory are not full;
[0270] A scalar processor, configured to store the memory address of the instruction into the instruction FIFO memory and store the memory address of the parameter into the data FIFO memory;
[0271] The vector processor is configured to read instructions from the instruction FIFO memory and data from the data FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0272] The task includes the identifier of the thread to which the task belongs.
[0273] The processor cluster provided in this embodiment includes a high-performance processor including: a scalar processor, a vector processor, and a local memory; the scalar processor and the vector processor share the local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; a connection is established between the scalar processor and the vector processor; the scalar processor establishes a connection with the global memory; the scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute a task based on the parameters, thereby coordinately using multiple heterogeneous cores to meet different computing needs.
[0274] Based on the same inventive concept of a high-performance processor, this embodiment provides an electronic device, which includes Figure 1 or Figure 2 The high performance processor shown, or, includes one or more processor clusters, wherein each processor cluster includes multiple processors such as Figure 1 or Figure 2 High-performance processor shown.
[0275] For example, the high-performance processor includes: a scalar processor, a vector processor, and a local memory;
[0276] The scalar processor and the vector processor share local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor;
[0277] Establish connections between scalar processors and vector processors;
[0278] The scalar processor establishes a connection with the global memory;
[0279] The scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute the task based on the parameters.
[0280] Wherein, the scalar processor is an out-of-order multi-issue scalar processor;
[0281] The vector processor is a very long instruction word vector processor.
[0282] Among them, high-performance processors also include:
[0283] Instruction first-in-first-out FIFO memory;
[0284] Among them, the depth of the instruction FIFO memory is 32 bits.
[0285] Wherein, the high performance processor further comprises: an instruction FIFO memory storage and a data FIFO memory storage;
[0286] The depth of the instruction FIFO memory and the data FIFO memory are both 32 bits.
[0287] The execution condition is that the vector processor is not processing a task.
[0288] Among them, the execution condition is that the instruction FIFO memory is not full;
[0289] A scalar processor for storing the storage addresses of instructions and parameters into an instruction FIFO memory;
[0290] The vector processor is used to read instructions and data from the instruction FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0291] The execution condition is that both the instruction FIFO memory and the data FIFO memory are not full;
[0292] A scalar processor, configured to store the memory address of the instruction into the instruction FIFO memory and store the memory address of the parameter into the data FIFO memory;
[0293] The vector processor is configured to read instructions from the instruction FIFO memory and data from the data FIFO memory when no task processing is being performed, and execute instructions based on the data.
[0294] The task includes the identifier of the thread to which the task belongs.
[0295] The electronic device provided by this embodiment includes a high-performance processor including: a scalar processor, a vector processor, and a local memory; the scalar processor and the vector processor share the local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; a connection is established between the scalar processor and the vector processor; the scalar processor establishes a connection with the global memory; the scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution conditions are met, it calls the vector processor to execute a task based on the parameters, thereby coordinately using multiple heterogeneous cores to meet different computing needs.
[0296] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.
[0297] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0298] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0299] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0300] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0301] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A high-performance processor, characterized in that: The high-performance processor includes: a scalar processor, a vector processor, and a local memory; The scalar processor and the vector processor share a local memory; and the vector processor only accesses the local memory and is only called and executed by the scalar processor; Establishing a connection between the scalar processor and the vector processor; The scalar processor establishes a connection with the global memory; The scalar processor is used to obtain instructions and parameters from the global memory; after the scalar processor determines that the execution condition is met, it calls the vector processor to execute the task based on the parameters.
2. The high-performance processor according to claim 1, wherein: The scalar processor is an out-of-order multi-issue scalar processor; The vector processor is a very long instruction word vector processor.
3. The high-performance processor according to claim 1, wherein: The high-performance processor further includes: Instruction first-in-first-out FIFO memory; Among them, the depth of the instruction FIFO memory is 32 bits.
4. The high-performance processor according to claim 1, wherein: The high-performance processor further includes: an instruction FIFO memory storage and a data FIFO memory storage; The depth of the instruction FIFO memory and the data FIFO memory are both 32 bits.
5. The high-performance processor according to claim 1, wherein: The execution condition is that the vector processor is not performing task processing.
6. The high-performance processor according to claim 3, characterized in that The execution condition is that the instruction FIFO memory is not full; The scalar processor is used to store the storage addresses of the instructions and parameters into the instruction FIFO memory; The vector processor is configured to read instructions and data from the instruction FIFO memory when no task processing is being performed, and execute the instructions based on the data.
7. The high-performance processor according to claim 4, characterized in that The execution condition is that both the instruction FIFO memory and the data FIFO memory are not full; The scalar processor is configured to store the storage address of the instruction into an instruction FIFO memory and store the storage address of the parameter into a data FIFO memory; The vector processor is configured to read instructions from the instruction FIFO memory and data from the data FIFO memory when no task processing is being performed, and execute the instructions based on the data.
8. The high-performance processor according to claim 1, wherein: The task includes an identifier of the thread to which the task belongs.
9. A processor cluster, comprising a plurality of high-performance processors according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: The high-performance processor according to any one of claims 1 to 7, or comprising one or more processor clusters according to claim 9.
Citation Information
Patent Citations
Video processor having scalar and vector components for controlling video processing
CN101371233A
An apparatusDevice and method for executing performing vector transcendental function operations
CN111651200A
Tensor, vector and scalar calculation acceleration and data scheduling system
CN115169541A
Low-hardware-overhead vector processor architecture based on RISC-V vector instruction extension
CN116521229A
Data processing method applied to processor, processor and storage medium
CN117369871A
Cited By
High-performance processing method and electronic equipment
CN120540706A
High performance processing method and electronic device
CN120540706B