A matrix-vector processor and a matrix-vector collaborative computing method

By adopting a one-main dual-cooperation architecture and a RISC-V compatible design in the processor, the problems of multi-core resource redundancy, inefficient communication between cores and closed private instruction sets are solved, and efficient matrix-vector computing and good ecological compatibility are achieved.

CN120067036BActive Publication Date: 2025-07-01SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510550151.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-01
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

In the prior art, multi-core parallel architecture and dedicated accelerators have problems such as resource redundancy, inefficient communication between cores and ecological closure of private instruction sets when improving computing efficiency.

Method used

The matrix-vector processor adopts a main dual-cooperation architecture, through the division of labor between the main processor core, matrix and vector coprocessor core, it reduces resource redundancy, improves communication efficiency between cores, and is compatible with the RISC-V open source ecosystem, breaking closed restrictions.

Benefits of technology

It significantly improves the utilization rate and energy efficiency ratio of computing resources, reduces communication latency between cores, improves the efficiency of computing pipelines, and enhances the technical universality and ecological compatibility of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067036B_ABST
    Figure CN120067036B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of processor design, and particularly to a matrix-vector processor and a matrix-vector collaborative computing method. The matrix-vector processor includes a main processor core, a matrix coprocessor core and a vector coprocessor core connected thereto, and an L2 cache connected to the main processor core, the matrix coprocessor core and the vector coprocessor core. The matrix-vector collaborative computing method includes the L1 cache obtaining instructions and data from the L2 cache and initially distinguishing the instruction types. For conventional computing instructions, matrix computing instructions and vector computing instructions, the main processor core, the matrix coprocessor core and the vector coprocessor core respectively execute the computations to obtain computation results. The computation results are written back to the L2 cache to update the resource status of the scoreboard module. This application can reduce resource redundancy, improve communication efficiency, break the closed restrictions, and achieve efficient elastic computing by combining hybrid scheduling and modular expansion.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of artificial intelligence, high-performance computing, and machine learning applications, matrix and vector operations have become the core load scenarios in processor design. Traditional general-purpose processors are limited by the memory wall problem of the von Neumann architecture and the parallelism bottleneck of scalar computing units, making it difficult to meet the high-throughput requirements of large-scale matrix multiplication, tensor operations, etc. At the same time, the rise of the RISC-V open instruction set ecosystem (such as RVV1.0) provides a new direction for customized computing architectures, but there are still technical challenges in achieving efficient matrix-vector collaborative computing based on open instruction sets.

[0003] In current technologies, the industry mainly improves computing efficiency through multi-core parallel architectures and dedicated accelerators. For example, multi-core SIMD architectures achieve thread-level parallelism by deploying multiple computing cores (such as the ARM Neoverse V series), or accelerate matrix operations by extending the RVV1.0 instruction set combined with private hardware units (such as Tenstorrent's Grayskull architecture). In addition, heterogeneous computing architectures (such as the CPU+GPU collaboration of NVIDIA Grace Hopper) separate the functions of the main processor and the accelerator through task offloading and high-speed interconnection technologies, further improving computing throughput. These solutions significantly optimize computing performance in specific scenarios.

[0004] However, the existing technologies still have the following bottlenecks: First, multiple computing cores in the multi-core architecture need to repeatedly execute the same instruction set, resulting in waste of on-chip resources and a decrease in energy efficiency ratio; Second, the inter-core instruction scheduling depends on complex bus protocols and software layer coordination, leading to low instruction fetch and decoding efficiency and computing latency problems; Third, mainstream solutions rely on private instruction sets (such as NVIDIA CUDA), which have poor compatibility with the RISC-V open ecosystem and limit the universality of the technology; Fourth, in traditional heterogeneous architectures, data needs to be transferred through the main memory multiple times, and the inter-core communication bandwidth becomes a performance bottleneck. Summary of the Invention

[0005] Aiming at the technical problems of multi-core resource redundancy, low inter-core communication efficiency, and closed private instruction set ecosystem dependence existing in the existing methods for improving computing efficiency through multi-core parallel architectures and dedicated accelerators, this application provides a matrix-vector processor and a matrix-vector collaborative computing method, which reduce resource redundancy through a one-master-two-slave architecture, improve communication efficiency through direct inter-core channels, break the closed restrictions by being compatible with the RISC-V ecosystem, and achieve efficient elastic computing by combining hybrid scheduling and modular expansion.

[0006] In a first aspect, this application provides a matrix-vector processor, including:

[0007] A main processor core for executing general computing instructions and performing preliminary decoding and scheduling for matrix computing instructions and vector computing instructions, including an L1 cache, an instruction fetch module, a warp allocation module, a decoding module, an issue cache module, a scoreboard module, a main processor register file module, a task issue module, a main processor computing unit, a coprocessor task scheduling module, and a computation result write-back module;

[0008] A matrix coprocessor core connected to the main processor core for executing matrix computing instructions, including a matrix secondary decoding module, a scalable number of matrix computing unit groups, and a matrix LSU storage unit;

[0009] A vector coprocessor core connected to the main processor core for executing vector computing instructions, including a vector secondary decoding module, a scalable number of vector computing unit groups, and a vector LSU storage unit;

[0010] An L2 cache connected to the L1 cache, the matrix LSU storage unit, and the vector LSU storage unit respectively;

[0011] Wherein, a one-way direct inter-core connection channel is provided between the matrix LSU storage unit and the vector LSU storage unit, and only allows the vector coprocessor core to transmit data to the matrix coprocessor core;

[0012] The scoreboard module is jointly maintained by the main processor core, the matrix coprocessor core, and the vector coprocessor core, and is configured such that the main processor core supports out-of-order instruction issue, and the matrix coprocessor core and the vector coprocessor core only support in-order instruction issue.

[0013] Furthermore, it should be noted that the main processor core adopts a five-stage pipeline structure, including instruction fetch, decoding, issue, execution, result collection, and write-back stages; the matrix coprocessor core and the vector coprocessor core both adopt a three-stage pipeline structure, including decoding-secondary scheduling, execution, result collection, and write-back stages.

[0014] Furthermore, it should be noted that the decoding module is configured to:

[0015] Parse the opcode, operands, and addressing mode of the instruction;

[0016] Match the instruction type through the opcode field, distinguish general computing instructions, vector computing instructions, or matrix computing instructions, and add an extended instruction set mark to the vector computing instructions or matrix computing instructions.

[0017] Furthermore, it should be noted that the coprocessor task scheduling module is used to select the target coprocessor core according to the type of the extended instruction set mark, and is respectively connected to the matrix secondary decoding module and the vector secondary decoding module through the RoCC interface to form a one-to-two instruction distribution architecture.

[0018] Further, it should be noted that the L1 cache includes an instruction cache and a data cache.

[0019] Further, it should be noted that the L2 cache is connected to the L1 cache of the main processor core, the LSU storage unit of the matrix co-processor core, and the LSU storage unit of the vector co-processor core through the AXI4 bus.

[0020] Further, it should be noted that the matrix calculation unit group and the vector calculation unit group adopt a modular array structure. Each matrix calculation unit group or vector calculation unit group contains multiple parallel calculation units, and data interconnection between the calculation unit groups is achieved through a data distribution unit and a process monitoring and distribution unit.

[0021] Further, it should be noted that the matrix calculation unit group includes 32 matrix calculation units, and the vector calculation unit group includes 32 vector calculation units.

[0022] Further, it should be noted that the matrix calculation unit includes a matrix processing register file module, a matrix quasi-processing task queue module, a matrix quasi-processing data queue module, a matrix calculation data splicing module, a calculation result matrix splitting module, a matrix calculation scheduling arbitration module, and 4 matrix calculation iterative execution modules;

[0023] The vector calculation unit includes a vector processing register file module, a vector quasi-processing task queue module, a vector quasi-processing data queue module, a vector calculation scheduling arbitration module, and 4 vector calculation iterative execution modules.

[0024] Further, it should be noted that the matrix processing register file module is constructed by SRAM and is used to store the operand data necessary for matrix calculation. The SRAM contains 16 storage partitions, and each group of 4 partitions stores 4 types of operand categories of a matrix calculation iterative execution module;

[0025] The matrix calculation iterative execution module is used to execute matrix calculation and includes a matrix ALU, a matrix FPU, and a matrix MUL;

[0026] The matrix quasi-processing task queue module and the matrix quasi-processing data queue module are used to sequentially cache the task information and quasi-processing data involved in the matrix calculation tasks according to the instruction execution order;

[0027] The matrix calculation scheduling arbitration module is used to verify whether the ID of this group of matrix calculation units allocated by the matrix co-processor process monitoring and distribution unit is idle, and sequentially read the quasi-processing instructions from the instruction data queue, read the quasi-processing data from the quasi-processing data queue, splice them, and send them to the target matrix calculation unit;

[0028] The matrix calculation data splicing module is used to read the original operation data from the data queue to be processed according to the decoded calculation instructions, and splice the data to be processed into the matrix calculation format according to the instruction requirements;

[0029] The calculation result matrix splitting module is used to perform format splitting and conversion on the calculation results received from the 4 matrix calculation iterative execution modules, and split the calculation results in matrix format into the sequential association format that can be directly stored in SRAM.

[0030] It should be further noted that the vector processing register file module is constructed by SRAM and is used to store the operand data necessary for vector calculation. SRAM contains 16 storage partitions, and each group of 4 partitions stores 4 types of operand categories of a vector calculation iterative execution module;

[0031] The vector calculation iterative execution module is used to perform vector calculation and includes a vector ALU, a vector FPU, and a vector MUL;

[0032] The vector data to be processed task queue module and the vector data to be processed queue module are used to sequentially cache the task information and the data to be processed involved in the vector calculation tasks according to the instruction execution order;

[0033] The vector calculation scheduling arbitration module is used to verify whether the group of vector calculation unit IDs assigned by the vector coprocessor process monitoring and allocation unit is idle, and sequentially read the instructions to be processed from the instruction data queue, read the data to be processed from the data queue to be processed, splice them and send them to the target vector calculation unit.

[0034] In a second aspect, the present application provides a matrix-vector collaborative calculation method, which uses the above matrix-vector processor to perform matrix-vector collaborative calculation, including:

[0035] S1. The L1 cache obtains instructions and data from the L2 cache and preliminarily distinguishes the instruction types through the decoding module;

[0036] S2. If the instruction is a conventional calculation instruction, it is executed by the main processor calculation unit to obtain a conventional calculation result;

[0037] If the instruction is a matrix calculation instruction, it is distributed by the coprocessor task scheduling module to the matrix coprocessor core. The matrix coprocessor core performs secondary decoding on the received instruction and allocates it to the matrix calculation unit group for execution to obtain a matrix calculation result;

[0038] If the instruction is a vector calculation instruction, it is distributed by the coprocessor task scheduling module to the vector coprocessor core. The vector coprocessor core performs secondary decoding on the received instruction and allocates it to the vector calculation unit group for execution to obtain a vector calculation result;

[0039] S3. Write the calculation result back to the L2 cache and update the resource status of the scoreboard module.

[0040] Furthermore, it should be noted that in step S1, the preliminary differentiation of instruction types includes: parsing the opcode field of the instruction. If it matches the vector extension instruction set format, it is marked as a vector calculation instruction; if it matches the matrix extension instruction set format, it is marked as a matrix calculation instruction; if it does not match both the vector extension instruction set format and the matrix extension instruction set format, it is determined as a regular calculation instruction.

[0041] Furthermore, it should be noted that during the calculation process of the vector coprocessor core, if matrix calculation is required, the intermediate data is directly transmitted to the LSU storage unit of the matrix coprocessor core through the direct inter-core connection channel for further matrix calculation.

[0042] Furthermore, it should be noted that in the matrix calculation unit group, the steps for the matrix calculation unit to perform calculations include:

[0043] S201. Receive the data to be processed from the matrix data to be processed queue module and store it in the SRAM storage partition of the matrix processing register file module;

[0044] S202. Receive the tasks and instructions to be processed through the matrix task to be processed queue module;

[0045] S203. The matrix calculation data splicing module reads data from the matrix data to be processed queue module and splices it;

[0046] S204. The matrix calculation scheduling arbitration module verifies the idle status of the current matrix calculation unit ID;

[0047] S205. Distribute the spliced instruction data to the target matrix calculation iterative execution module;

[0048] S206. The calculation result matrix splitting module converts the calculation result and writes it back to the matrix processing register file module.

[0049] Furthermore, it should be noted that in the vector calculation unit group, the steps for the vector calculation unit to perform calculations include:

[0050] S211. Receive the data to be processed from the vector data to be processed queue module and store it in the SRAM storage partition of the vector processing register file module;

[0051] S212. Receive the tasks and instructions to be processed through the vector task to be processed queue module;

[0052] S213. The vector calculation scheduling arbitration module verifies the idle status of the current vector calculation unit ID;

[0053] S214. Distribute the instruction data to the target vector calculation iterative execution module;

[0054] S215. Write the calculation result back to the corresponding SRAM storage partition of the vector processing register file module.

[0055] As can be seen from the above technical solutions, the present application has the following advantages:

[0056] 1. Through the "one master - two co - processors" architecture of the main processor core and the dual co - processor cores, the matrix and vector calculation tasks are separated from the main processor and executed by the dedicated co - processor cores. The main processor is only responsible for preliminary decoding and global scheduling, avoiding resource redundancy caused by multiple computing cores repeatedly executing the same instruction set in the multi - core architecture, reducing the power consumption of idle logical units, and significantly improving the utilization rate of computing resources and the energy efficiency ratio.

[0057] 2. The present application constructs a one - way direct inter - core connection channel between the matrix co - processor core and the vector co - processor core. The intermediate calculation results of the vector co - processor can be directly transmitted to the matrix co - processor core, bypassing the transfer bottleneck of the main memory or bus protocol in the traditional heterogeneous architecture, realizing seamless connection of matrix - vector cascade operations, significantly reducing the inter - core communication delay, and improving the efficiency of the computing pipeline.

[0058] 3. In the present application, the instruction parsing, task distribution, and interface protocol of the main processor core and the co - processor cores are compatible with the RISC - V open - source ecosystem standard, support standard extensions including the RVV1.0 vector instruction set, and support developers to directly call the RISC - V standard toolchain for programming and function extension, solving problems such as high development threshold and difficult cross - platform adaptation caused by the closed private instruction set ecosystem, and significantly improving the technical universality and ecological compatibility of the processor.

[0059] 4. Through the scoreboard module jointly maintained by the main processor core and the dual co - processor cores, the present application synchronizes the task status and resource occupancy of the three types of processor cores in real - time. In the mixed scheduling mode of out - of - order issue of the main processor and in - order execution of the co - processor, it avoids pipeline stalls caused by task conflicts, ensuring stable execution and high - efficiency throughput of large - scale mixed operations.

[0060] 5. The present application sets up a modular and extensible matrix / vector calculation unit group. By dynamically allocating the number of calculation units, it flexibly adapts to matrix - vector operation requirements of different scales, maximizes the utilization rate of hardware resources while ensuring calculation efficiency, and meets the elastic computing power requirements from edge computing to cloud servers. Description of the Drawings

[0061] To more clearly illustrate the technical solutions of the present application, the accompanying drawings required for the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0062] Figure 1 It is a schematic diagram of the architecture of a matrix-vector processor in an embodiment of the present application.

[0063] Figure 2 It is a schematic diagram of the architecture of a matrix calculation unit in an embodiment of the present application.

[0064] Figure 3 It is a schematic diagram of the architecture of a vector calculation unit in an embodiment of the present application.

[0065] Figure 4 It is a flowchart of a matrix-vector processing method in an embodiment of the present application.

[0066] Figure 5 It is a flowchart of a matrix calculation unit performing calculations in an embodiment of the present application.

[0067] Figure 6 It is a flowchart of a moment vector calculation unit performing calculations in an embodiment of the present application. Detailed implementation manners

[0068] To make the application purpose, features, and advantages of the present application more obvious and understandable, the technical solutions protected by the present application will be clearly and completely described below by using specific embodiments and the accompanying drawings. Obviously, the embodiments described below are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in this patent, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this patent.

[0069] The matrix-vector processor and matrix-vector collaborative computing method involved in this application mainly target the field of processor design technology. Through the "one master - two co-processors" architecture of the main processor core and dual co-processor cores, the matrix and vector computing tasks are separated from the main processor and executed by dedicated co-processor cores. The main processor is only responsible for preliminary decoding and global scheduling, avoiding resource redundancy caused by multiple computing cores repeating the same instruction set in a multi-core architecture, reducing the power consumption of idle logical units, and significantly improving the utilization rate of computing resources and energy efficiency ratio; A one-way direct inter-core connection channel is built between the matrix co-processor core and the vector co-processor core. The intermediate calculation results of the vector co-processor can be directly transmitted to the matrix co-processor core, bypassing the transfer bottleneck of the main memory or bus protocol in the traditional heterogeneous architecture, realizing seamless connection of matrix-vector cascade operations, significantly reducing inter-core communication latency, and improving the efficiency of the computing pipeline; The instruction parsing, task distribution, and interface protocol of the main processor core and co-processor cores are compatible with the RISC-V open source ecosystem standard, supporting standard extensions including the RVV1.0 vector instruction set, and supporting developers to directly call the RISC-V standard toolchain for programming and function extension, solving problems such as high development thresholds and difficult cross-platform adaptation caused by the closed private instruction set ecosystem, and significantly improving the technical universality and ecological compatibility of the processor; Through the scoreboard module jointly maintained by the main processor core and dual co-processor cores, the task status and resource occupancy of the three types of processor cores are synchronized in real time. In the mixed scheduling mode of out-of-order issue of the main processor and in-order execution of the co-processor, pipeline stalls caused by task conflicts are avoided, ensuring the stable execution and high throughput of large-scale mixed operations; Modular and extensible matrix / vector computing unit groups are set up. By dynamically allocating the number of computing units, different scales of matrix-vector operation requirements are flexibly adapted, maximizing the utilization rate of hardware resources while ensuring computing efficiency, and meeting the elastic computing power requirements from edge computing to cloud servers.

[0070] The matrix-vector processor and matrix-vector collaborative computing method involved in this application mainly address the technical problems of multi-core resource redundancy, low inter-core communication efficiency, and dependence on the closed private instruction set ecosystem existing in the existing methods of improving computing efficiency through multi-core parallel architectures and dedicated accelerators.

[0071] The matrix-vector processor and matrix-vector collaborative computing method involved in this application will be described in detail below. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details.

[0072] In the matrix-vector processor and the matrix-vector collaborative computing method involved in this application, the term "including" indicates the existence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their sets. The terms "including", "comprising", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0073] For the convenience of clearly describing the technical solutions of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily limit to be different.

[0074] The statements such as "an embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of this application. Thus, the statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different places in this application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0075] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0076] The matrix-vector collaborative computing method provided by the embodiments of this application is executed by a computer device. Correspondingly, the matrix-vector processor runs in the computer device.

[0077] The following are some noun explanations in this solution for better understanding of this solution:

[0078] L1 cache: Refers to the primary high-speed cache inside the CPU core, directly connected to the processor core, used to store recently used instructions and data, divided into instruction cache (L1i) and data cache (L1d) to support the Harvard architecture. Its access latency is extremely low (usually 1-4 clock cycles), the capacity is small (about 16KB-64KB per core), implemented by SRAM, and the physical address or virtual address mapping strategy (such as set associative) is used to balance the hit rate and power consumption. It is the fastest layer in the multi-level cache system and directly affects the single-threaded performance.

[0079] L2 Cache: Refers to the intermediate-level cache located between the CPU core and the shared last-level cache (LLC). It has a relatively large capacity (about 256KB - 1MB per core), and can be exclusive to a single core or shared by multiple cores. Its latency is higher than that of L1 (about 10 - 20 cycles), but its bandwidth and capacity advantages enable it to buffer requests that miss in L1, reducing the frequency of accessing the main memory. In modern multi-core processors, L2 often serves as a bridge for data sharing between cores and maintains multi-core cache coherence through coherence protocols (such as MESI), achieving a trade-off between energy consumption and performance.

[0080] LSU (Load-Store Unit): It is a functional unit in the CPU pipeline that specifically manages memory accesses, responsible for performing load (reading data from memory into registers) and store (writing register data to memory) operations. Its core functions include generating valid addresses, handling memory alignment, initiating cache or bus requests, and translating virtual addresses to physical addresses through address translation (MMU). High-performance LSUs support out-of-order execution, prefetching, and memory dependence prediction to hide memory access latency, which is crucial for applications that rely on memory bandwidth (such as databases and scientific computing).

[0081] SRAM (Static Random-Access Memory): A type of static memory based on flip-flop circuits that can retain data without dynamic refreshing. It has extremely fast access speeds (nanosecond level) and relatively high power consumption. Each memory cell consists of 6 transistors, and its complex structure results in lower density and higher cost. It is commonly used in CPU caches, register files, and speed-sensitive on-chip storage (such as Block RAM in FPGAs). Compared with DRAM, SRAM loses data after power-off, but does not require complex controller design and is suitable for small-capacity high-speed cache scenarios.

[0082] CoCC Interface (Coherent Cache Controller Interface): The CoCC interface is the core mechanism for implementing cache coherence in multi-core processors or systems-on-chip (SoCs). It is mainly used to coordinate data consistency among multiple processor cores, caches, or hardware accelerators. In heterogeneous computing scenarios, different cores may have independent caches, and data inconsistency can lead to logical errors (such as dirty data or reading stale values). CoCC solves this problem by implementing cache coherence protocols (such as MESI, MOESI). Its functions include snooping on the cache operations of other cores, managing the states of cache lines (such as invalidation or update), and handling read / write request conflicts. At the same time, the CoCC interface needs to consider low-latency and high-throughput designs, such as optimizing bus arbitration, supporting multi-level cache expansion, or parallel transaction processing, to meet the performance requirements of multi-core chips. Typical applications include multi-core CPUs, GPU clusters, AI acceleration chips, etc., which require efficient shared memory systems.

[0083] AXI4 Bus (Advanced eXtensible Interface 4): AXI4 is a high-performance on-chip interconnection protocol in the AMBA 4.0 standard introduced by ARM. It is designed for communication between processors, memory controllers, peripherals, and hardware accelerators. Its core features include separate read and write channels (independent transmission of address, data, and response), support for burst transmission (the maximum burst length of AXI4 is 256), out-of-order transaction processing (identified by transaction ID), multi-master multi-slave architecture, and low-power control mechanism. AXI4 provides three variants: standard AXI4 for high-bandwidth memory access, AXI4-Lite simplified version for simple scenarios such as register configuration, and AXI4-Stream for addressless streaming data transmission (such as video streams). This bus significantly improves throughput through designs such as parallelizing transmission paths and supporting outstanding transactions, and ensures reliability through standardized handshake signals (VALID / READY). It is widely used in mobile phone SoCs, FPGA heterogeneous systems, and data center chips, becoming the de facto standard for complex chip interconnection.

[0084] ALU (Arithmetic Logic Unit): Refers to the arithmetic logic unit, which is the core component of the CPU and is responsible for performing integer arithmetic operations (such as addition, subtraction, multiplication, and division) and logical operations (such as AND, OR, NOT, shift). Its inputs come from registers and immediate values, and the output results directly affect the program status register (such as the overflow flag). Modern ALUs usually support single-cycle completion of simple operations, and complex operations (such as division) may require multiple cycles or co-processor assistance. Multiple ALUs can be integrated in a multi-issue architecture to achieve instruction-level parallelism and improve IPC (instructions per cycle).

[0085] FPU (Floating - Point Unit): Refers to the floating - point arithmetic unit, which is designed specifically for high - precision floating - point calculations and supports addition, subtraction, multiplication, division, square root, and transcendental functions (such as trigonometric functions) according to the IEEE 754 standard. In the early days, it existed as an independent coprocessor (such as x87), and modern CPUs mostly integrate it into the core, adopting pipelined design and SIMD instruction sets (such as AVX, NEON) to accelerate vectorized floating - point operations. The performance of the FPU directly affects scenarios such as scientific computing, graphics rendering, and AI inference, and its hardware implementation needs to balance precision, throughput, and power consumption (such as optimizing with fused multiply - add FMA instructions).

[0086] MUL (Multiplier Unit): Refers to the multiplier, which is a dedicated hardware unit for performing integer multiplication operations and can exist independently or be integrated into the ALU / FPU. Its implementation methods include iterative multiplication (completed in multiple cycles), array multiplication (parallel partial product accumulation), and look - up table methods (such as optimized by Booth encoding). High - performance designs use carry - save adders or Wallace trees to shorten the critical path. In modern processors, MUL often supports single - cycle completion of small - bit - width multiplications (such as 32x32 bits), and large - bit - width operations (such as 64x64 bits) may be split into multiple cycles or rely on SIMD extension instructions (such as SSE / AVX).

[0087] SFU (Special Function Unit): It is a dedicated hardware unit in the processor for performing complex mathematical operations, mainly dealing with non - scalar calculations such as trigonometric functions (sin / cos), exponential functions, and logarithmic operations. It is particularly important in GPUs and scientific computing accelerators, as these operations would cause performance bottlenecks if processed by a general - purpose ALU. The SFU achieves low latency and high throughput through highly optimized hardware circuits (such as polynomial approximation or look - up table methods), significantly improving the efficiency of graphics rendering, physical simulation, and AI inference. For example, in the GPU shader core, the SFU is responsible for lighting calculations and geometric transformations, directly determining the rendering speed and quality.

[0088] LD (Load): The LD instruction is responsible for loading data from memory or cache into registers and is a key entry point for the data stream. Its operations include generating a valid address, accessing the memory hierarchy (which may trigger cache hits / misses), and handling alignment and permission checks (such as MMU address translation). Modern processors hide the LD latency through prefetching, out - of - order execution, and non - blocking cache design. For example, in a superscalar pipeline, the LD instruction can be issued in parallel with other operations to reduce stalls. The performance of LD directly affects the response speed of data - intensive applications (such as databases, real - time signal processing).

[0089] ST (Store): The ST instruction writes the register data back to memory or cache to persist the calculation results. This process needs to handle write combining (combining multiple small writes), cache coherence (such as the MESI protocol maintaining the multi-core state), and write buffer management. To reduce the blocking of the ST on the pipeline, the processor often adopts the Write-Through or Write-Back strategy and combines a write queue to buffer multiple write requests. In a weak memory order architecture, the ST may need barrier instructions to ensure data visibility and avoid race conditions in a multi-threaded environment. The efficiency of ST is crucial for high-frequency update scenarios (such as high-frequency trading, logging systems).

[0090] In some specific embodiments, the matrix-vector processor includes: a main processor core for executing general computing instructions and completing preliminary decoding and scheduling for matrix calculation instructions and vector calculation instructions, including an L1 cache, an instruction fetch module, a warp allocation module, a decoding module, an issue cache module, a scoreboard module, a main processor register file module, a task issue module, a main processor computing unit, a coprocessor task scheduling module, and a calculation result write-back module;

[0091] A matrix coprocessor core connected to the main processor core for executing matrix calculation instructions, including a matrix secondary decoding module, an expandable number of matrix computing unit groups, and a matrix LSU storage unit;

[0092] A vector coprocessor core connected to the main processor core for executing vector calculation instructions, including a vector secondary decoding module, an expandable number of vector computing unit groups, and a vector LSU storage unit;

[0093] An L2 cache connected to the L1 cache, the matrix LSU storage unit, and the vector LSU storage unit respectively;

[0094] Among them, a one-way direct inter-core connection channel is set between the matrix LSU storage unit and the vector LSU storage unit, allowing only the vector coprocessor core to transmit data to the matrix coprocessor core;

[0095] The scoreboard module is jointly maintained by the main processor core, the matrix coprocessor core, and the vector coprocessor core, and is configured such that the main processor core supports out-of-order instruction issue, and the matrix coprocessor core and the vector coprocessor core only support in-order instruction issue.

[0096] In this embodiment, the L1 cache is used to provide fast data access and reduce the memory access latency of the main processor computing unit, and can be implemented relying on SRAM;

[0097] The instruction fetch module is used to prefetch the instruction stream from the memory to ensure the continuity of the instruction supply, and can be implemented relying on an instruction prefetch buffer / PC register;

[0098] The warp allocation module is used to dynamically allocate thread resources to computing units and can be implemented relying on a multi-threaded scheduler / state machine;

[0099] The decoding module is used to parse the instruction type and mark the extended instruction set and can be implemented relying on an instruction decoder / u-code controller;

[0100] The issue cache module is used to temporarily store the decoded instructions waiting for issue and can be implemented relying on an instruction queue / FIFO buffer;

[0101] The scoreboard module is used to track the resource occupancy status to avoid conflicts and can be implemented relying on status registers / resource dependency tables;

[0102] The main processor register file is used to store the operands and status of scalar operations and can be implemented relying on a general-purpose register file;

[0103] The task issue module is used to distribute instructions to computing units or coprocessors and can be implemented relying on an instruction distributor / crossbar switch;

[0104] The main processor computing unit is used to execute scalar operations and can be implemented relying on an ALU / FPU;

[0105] The coprocessor task scheduling module is used to select the target coprocessor and transfer instructions and can be implemented relying on a RoCC interface controller;

[0106] The computation result write-back module is used to write the results into the cache or registers and can be implemented relying on a write-back bus / data path;

[0107] The matrix secondary decoding module is used to parse the detailed operation mode of matrix instructions and can be implemented relying on dedicated decoding logic / u-op generator;

[0108] The matrix computing unit group is used to perform parallel computations such as matrix multiplication / decomposition and can be implemented relying on a systolic array / matrix ALU cluster;

[0109] The matrix LSU storage unit is used to manage matrix data access and can be implemented relying on a matrix-specific DMA controller;

[0110] The vector secondary decoding module is used to parse the data bit width and operation type of vector instructions and can be implemented relying on a vector instruction decoder;

[0111] The vector computing unit group is used to perform SIMD vector operations and can be implemented relying on a vector ALU / SIMD processing unit;

[0112] The vector LSU storage unit is used to handle vector data load / store and can be implemented relying on a vector DMA / stride access controller;

[0113] The L2 cache is used as a data transfer medium between the main memory and the L1 cache, and can be implemented relying on SRAM / 3D stacked memory;

[0114] The inter-core direct connection channel is used to achieve zero-copy data transfer from vector to matrix, and can be implemented relying on NoC routing nodes / specialized data buses.

[0115] Through the division of labor architecture design of the main processor core, matrix, and vector coprocessor cores, the main processor core is responsible for the execution of regular computing instructions and preliminary decoding scheduling. The matrix and vector coprocessor cores respectively process dedicated computing tasks, combine the inter-core unidirectional direct connection channel to achieve zero-copy data transfer from vector to matrix, and use the shared scoreboard module to collaboratively track the resource status, realizing the efficient collaborative scheduling of heterogeneous computing resources, reducing task conflicts and redundant data transfer, and significantly improving the overall parallel efficiency and computing throughput of the system.

[0116] In some specific embodiments, the main processor core adopts a five-stage pipeline structure, including instruction fetch, decoding, issue, execution, result collection, and write-back stages; both the matrix coprocessor core and the vector coprocessor core adopt a three-stage pipeline structure, including decoding-secondary scheduling, execution, result collection, and write-back stages.

[0117] In some specific embodiments, the decoding module is configured to:

[0118] Parse the opcode, operands, and addressing mode of the instruction;

[0119] Match the instruction type through the opcode field, distinguish regular computing instructions, vector computing instructions, or matrix computing instructions, and add extended instruction set tags to vector computing instructions or matrix computing instructions.

[0120] Through the opcode parsing and extended instruction set tagging mechanism of the decoding module, accurately distinguish the types of regular, matrix, and vector instructions at the instruction processing front end of the main processor core, provide clear instruction classification identifiers for subsequent coprocessor task scheduling, reduce the instruction interaction delay between the main processor core and the coprocessor core, and at the same time avoid resource waste caused by incorrect instruction judgment, improving the processing efficiency of heterogeneous computing instruction streams and the system response speed.

[0121] In some specific embodiments, the coprocessor task scheduling module is used to select the target coprocessor core according to the type of the extended instruction set tag, and is respectively connected to the matrix secondary decoding module and the vector secondary decoding module through the RoCC interface, forming an instruction distribution architecture of one hanging two.

[0122] Adopt a one-for-two instruction distribution architecture based on the RoCC interface. The matrix and vector instructions are respectively routed to the corresponding coprocessor cores through the coprocessor task scheduling module, realizing the decoupled control of the main processor core and the coprocessor cores, simplifying the instruction distribution path and reducing the communication complexity of multi-core collaboration. At the same time, the interface compatibility for expanding more coprocessor cores is retained, enhancing the system scalability and the flexibility of modular design.

[0123] In some specific embodiments, the L1 cache includes an instruction cache and a data cache.

[0124] In some specific embodiments, the L2 cache is connected to the L1 cache, the matrix LSU storage unit, and the vector LSU storage unit through the AXI4 bus.

[0125] Through the AXI4 bus protocol, the efficient interconnection between the L2 cache, the main processor L1 cache, and the coprocessor storage unit is realized. The burst transmission and multi-channel parallel characteristics of the standardized bus are utilized to meet the large-scale data block transmission requirements of matrix and vector calculations, ensuring the storage access consistency between the main processor and the coprocessor. At the same time, it supports dynamically adjusting the transmission parameters according to the coprocessor characteristics, maximizing the bandwidth utilization rate and data supply stability of the heterogeneous storage subsystem.

[0126] In some specific embodiments, the matrix calculation unit group and the vector calculation unit group adopt a modular array structure. Each matrix calculation unit group or vector calculation unit group contains multiple parallel calculation units, and the data interconnection between the calculation unit groups is realized through the data distribution unit and the process monitoring and distribution unit.

[0127] Through the modular array structure and crossbar interconnection design, the matrix and vector calculation unit groups are constructed into a dynamically reconfigurable parallel computing resource pool, supporting the flexible allocation of the data flow path between the calculation units according to the task requirements, realizing multi-task parallel execution and resource load balancing, breaking through the calculation granularity limitation of the traditional fixed pipeline, and improving the hardware resource utilization rate and task parallelism in intensive computing scenarios.

[0128] In some specific embodiments, the matrix calculation unit group includes 32 matrix calculation units, and the vector calculation unit group includes 32 vector calculation units.

[0129] By configuring large-scale parallel calculation units for the matrix and vector calculation unit groups, a high-density computing resource cluster is formed, supporting the task-level fine-grained decomposition of complex matrix block operations and long vector element-level operations, accelerating the execution of computing-intensive tasks by using the hardware-level parallelism. At the same time, the unified calculation unit scale simplifies the task scheduling logic, reduces the multi-core collaboration management overhead, and improves the overall energy efficiency ratio of the system.

[0130] In some specific embodiments, the matrix calculation unit includes a matrix processing register file module, a matrix to-be-processed task queue module, a matrix to-be-processed data queue module, a matrix calculation data splicing module, a calculation result matrix splitting module, a matrix calculation scheduling arbitration module, and 4 matrix calculation iterative execution modules;

[0131] The vector calculation unit includes a vector processing register file module, a vector to-be-processed task queue module, a vector to-be-processed data queue module, a vector calculation scheduling arbitration module, and 4 vector calculation iterative execution modules.

[0132] Through the pipelined design of the task queue module, the scheduling arbitration module, and the iterative execution module, pre-storage of instructions and data, dynamic scheduling, and seamless multi-cycle execution are realized inside the calculation unit, reducing the idle time of the calculation unit caused by resource competition or data waiting. At the same time, through the collaborative management of the dedicated register file and the data queue, the stability of the operand supply and the calculation process is ensured, and the task execution efficiency and reliability of the calculation unit are improved.

[0133] In some specific embodiments, the matrix processing register file module is constructed by SRAM and is used to store the operand data necessary for matrix calculation. The SRAM contains 16 storage partitions, and every 4 partitions store 4 types of operand categories of a group of matrix calculation iterative execution modules respectively;

[0134] The matrix calculation iterative execution module is used to execute matrix calculation and includes a matrix ALU, a matrix FPU, and a matrix MUL;

[0135] The matrix to-be-processed task queue module and the matrix to-be-processed data queue module are used to sequentially cache the task information and to-be-processed data involved in the matrix calculation task according to the instruction execution order;

[0136] The matrix calculation scheduling arbitration module is used to verify whether the group of matrix calculation unit IDs assigned by the matrix co-processor process monitoring and allocation unit is idle, and sequentially reads the to-be-processed instructions from the instruction data queue and the to-be-processed data from the to-be-processed data queue, splices them, and sends them to the target matrix calculation unit;

[0137] The matrix calculation data splicing module is used to read the original operand data from the to-be-processed data queue according to the decoded calculation instruction, and splice the to-be-processed data into the matrix calculation format according to the instruction requirements;

[0138] The calculation result matrix splitting module is used to perform format splitting and conversion on the calculation results received from the 4 matrix calculation iterative execution modules, and split the matrix-format calculation results into the sequential association format that can be directly stored in the SRAM.

[0139] Through the hardware-level data format conversion design of the data splicing module and the calculation result splitting module in matrix calculation, the automatic reorganization of the original operands and the storage optimization of the calculation results are realized, the overhead of the software preprocessing link is eliminated, and at the same time, combined with the parallel operation of the multi-iteration execution module, the sub-block decomposition and pipeline execution of the matrix calculation task are supported, ensuring the efficient coordination of the high-throughput matrix operation and the storage system, and reducing the end-to-end processing delay of complex matrix algorithms.

[0140] In some specific embodiments, the vector processing register file module is constructed by SRAM and is used to store the operand data necessary for vector calculation. The SRAM contains 16 storage partitions, and each group of 4 partitions stores 4 types of operand categories of a vector calculation iteration execution module.

[0141] The vector calculation iteration execution module is used to execute vector calculations and includes a vector ALU, a vector FPU, and a vector MUL.

[0142] The vector quasi-processing task queue module and the vector quasi-processing data queue module are used to cache the task information and quasi-processing data involved in the vector calculation task in sequence according to the instruction execution order.

[0143] The vector calculation scheduling arbitration module is used to verify whether the group of vector calculation unit IDs assigned by the vector coprocessor process monitoring and allocation unit is idle, and sequentially reads the quasi-processing instructions from the instruction data queue and the quasi-processing data from the quasi-processing data queue, and then sends them to the target vector calculation unit ID.

[0144] Through the centralized scheduling and distributed execution architecture of the vector calculation unit, combined with the partition management of the register file and the load-aware arbitration strategy, the fast dispatch of vector instructions and the balanced load of the multi-iteration execution module are realized, supporting the parallel processing of vector operations with mixed bit widths and operation types. At the same time, the data alignment and storage access processes are simplified, and the execution efficiency of regular vector operations and the hardware resource reuse rate are improved.

[0145] Figure 1 is a schematic diagram of the architecture of a matrix-vector processor in an embodiment of the present application, as Figure 1 shown, the matrix-vector processor includes:

[0146] The main processor core is used to execute general computing instructions and complete preliminary decoding and scheduling for matrix computing instructions and vector computing instructions. The main processor core adopts a five-stage pipeline structure, including instruction fetching, decoding, issuing, execution, result collection, and write-back stages. The specific architecture includes an L1 cache, an instruction fetching module, a warp allocation module, a decoding module, an issue cache module, a scoreboard module, a main processor register file module, a task issuing module, a main processor computing unit, a coprocessor task scheduling module, and a computing result write-back module. The L1 cache includes an instruction cache and a data cache. The decoding module is configured to:

[0147] Parse the opcode, operands, and addressing mode of the instruction;

[0148] Match the instruction type through the opcode field, distinguish general computing instructions, vector computing instructions, or matrix computing instructions, and add extended instruction set tags to vector computing instructions or matrix computing instructions;

[0149] The matrix coprocessor core is connected to the main processor core and is used to execute matrix computing instructions. It adopts a three-stage pipeline structure, including decoding-secondary scheduling, execution, result collection, and write-back stages. The specific architecture includes a matrix secondary decoding module, a matrix secondary scheduling module, a matrix data allocation module, a matrix process monitoring and allocation module, an expandable number of matrix computing unit groups, and a matrix LSU storage unit;

[0150] The vector coprocessor core is connected to the main processor core and is used to execute vector computing instructions. It adopts a three-stage pipeline structure, including decoding-secondary scheduling, execution, result collection, and write-back stages. The specific architecture includes a vector secondary decoding module, a vector secondary scheduling module, a vector data allocation module, a vector process monitoring and allocation module, an expandable number of vector computing unit groups, and a vector LSU storage unit;

[0151] The L2 cache is connected to the L1 cache, the matrix LSU storage unit, and the vector LSU storage unit through the AXI4 bus;

[0152] Among them, the coprocessor task scheduling module is used to select the target coprocessor core according to the type of the extended instruction set tag, and is connected to the matrix secondary decoding module and the vector secondary decoding module respectively through the RoCC interface to form a one-to-two instruction distribution architecture;

[0153] A unidirectional direct inter-core channel is set between the matrix LSU storage unit and the vector LSU storage unit, and only allows the vector coprocessor core to transmit data to the matrix coprocessor core;

[0154] The scoreboard module is jointly maintained by the main processor core, the matrix coprocessor core, and the vector coprocessor core, and is configured such that the main processor core supports out-of-order instruction issue, and the matrix coprocessor core and the vector coprocessor core only support in-order instruction issue;

[0155] The matrix calculation unit group and the vector calculation unit group adopt a modular array structure. Each matrix calculation unit group or vector calculation unit group contains multiple parallel calculation units. Data interconnection is achieved between the calculation unit groups through a data distribution unit and a process monitoring and distribution unit. Each matrix calculation unit group includes 32 matrix calculation units, and each vector calculation unit group includes 32 vector calculation units;

[0156] Figure 2 This is the architecture schematic diagram of the matrix calculation unit in this embodiment. As Figure 2 shown, the matrix calculation unit includes a matrix processing register file module, a matrix to-be-processed task queue module, a matrix to-be-processed data queue module, a matrix calculation data splicing module, a calculation result matrix splitting module, a matrix calculation scheduling and arbitration module, and 4 matrix calculation iterative execution modules;

[0157] Among them, the matrix processing register file module is constructed by SRAM and is used to store the operand data necessary for matrix calculation. The SRAM contains 16 storage partitions, and every 4 partitions store one of the 4 types of operand categories of a group of matrix calculation iterative execution modules respectively;

[0158] The matrix calculation iterative execution module is used to execute matrix calculation and includes a matrix ALU, a matrix FPU, and a matrix MUL;

[0159] The matrix to-be-processed task queue module and the matrix to-be-processed data queue module are used to cache the task information and to-be-processed data involved in the matrix calculation tasks in sequence according to the instruction execution order;

[0160] The matrix calculation scheduling and arbitration module is used to verify whether the ID of this group of matrix calculation units assigned by the matrix co-processor process monitoring and distribution unit is idle, and sequentially reads the to-be-processed instructions from the instruction data queue, reads the to-be-processed data from the to-be-processed data queue, splices them and sends them to the target matrix calculation unit;

[0161] The matrix calculation data splicing module is used to read the original operation data from the to-be-processed data queue according to the decoded calculation instructions, and splice the to-be-processed data into the matrix calculation format according to the instruction requirements;

[0162] The calculation result matrix splitting module is used to perform format splitting and conversion on the calculation results received from the 4 matrix calculation iterative execution modules, and split the matrix-format calculation results into the sequential association format that can be directly stored in the SRAM;

[0163] Figure 3 This is the architecture schematic diagram of the vector calculation unit in this embodiment. As Figure 3As shown in the figure, the vector calculation unit includes a vector processing register file module, a vector quasi-processing task queue module, a vector quasi-processing data queue module, a vector calculation scheduling arbitration module, and four vector calculation iterative execution modules;

[0164] Among them, the vector processing register file module is built by SRAM and is used to store the operand data necessary for vector calculation. The SRAM contains 16 storage partitions, and every 4 partitions store one of the 4 types of operand categories of a group of vector calculation iterative execution modules;

[0165] The vector calculation iterative execution module is used to execute vector calculations and includes a vector ALU, a vector FPU, and a vector MUL;

[0166] The vector quasi-processing task queue module and the vector quasi-processing data queue module are used to sequentially cache the task information and quasi-processing data involved in the vector calculation tasks according to the instruction execution order;

[0167] The vector calculation scheduling arbitration module is used to verify whether the ID of this group of vector calculation units allocated by the vector co-processor process monitoring and allocation unit is idle, and sequentially read the quasi-processing instructions from the instruction data queue and the quasi-processing data from the quasi-processing data queue, splice them, and send them to the target vector calculation unit.

[0168] The following is an embodiment of the matrix-vector collaborative calculation method provided by the present disclosure. This matrix-vector collaborative calculation method and the matrix-vector processor in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the matrix-vector collaborative calculation method, reference can be made to the embodiments of the above matrix-vector processor.

[0169] Figure 4 is a flowchart of the matrix-vector processing method in an embodiment of the present application. The above matrix-vector processor is used for matrix-vector collaborative calculation, as Figure 4 shown, the steps include:

[0170] S1. The L1 cache obtains instructions and data from the L2 cache and preliminarily distinguishes the instruction types through the decoding module;

[0171] S2. If the instruction is a conventional calculation instruction, it is executed by the main processor calculation unit to obtain a conventional calculation result;

[0172] If the instruction is a matrix calculation instruction, it is distributed by the co-processor task scheduling module to the matrix co-processor core. The matrix co-processor core performs secondary decoding on the received instruction and allocates it to the matrix calculation unit group for execution to obtain a matrix calculation result;

[0173] If the instruction is a vector calculation instruction, it is distributed by the coprocessor task scheduling module to the vector coprocessor core. The vector coprocessor core performs secondary decoding on the received instruction, allocates it to the vector calculation unit group for execution, and obtains the vector calculation result;

[0174] S3. Write the calculation result back to the L2 cache and update the resource status of the scoreboard module.

[0175] It should be further noted that in step S1, the preliminary distinction of instruction types includes: parsing the opcode field of the instruction. If it matches the vector extension instruction set format, it is marked as a vector calculation instruction; if it matches the matrix extension instruction set format, it is marked as a matrix calculation instruction; if it does not match both the vector extension instruction set format and the matrix extension instruction set format at the same time, it is determined as a regular calculation instruction.

[0176] Through the collaborative mechanism of instruction classification execution by the main processor core and secondary decoding by the coprocessor core, dynamic task allocation and pipelined execution of heterogeneous computing resources are achieved. Combining the unified write-back of the calculation result and the update of the global status of the scoreboard ensures data consistency and resource visibility during the multi-core calculation process. At the same time, the data interaction path for matrix and vector hybrid calculations is shortened through the inter-core direct connection channel, improving the overall execution efficiency of complex calculation workflows.

[0177] In some specific embodiments, in step S1, the preliminary distinction of instruction types includes: parsing the opcode field of the instruction. If it matches the vector extension instruction set format, it is marked as a vector calculation instruction; if it matches the matrix extension instruction set format, it is marked as a matrix calculation instruction; if it does not match both the vector extension instruction set format and the matrix extension instruction set format at the same time, it is determined as a regular calculation instruction.

[0178] Through the fast matching and marking mechanism of the extended instruction set, accurate identification of instruction types is completed during the decoding stage of the main processor core, avoiding the coprocessor core from participating in the parsing of irrelevant instructions, reducing redundant decoding operations and power consumption overhead, and at the same time providing a low-latency instruction classification signal for subsequent task scheduling, ensuring the efficient connection of the instruction streams of the main processor and the coprocessor and the optimization of system energy efficiency.

[0179] In some specific embodiments, during the execution of calculations by the vector coprocessor core, if matrix calculations are required, the intermediate data is directly transmitted through the inter-core direct connection channel to the LSU storage unit of the matrix coprocessor core for further matrix calculations.

[0180] Through the design of a one-way inter-core direct connection channel from the vector to the matrix coprocessor, hardware-level zero-copy transmission of intermediate calculation results is achieved, bypassing the data write-back and reloading links of the traditional storage hierarchy, shortening the data transmission path for matrix-vector hybrid calculation tasks, reducing the end-to-end processing delay, and at the same time, through the matching design of the channel bandwidth and the throughput capacity of the calculation unit, avoiding data transmission from becoming a system performance bottleneck.

[0181] In some specific embodiments, in the matrix calculation unit group, the steps for the matrix calculation unit to perform calculations include:

[0182] S201. Receive the data to be processed from the matrix data queue module to be processed, and store it in the SRAM storage partition of the matrix processing register file module;

[0183] S202. Receive the task and instruction to be processed through the matrix task queue module to be processed;

[0184] S203. The matrix calculation data splicing module reads data from the matrix data queue module to be processed and splices it;

[0185] S204. The matrix calculation scheduling arbitration module verifies the idle state of the current matrix calculation unit ID;

[0186] S205. Distribute the spliced instruction data to the target matrix calculation iterative execution module;

[0187] S206. The calculation result matrix splitting module converts the calculation result and writes it back to the matrix processing register file module.

[0188] Through the multi-level pipelined task management mechanism of the matrix calculation unit, the full-process hardware automation from data pre-storage, format conversion to scheduling execution is realized. Combining the dynamic allocation of computing resources and the parallel processing of subtasks, it ensures the efficient execution of matrix operations and real-time result feedback. At the same time, through the optimization of the storage format, the result write-back overhead is reduced, and the continuous processing ability and system stability of large-scale matrix calculations are improved.

[0189] In some specific embodiments, in the vector calculation unit group, the steps for the vector calculation unit to perform calculations include:

[0190] S211. Receive the data to be processed from the vector data queue module to be processed, and store it in the SRAM storage partition of the vector processing register file module;

[0191] S212. Receive the task and instruction to be processed through the vector task queue module to be processed;

[0192] S213. The vector calculation scheduling arbitration module verifies the idle state of the current vector calculation unit ID;

[0193] S214. Distribute the instruction data to the target vector calculation iterative execution module;

[0194] S215. Write the calculation result back to the corresponding SRAM storage partition of the vector processing register file module.

[0195] Through the pre-verification mechanism of the vector calculation unit and the centralized task scheduling strategy, ensure the compliance of input data and the load balancing of computing resources. Combine the independent operation of the multi-iteration execution module and the partitioned write-back management to support the fast response of high-concurrency vector operations and the efficient storage of results, reduce the micro-management dependence of the software layer on hardware resources, and improve the processing efficiency and system maintainability of regular vector batch processing tasks.

[0196] In a specific embodiment, the steps of the matrix-vector processing method include:

[0197] S1. The instruction cache and data cache of the L1 cache respectively receive and cache the instructions to be executed in the current calculation and the data to be processed from the corresponding storage partitions of the L2 cache.

[0198] The instruction fetch module fetches the address of the instruction to be executed currently, reads the corresponding instruction from the storage unit, and outputs it to the warp allocation module. The warp allocation module groups the instructions for subsequent calculation in units of warps.

[0199] The decoding module receives the instructions from the instruction bus, parses them, and preliminarily distinguishes the instruction types. The decoding results are cached sequentially by the issue cache module.

[0200] Preliminarily distinguishing the instruction types includes: parsing the opcode field of the instruction. If it matches the vector extension instruction set format, it is marked as a vector calculation instruction; if it matches the matrix extension instruction set format, it is marked as a matrix calculation instruction; if it does not match both the vector extension instruction set format and the matrix extension instruction set format, it is determined as a regular calculation instruction.

[0201] The scoreboard module performs appropriate dynamic instruction scheduling on the decoding results to enable multiple instructions to be executed simultaneously as much as possible; then reads the operands involved in the instructions from the main processor register file module, and sends the instructions and the operands to be processed to the task issue module in sequence; the task issue module determines whether the preliminary recognition result of the decoding module is a matrix calculation instruction / vector calculation instruction.

[0202] S2. If the instruction is a regular calculation instruction, the task issue module sends the current task instruction and the relevant data to be processed to the main processor calculation unit. The main processor allocates the task to the corresponding sub-calculation unit according to the instruction category to perform the corresponding ALU / FPU / SFU calculation, obtains the regular calculation result, and stores the regular calculation result in the L1 cache.

[0203] If the instruction is a matrix calculation instruction, the co-processor task scheduling module sends the current instruction and the data to be processed to the matrix co-processor core through the RoCC interface. The secondary decoding module of the matrix co-processor core performs secondary decoding on the received instruction. The matrix secondary scheduling module reads the real-time calculation status of the process monitoring and allocation unit, allocates the received instruction to the idle matrix calculation unit group, and the data allocation unit sends the instruction and data to be processed to the specific matrix calculation unit included in the matrix calculation unit group to perform the corresponding calculation to obtain the matrix calculation result;

[0204] After the calculation is completed, the calculation result is returned to the matrix LUS storage unit in sequence in units of the matrix calculation unit group. If the main processor core needs the calculation result, the matrix LUS storage unit sends the calculation result to the main processor core through the AXI bus, otherwise the calculation result is directly sent back to the L2 cache;

[0205] Figure 5 It is the flowchart of the matrix calculation unit performing calculations in this embodiment. As Figure 5 shown, in the matrix calculation unit group, the steps for the matrix calculation unit to perform calculations include:

[0206] S201. Receive the data to be processed from the matrix data queue module to be processed, and store it in the SRAM storage partition of the matrix processing register file module;

[0207] S202. Receive the task and instruction to be processed through the matrix task queue module to be processed;

[0208] S203. The matrix calculation data splicing module reads and splices the data from the matrix data queue module to be processed;

[0209] S204. The matrix calculation scheduling arbitration module verifies the idle status of the current matrix calculation unit ID;

[0210] S205. Distribute the spliced instruction data to the target matrix calculation iterative execution module;

[0211] S206. The calculation result matrix splitting module converts the calculation result and writes it back to the matrix processing register file module;

[0212] If the instruction is a vector calculation instruction, the co-processor task scheduling module sends the current instruction and the data to be processed to the vector co-processor core through the RoCC interface. The vector co-processor core performs secondary decoding on the received instruction. The vector secondary scheduling module reads the real-time calculation status of the process monitoring and allocation unit, allocates the received instruction to the idle vector calculation unit group, and the data allocation unit sends the instruction and data to be processed to the specific vector calculation unit included in the vector calculation unit group to perform the corresponding calculation to obtain the vector calculation result;

[0213] Among them, during the execution of calculations by the vector co-processor core, if matrix calculations need to be invoked, the intermediate data is directly transmitted to the LSU storage unit of the matrix co-processor core through the direct inter-core connection channel for further matrix calculations;

[0214] After the calculation is completed, the calculation results are returned to the vector LUS storage unit in sequence in units of the vector calculation unit group. If the main processor core needs the calculation results, the vector LUS storage unit sends the calculation results to the main processor core via the AXI bus; otherwise, the calculation results are directly sent back to the L2 cache.

[0215] Figure 6 is the flowchart of the calculation executed by the matrix-vector calculation unit in this embodiment. As Figure 6 shown, in the vector calculation unit group, the steps for the vector calculation unit to execute calculations include:

[0216] S211. Receive the data to be processed from the vector data to be processed queue module and store it in the SRAM storage partition of the vector processing register file module;

[0217] S212. Receive the tasks and instructions to be processed through the vector task to be processed queue module;

[0218] S213. The vector calculation scheduling arbitration module verifies the idle state of the current vector calculation unit ID;

[0219] S214. Distribute the instruction data to the target vector calculation iterative execution module;

[0220] S215. Write the calculation results back to the corresponding SRAM storage partition of the vector processing register file module;

[0221] S3. Write the calculation results back to the L2 cache and update the resource status of the scoreboard module.

[0222] This application also provides an electronic device for implementing each embodiment of this application. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is a matrix-vector processor.

[0223] Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of this application does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0224] In embodiments of the present application, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.

[0225] For software implementation, the implementation of a process or function may be implemented with separate software modules that allow the execution of at least one function or operation. The software code may be implemented by a software application (or program) written in any suitable programming language, and the software code may be stored in a memory and executed by a controller.

[0226] In addition, the electronic device includes some functional modules not shown herein, which will not be elaborated herein.

[0227] Those skilled in the art can understand that various aspects of the electronic device provided by the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" herein.

[0228] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A matrix-vector processor, characterized in that: include: The main processor core is used to execute conventional computing instructions and complete preliminary decoding and scheduling for matrix computing instructions and vector computing instructions, including L1 cache, instruction capture module, thread bundle allocation module, decoding module, emission cache module, scoreboard module, main processor register file module, task emission module, main processor computing unit, coprocessor task scheduling module, and calculation result write-back module; A matrix coprocessor core, connected to the main processor core, is used to execute matrix calculation instructions, including a matrix secondary decoding module, an expandable number of matrix calculation unit groups and a matrix LSU storage unit; A vector coprocessor core, connected to the main processor core, for executing vector computing instructions, including a vector secondary decoding module, an expandable number of vector computing unit groups and a vector LSU storage unit; The L2 cache is connected to the L1 cache, the matrix LSU storage unit and the vector LSU storage unit respectively; Among them, a unidirectional inter-core direct connection channel is set between the matrix LSU storage unit and the vector LSU storage unit, which only allows the vector coprocessor core to transmit data to the matrix coprocessor core; The scoreboard module is jointly maintained by the main processor core, the matrix coprocessor core and the vector coprocessor core, and is configured so that the main processor core supports out-of-order instruction issuance, while the matrix coprocessor core and the vector coprocessor core only support sequential instruction issuance.

2. The matrix-vector processor according to claim 1, wherein: The decoding module is configured as: Parse the instruction's opcode, operand, and addressing mode; The instruction type is matched through the opcode field to distinguish between conventional computing instructions, vector computing instructions or matrix computing instructions, and an extended instruction set tag is added to the vector computing instructions or matrix computing instructions.

3. The matrix-vector processor according to claim 1, wherein: The coprocessor task scheduling module is used to select the target coprocessor core according to the type of extended instruction set tag, and connect the matrix secondary decoding module and the vector secondary decoding module through the RoCC interface to form a one-to-two instruction distribution architecture.

4. The matrix-vector processor of claim 1, wherein: The L2 cache is connected to the L1 cache, the matrix LSU storage unit and the vector LSU storage unit through the AXI4 bus.

5. The matrix-vector processor of claim 1, wherein: The matrix computing unit group and the vector computing unit group adopt a modular array structure. Each matrix computing unit group or vector computing unit group includes multiple parallel computing units. Data interconnection is achieved between the computing unit groups through a data distribution unit and a process monitoring distribution unit.

6. The matrix-vector processor of claim 5, wherein: The matrix calculation unit group includes 32 matrix calculation units, and the vector calculation unit group includes 32 vector calculation units.

7. The matrix-vector processor of claim 5, wherein: The matrix calculation unit includes a matrix processing register file module, a matrix processing task queue module, a matrix processing data queue module, a matrix calculation data splicing module, a calculation result matrix splitting module, a matrix calculation scheduling arbitration module and four matrix calculation iterative execution modules; The vector computing unit includes a vector processing register file module, a vector task queue module, a vector data queue module, a vector computing scheduling arbitration module and four vector computing iteration execution modules.

8. A matrix-vector collaborative computing method, characterized in that: Using the matrix-vector processor as described in any one of claims 1 to 7 to perform matrix-vector collaborative computing, comprising: S1. L1 cache obtains instructions and data from L2 cache and initially distinguishes instruction types through the decoding module; S2. If the instruction is a conventional computing instruction, the main processor computing unit performs the calculation to obtain a conventional calculation result; If the instruction is a matrix calculation instruction, it is distributed to the matrix coprocessor core by the coprocessor task scheduling module. The matrix coprocessor core performs secondary decoding on the received instruction and distributes it to the matrix calculation unit group to perform calculation and obtain the matrix calculation result. If the instruction is a vector calculation instruction, the coprocessor task scheduling module distributes it to the vector coprocessor core, and the vector coprocessor core performs secondary decoding on the received instruction and distributes it to the vector calculation unit group to perform calculation and obtain the vector calculation result; S3. Write the calculation result back to the L2 cache and update the resource status of the scoreboard module.

9. The matrix-vector collaborative computing method according to claim 8, characterized in that: In step S1, the initial distinction of instruction types includes: parsing the opcode field of the instruction, marking it as a vector calculation instruction if it matches the vector extension instruction set format; marking it as a matrix calculation instruction if it matches the matrix extension instruction set format; and determining it as a regular calculation instruction if it does not match both the vector extension instruction set format and the matrix extension instruction set format.

10. The matrix-vector collaborative computing method according to claim 8, characterized in that: During the calculation process of the vector coprocessor core, if matrix calculation needs to be called, the intermediate data is directly transferred to the LSU storage unit of the matrix coprocessor core through the inter-core direct connection channel for further matrix calculation.

Citation Information

Patent Citations

  • Data processing system for collaborative operation of vector DSP and coprocessors

    CN103793208A

  • Cooperative computing method, system and equipment based on SIMD and SIMT

    CN119536820A