Data processing task execution method and device, equipment and medium

By setting the target hardware accelerator outside the processor and using two-stage decoding and instruction scheduler, the problem of hardware accelerator's computing power bottleneck and poor scalability in data processing tasks is solved, and efficient data processing task execution is achieved.

CN120564009APending Publication Date: 2025-08-29SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510694059.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional CPUs cannot meet the growing demand for data task processing computing power, and hardware accelerators have bottlenecks and poor scalability when performing data processing tasks.

Method used

The target hardware accelerator independently set outside the processor is used to decode the data processing tasks issued by the processor first and again through two-stage decoding and instruction scheduler, and the target processing instructions of vector type and matrix type are filtered out, and each operation executor is scheduled to perform operations according to the execution order.

Benefits of technology

It improves the scalability of the hardware accelerator, avoids computing power bottlenecks, improves data processing speed and overall computing power, and achieves efficient data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564009A_ABST
    Figure CN120564009A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing task execution method and device, equipment and a medium. The method and device are applied to a target hardware accelerator independently arranged outside a processor. The method comprises the following steps: performing primary decoding on each processing instruction in a current data processing task issued by the processor to screen out a target processing instruction of a target type; wherein the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type comprises a vector type and a matrix type; determining an execution sequence of the re-decoded target processing instructions; and according to the execution sequence, scheduling a target executor in each operation executor to execute an operation corresponding to each re-decoded target processing instruction so as to complete the current data processing task. By means of the scheme, the problems that the computing power of the hardware accelerator is bottleneck when the hardware accelerator executes the data processing task and the expandability of the hardware accelerator is poor can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and medium for executing data processing tasks. Background Art

[0002] In recent years, driven by emerging application scenarios such as big data, 5G communications, and large language models, artificial intelligence technology has penetrated into all aspects of human life. At the same time, with the rapid rise of generative artificial intelligence (AIGC), companies are increasingly adopting AIGC, natural language processing, and neural networks to expand functionality and enhance user experience. The current largest neural network model is approximately 150,000 to 300,000 times larger than its predecessor. Traditional CPUs (Central Processing Units) can no longer meet the computing power requirements of the growing data task processing.

[0003] To break through the computing bottleneck of traditional computing systems, researchers have begun focusing on new hardware architectures. To balance application cost and flexibility, many mainstream processors have implemented the Single Instruction Multiple Data (SIMD) instruction set extension. SIMD instructions can perform the same operation on multiple sets of data simultaneously, saving dynamic instruction bandwidth and program space.

[0004] The conventional vector instruction set extension structure based on the RISC-V architecture uses a VPU (Video Processing Unit) solution in which a hardware accelerator is built into the CPU pipeline. This solution faces bottlenecks in the computing power requirements of the growing data processing tasks of artificial intelligence and has poor scalability.

[0005] It can be seen that how to avoid the bottleneck problem of computing power when the hardware accelerator performs data processing tasks and the poor scalability of the hardware accelerator are problems that technical personnel in this field need to solve. Summary of the Invention

[0006] The purpose of the embodiments of the present invention is to provide a method, apparatus, device, and medium for executing data processing tasks to avoid bottlenecks in computing power and poor scalability of hardware accelerators when executing data processing tasks. The specific solution is as follows:

[0007] In a first aspect, the present invention discloses a method for executing a data processing task, which is applied to a target hardware accelerator independently provided outside a processor; the method comprises:

[0008] Performing initial decoding on each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of a target type; wherein the current data processing task is a task in an image recognition model built based on a neural network, and the target type includes a vector type and a matrix type;

[0009] determining an execution order of the re-decoded target processing instructions;

[0010] The target executors in each operation executor are scheduled to execute operations corresponding to each re-decoded target processing instruction according to the execution order to complete the current data processing task.

[0011] Optionally, the performing initial decoding on each processing instruction in the current data processing task issued by the processor includes:

[0012] Acquiring a current data processing task issued by the processor through a preset standardized interface between the processor and the target hardware accelerator;

[0013] Performing initial decoding on the current data processing task.

[0014] Optionally, the target hardware accelerator includes an instruction decoder, and the instruction decoder includes a first-level instruction decoder and a second instruction decoder;

[0015] The initial decoding of each processing instruction in the current data processing task issued by the processor includes:

[0016] Using the first-level instruction decoder to initially decode each processing instruction in the current data processing task issued by the processor;

[0017] Accordingly, determining the execution order of the target processing instructions after re-decoding includes:

[0018] The second instruction decoder is used to decode each target processing instruction again to obtain a micro-operation corresponding to each target processing instruction and determine an execution order of each micro-operation.

[0019] Optionally, the target hardware accelerator includes an instruction scheduler; and determining the execution order of the re-decoded target processing instructions includes:

[0020] Re-decoding each of the target processing instructions to obtain a micro-operation and a required operand corresponding to each of the target processing instructions, and obtaining the required operand from the processor;

[0021] Utilizing the instruction scheduler to analyze dependencies and conflicts between the micro-operations, and determining an execution order of the micro-operations based on the dependencies, conflicts, and computing resources of the operation executors;

[0022] Accordingly, scheduling the target executors in each operation executor to execute operations corresponding to each re-decoded target processing instruction according to the execution order includes:

[0023] The target executors in each operation executor are scheduled according to the execution order so that the target executor executes each micro-operation using the required operands.

[0024] Optionally, each of the operation executors includes a vector matrix operator, a vector mask controller, a cross-channel instruction processor, and a vector memory access manager, and the target executor is any one or more executors among the operation executors.

[0025] Optionally, the vector-matrix operator includes a vector register, a matrix register, a first operator for vector arithmetic instruction operations, a second operator for floating-point and multiplication and division operations, a third operator for matrix arithmetic instruction operations, and a fourth operator for matrix multiplication and accumulation operations; wherein each channel of the vector register includes multiple single ports for parallel storage of data blocks; the matrix register includes multiple two-dimensional matrix registers, and the number of rows and columns of each of the two-dimensional matrix registers is determined based on the row length of the two-dimensional matrix register.

[0026] Optionally, scheduling the target executors in each operation executor to execute operations corresponding to each re-decoded target processing instruction according to the execution order includes:

[0027] The target executors in each operation executor are scheduled as the current executor in turn according to the execution order;

[0028] If the vector mask controller is the current executor, the vector mask controller performs a bit-level operation on the target element corresponding to the current instruction to complete the mask operation on the current instruction; the bit-level operation is any one or more of a bit-and operation, a bit-or operation, and a bit-exclusive-or operation;

[0029] If the cross-channel instruction processor is the current executor, the cross-channel instruction processor performs any one or more operations of a data fusion operation, a data extraction operation, a reduction calculation operation, a data rearrangement operation, and a data shift operation on the current instruction;

[0030] If the vector memory access manager is the current executor, the vector memory access manager obtains a memory access operation type of a current instruction, performs address calculation on the current instruction according to a memory access mode supported by the extensible vector instruction set, to obtain a memory access request including a memory access address, merges the memory access requests that meet a preset condition, and performs a memory access operation corresponding to the memory access operation type on data at the memory address in a target storage area through an advanced extensible interface to respond to the merged memory access request;

[0031] The current instruction is any one of the target processing instructions after being decoded again.

[0032] In a second aspect, the present invention discloses a data processing task execution device, which is applied to a target hardware accelerator independently arranged outside a processor; the device comprises:

[0033] a primary decoding module, configured to perform primary decoding on each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of a target type; wherein the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type;

[0034] An order determination module, configured to determine an execution order of the target processing instructions after being re-decoded;

[0035] The operation scheduling module is used to schedule the target executors in each operation executor to execute the operations corresponding to each re-decoded target processing instruction according to the execution order, so as to complete the current data processing task.

[0036] In a third aspect, the present invention discloses an electronic device, comprising:

[0037] Memory, used to store computer programs;

[0038] The processor is used to execute a computer program to implement the steps of the aforementioned disclosed method for executing data processing tasks.

[0039] In a fourth aspect, the present invention discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the steps of the aforementioned disclosed method for executing data processing tasks are implemented.

[0040] It can be seen that the present invention is applied to a target hardware accelerator independently arranged outside a processor; the method includes: performing initial decoding on each processing instruction in the current data processing task issued by the processor to screen out target processing instructions of the target type; wherein, the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type; determining the execution order of each re-decoded target processing instruction; scheduling the target executor in each operation executor to execute the operation corresponding to each re-decoded target processing instruction according to the execution order to complete the current data processing task.

[0041] The beneficial effects are as follows: the present invention abandons the conventional solution of integrating the hardware accelerator into the processor pipeline and is applied to a target hardware accelerator independently set outside the processor, that is, the hardware accelerator is placed outside the processor. This layout makes the hardware accelerator independent of the processor, reduces the dependence on the processor pipeline, improves the scalability of the hardware accelerator, avoids the computing power bottleneck caused by the internal resource limitation of the processor, and provides more space for improving computing power; two-level decoding is adopted for the processing instructions of the data processing task. The initial decoding simply distinguishes the processing instructions, filters out the target processing instructions of the vector type and the matrix type, and then decodes the target processing instructions again, that is, performs full decoding. The two-level decoding has a clear division of labor, improves the instruction processing efficiency, and speeds up the data processing speed; further, the execution order of the target processing instructions after re-decoding is determined, and each target executor is scheduled according to the execution order. That is, different target executors are used to reasonably execute different operations. Each executor performs its duties and works together to efficiently process various instructions and improve the overall computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0043] Figure 1 A structural diagram of a specific vector instruction set extension based on the RISC-V architecture provided in an embodiment of the present invention;

[0044] Figure 2 A flow chart of a data processing task execution method provided by an embodiment of the present invention;

[0045] Figure 3 A specific vector register schematic diagram provided by an embodiment of the present invention;

[0046] Figure 4A schematic diagram of a specific matrix register provided by an embodiment of the present invention;

[0047] Figure 5 A flowchart of a specific data processing task execution method provided by an embodiment of the present invention;

[0048] Figure 6 A schematic diagram of a specific target hardware accelerator structure provided by an embodiment of the present invention;

[0049] Figure 7 A schematic diagram of a specific hardware accelerator data flow provided by an embodiment of the present invention;

[0050] Figure 8 A schematic structural diagram of a data processing task execution device provided by an embodiment of the present invention;

[0051] Figure 9 A structural diagram of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0053] In recent years, driven by emerging application scenarios such as big data, 5G communications, and large language models, artificial intelligence technology has penetrated into all aspects of human life. At the same time, with the rapid rise of generative artificial intelligence, companies are increasingly adopting AIGC, natural language processing, and neural networks to expand functions and enhance user experience. The current largest neural network model is about 150,000 to 300,000 times larger than that at the time. Traditional CPUs can no longer meet the computing power requirements of the growing data task processing.

[0054] To break through the computing power bottleneck of traditional computing systems, researchers have begun focusing on new hardware architectures. To balance application cost and flexibility, many mainstream processors have implemented single-instruction, multiple-data (SIMD) instruction set extensions. SIMD instructions can perform the same operation on multiple sets of data simultaneously, saving dynamic instruction bandwidth and program space.

[0055] For example Figure 1 As shown in the figure, the conventional vector instruction set extension structure based on the RISC-V architecture uses a VPU solution in which the hardware accelerator is built into the CPU pipeline. This solution will encounter bottlenecks in the computing power requirements of the growing data processing tasks of artificial intelligence, and its scalability is poor.

[0056] The terms "including" and "having," as used in the present description and accompanying drawings, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.

[0057] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0058] Next, a data processing task execution solution provided by an embodiment of the present invention is introduced in detail. Figure 2 A data processing task execution method provided in an embodiment of the present invention is applied to a target hardware accelerator independently provided outside a processor; the method comprises:

[0059] Step S11: Initially decode each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of the target type; wherein, the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type.

[0060] In this embodiment, the initial decoding of each processing instruction in the current data processing task issued by the processor includes: obtaining the current data processing task issued by the processor through a preset standardized interface between the processor and the target hardware accelerator; and initial decoding of the current data processing task.

[0061] The target hardware accelerator is independently set outside the processor, and there is a preset standardized interface between the processor and the target hardware accelerator. The preset standardized interface is used to transmit the current data processing task issued by the processor to the target hardware accelerator. The current data processing task is a task in the image recognition model built based on the neural network, so that the target hardware accelerator performs the initial decoding of the current data processing task.

[0062] In this embodiment, the target hardware accelerator includes an instruction decoder, and the instruction decoder includes a first-level instruction decoder and a second instruction decoder; the initial decoding of each processing instruction in the current data processing task issued by the processor includes: using the first-level instruction decoder to initially decode each processing instruction in the current data processing task issued by the processor.

[0063] The target hardware accelerator includes an instruction decoder, and the instruction decoder includes a first-level instruction decoder and a second instruction decoder. The first-level instruction decoder is placed at the ID level (instruction decoding stage) of the front end of the processor. That is to say, in this embodiment, the current data processing task is decoded at two levels. During the initial decoding process, only target processing instructions of the target type need to be screened out, where the target type includes vector type and matrix type. The first-level decoder is placed at the ID level of the front end of the central processing unit and is only used to simply distinguish whether the current instruction is a vector instruction or a matrix instruction. Target processing instructions of non-target types are processed by the processor.

[0064] Step S12: Determine the execution order of the re-decoded target processing instructions.

[0065] In this embodiment, determining the execution order of each target processing instruction after re-decoding includes: using the second instruction decoder to re-decode each target processing instruction to obtain the micro-operation corresponding to each target processing instruction, and determining the execution order of each micro-operation.

[0066] The second instruction decoder needs to decode each target processing instruction again to determine the micro-operation corresponding to each target processing instruction. It can be understood that a data processing task may contain multiple processing instructions, and each processing instruction may correspond to multiple micro-operations. There may be dependencies, contradictions, etc. between different micro-operations, and the execution order needs to be determined based on the relationship between each micro-operation.

[0067] Step S13: scheduling the target executors in each operation executor to execute operations corresponding to each re-decoded target processing instruction according to the execution order, so as to complete the current data processing task.

[0068] In this embodiment, each of the operation executors includes a vector matrix operator, a vector mask controller, a cross-channel instruction processor, and a vector memory access manager, and the target executor is any one or more executors among the operation executors.

[0069] Each operation executor includes a vector matrix operator, a vector mask controller, a cross-channel instruction processor and a vector memory access manager. The target executor in each operation executor is scheduled according to the execution order to perform the operation corresponding to each target processing instruction after re-decoding. That is, according to the execution order, each target executor acts as the current executor in turn to execute the current micro-operation, thereby completing each target processing instruction and then completing the current data processing task. The target executor is any one or more executors in each operation executor. That is, all operation executors can be target executors, or only the vector matrix operator can be the target executor, etc., which is determined according to the micro-operation to be executed. The current target executor can be one or more. The current target executor executes the current corresponding micro-operation and obtains the result of the micro-operation to perform the next micro-operation.

[0070] Among them, the vector mask controller is used to process vector mask instructions, which is very useful when bit-level operations on elements in a vector are required. Since they directly operate on mask registers, these operations are often used to control the flow of vector calculations, for example, in conditional execution or selective updating of vector elements.

[0071] The cross-lane instruction processor is responsible for processing cross-lane instructions, such as: inserting a scalar operand into a vector operand; grabbing a scalar operand from a vector operand; the Reduction instruction receives an element of a vector register group and a scalar placed at element 0 of the vector register, and obtains a scalar value through some reduction operation, which is also placed at element 0 of a vector register; the Shuffle instruction can rearrange the elements in the vector according to a specified pattern; the Slide instruction moves data elements up and down in the vector register group, which is very useful when processing sliding windows, overlapping windows or other scenarios that require relative displacement between elements.

[0072] The Vector Memory Manager is responsible for processing vector memory access instructions. It has a single external storage interface: the Advanced eXtensible Interface (AXI) with a bit width of 2 bytes per DP-FLOP. This choice was made to ensure a balance between computing power and bandwidth. The Vector Memory Manager includes an Address Generator (AGU) to support the memory access modes in the RVV instruction set, including unit-stride, strided, and indexed addressing. The Vector Memory Manager consolidates AGU memory access requests into bursts, accessing external storage through the AXI interface.

[0073] In this embodiment, the vector-matrix operator includes a vector register, a matrix register, a first operator for vector arithmetic instruction operations, a second operator for floating-point and multiplication and division operations, a third operator for matrix arithmetic instruction operations, and a fourth operator for matrix multiplication and accumulation operations; wherein each channel of the vector register includes multiple single ports for parallel storage of data blocks; the matrix register includes multiple two-dimensional matrix registers, and the number of rows and columns of each of the two-dimensional matrix registers is determined based on the row length of the two-dimensional matrix register.

[0074] Furthermore, the number of vector matrix operators can be expanded and configured to 2 to 16. The vector matrix operators are divided into registers and operators. Each vector matrix operator has its own lane sequencer (operation channel instruction sequencer) responsible for tracking 8 parallel instructions. Registers include vector register files (VRF) and matrix register files (MRF). For example, Figure 3 As shown in the figure, a specific vector register is implemented using a set of single-port (1RW) storage banks (data blocks). The width of each data block is 64 bits, which is the same as the width of the data path of each channel. Each channel has eight single-port storage data blocks. In other words, each channel of the vector register contains multiple single ports for parallel storage of data blocks. The width of the data block is the same as the path width of the channel. For example Figure 4 A specific matrix register schematic diagram is shown, where the matrix register includes multiple two-dimensional matrix registers, specifically 8 two-dimensional matrix registers, namely M0, M1, M2, M3, M4, M5, M6 and M7. RLEN represents the row length of each register in bit units, which is a constant value in any implementation. The available RLEN options are 128, 256 or 512 or more. The number of rows of the matrix register is RLEN divided by 32, and the result is 4, 8 or 16 or more. Therefore, each matrix register consists of [M×K]MSEW elements, where M is calculated as RLEN divided by 32, K is calculated as RLEN divided by MSEW, and MSEW represents the element size, that is, the number of rows and columns of each two-dimensional matrix register is determined based on the row length of the two-dimensional matrix register.

[0075] The first operator (Vector Arithmetic Logic Unit, VALU) is used for vector arithmetic instruction operations, the second operator (Vector Multi-Floating-Point Unit, VMFPU) is used for floating-point and multiplication and division operations, the third operator (Matrix Arithmetic Logic Unit, MALU) is used for matrix arithmetic instruction operations, and the fourth operator (Multiply-AccumulateUnit, MAC) is used for matrix multiplication and accumulation operations.

[0076] In this embodiment, the target executors in each operation executor are scheduled according to the execution order to execute the operations corresponding to the target processing instructions after each re-decoding, including: scheduling the target executors in each operation executor as the current executor in turn according to the execution order; if the vector mask controller is the current executor, the vector mask controller performs a bit-level operation on the target element corresponding to the current instruction to complete the mask operation on the current instruction; the bit-level operation is any one or more of a bit-and operation, a bit-or operation, and a bit-exclusive-or operation; if the cross-channel instruction processor is the current executor, the cross-channel instruction processor performs a data fusion operation, a data extraction operation, and a bit-exclusive-or operation on the current instruction. any one or more operations among fetch operations, reduction calculation operations, data rearrangement operations and data displacement operations; if the vector memory access manager is the current executor, the vector memory access manager obtains the memory access operation type of the current instruction, and performs address calculation on the current instruction according to the memory access mode supported by the extensible vector instruction set to obtain a memory access request including a memory access address, merges the memory access requests that meet the preset conditions, and performs a memory access operation corresponding to the memory access operation type on the data at the memory address in the target storage area through the high-level extensible interface to respond to the merged memory access request; wherein, the current instruction is any one of the target processing instructions after being decoded again.

[0077] When the system is scheduled according to the preset execution order, the target executor in each operation executor will be designated as the current executor in turn:

[0078] If the current executor is a vector mask controller, the controller will perform bit-level logical operations on the target element corresponding to the current instruction. The supported bit-level operations include basic logical operations such as bitwise AND, bitwise OR, and bitwise exclusive OR (XOR). These operations complete the mask control function of the current instruction.

[0079] If the current executor is a cross-channel instruction processor, the cross-channel instruction processor will perform cross-channel data processing operations on the current instruction. Supported operation types include: data fusion (such as inserting scalars into vectors), data extraction (such as obtaining scalars from vectors), reduction calculations (such as vector sum reduction), data rearrangement (such as element reordering), and data displacement (such as sliding window operations).

[0080] If the current executor is a vector memory manager, it first identifies the memory access operation type (load / store) of the current instruction, performs address calculation according to the memory access mode supported by the extensible vector instruction set, and generates a memory access request. The supported memory access modes include: unit stride, stride, index, etc. It merges and optimizes multiple memory access requests that meet the conditions, executes the actual memory access operation through the advanced extensible interface (i.e., AXI interface), reads or writes the data at the specified address in the target storage area, and finally completes the response to the merged memory access request.

[0081] It should be noted that the current instruction refers to any target processing instruction after secondary decoding by the instruction decoder. The entire scheduling process is coordinated by the instruction scheduler to ensure that each executor executes the instructions correctly in sequence.

[0082] In this embodiment, a power regulator can also be set in the hardware accelerator to determine the current target executor and the non-current target executor, and the power regulator is used to open the power path of the current target executor and close the power path of the non-current target executor, so as to power the current target executor instead of the non-current target executor. In this way, the power consumption of the target hardware accelerator when executing the current data processing task can be effectively reduced. For example, if the current target executor only has a vector memory manager, the power paths of the vector matrix operator, the vector mask controller, and the cross-channel instruction processor can all be turned off.

[0083] The beneficial effects are as follows: the present invention abandons the conventional solution of integrating the hardware accelerator into the processor pipeline and is applied to a target hardware accelerator independently set outside the processor, that is, the hardware accelerator is placed outside the processor. This layout makes the hardware accelerator independent of the processor, reduces the dependence on the processor pipeline, improves the scalability of the hardware accelerator, avoids the computing power bottleneck caused by the internal resource limitation of the processor, and provides more space for improving computing power; two-level decoding is adopted for the processing instructions of the data processing task. The initial decoding simply distinguishes the processing instructions, filters out the target processing instructions of the vector type and the matrix type, and then decodes the target processing instructions again, that is, performs full decoding. The two-level decoding has a clear division of labor, improves the instruction processing efficiency, and speeds up the data processing speed; further, the execution order of the target processing instructions after re-decoding is determined, and each target executor is scheduled according to the execution order. That is, different target executors are used to reasonably execute different operations. Each executor performs its duties and works together to efficiently process various instructions and improve the overall computing power.

[0084] See also Figure 5 This embodiment of the present invention discloses a specific method for executing a data processing task. Compared to the previous embodiment, this embodiment further illustrates and optimizes the technical solution. The method is applied to a target hardware accelerator independently disposed outside a processor; the target hardware accelerator includes an instruction scheduler; and the method includes:

[0085] Step S21: Initially decode each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of the target type; wherein, the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type.

[0086] An image recognition model built based on a neural network is used to process image recognition tasks, such as real-time recognition of facial images. The facial image captured by the camera is input into the image recognition model. The model recognizes the image and sends the data processing tasks in the recognition process from the processor to the target hardware accelerator. The data processing tasks include various processing instructions, among which vector type and matrix type processing instructions are target processing instructions, such as affine transformation matrix calculation instructions, depthwise separable convolution instructions, vector dot product instructions, etc. The target hardware accelerator filters out the target processing instructions.

[0087] Step S22: re-decode each of the target processing instructions to obtain the micro-operations and required operands corresponding to each of the target processing instructions, and obtain the required operands from the processor.

[0088] Each target processing instruction is decoded again to obtain the micro-operation and required operand corresponding to each target processing instruction, and the required operand is obtained from the processor. The required operand is specifically a scalar operand, a source operand that requires a scalar, or a destination operand that generates a scalar (i.e., an operand in the scalar x, f Register File), for example, reading the required operand from a register in the processor.

[0089] Step S23: Utilizing the instruction scheduler to analyze the dependency and conflict relationships among the micro-operations, and determining the execution order of the micro-operations based on the dependency, conflict, and computing resources of each operation executor.

[0090] The instruction scheduler is a global scheduler for the hardware accelerator. Its primary function is to manage and control the execution order, dependencies, and conflicts of instructions in parallel vector computations, ensuring coordinated interface operation with the various functional units in the hardware accelerator. By recording and analyzing instruction status and conflicts, it optimizes the allocation of computing resources and improves execution efficiency. Furthermore, the module implements a mechanism for waiting for and processing unfinished instructions to ensure correctness in complex data paths.

[0091] Step S24: scheduling the target executors in each operation executor according to the execution order, so that the target executor executes each micro-operation using the required operands to complete the current data processing task.

[0092] As can be seen, the hardware accelerator of the present invention utilizes an external, independent design, avoiding the performance bottlenecks of traditional integrated processor solutions. It implements efficient instruction scheduling through an instruction decoder and instruction scheduler, distributing vector and matrix instructions to dedicated functional modules, namely, operation executors, for parallel processing, supporting diverse tasks such as vector operations, matrix multiplication and accumulation, data reorganization, and conditional mask control. Furthermore, the accelerator is compatible with the RISC-V RVV instruction set through a two-stage decoding mechanism, ensuring ecosystem compatibility. Furthermore, memory access optimizations in the VLSU (such as AXI interface burst transfers) and cross-channel data processing in the SLDU (such as reduction and shuffle operations) further reduce latency and power consumption. Each channel of the vector register contains multiple single ports for parallel storage of data blocks. This design improves CPU performance in large-scale parallel computing, achieving high-performance, low-power real-time computing goals.

[0093] Below Figure 6 The present invention is described using a specific target hardware accelerator structure diagram as an example. The target hardware accelerator includes an instruction decoder (Dispatcher), an instruction scheduler (Sequencer), and an operation executor. The specific functions are as follows:

[0094] 1) The instruction decoder (Dispatcher) is responsible for connecting CPU requests with the hardware accelerator. The instruction decoder consists of a first-level instruction decoder and a second-level instruction decoder, performing two-stage decoding of instructions. The first-level decoder, located at the ID stage on the front end of the CPU, simply distinguishes whether the current instruction is a vector instruction or a matrix instruction. The second-level decoder, located at the front end of the hardware accelerator, fully decodes the vector instruction and determines whether it requires a scalar source operand or a destination operand that generates a scalar.

[0095] 2) The instruction scheduler (Sequencer) is a global scheduler for the hardware accelerator. Its primary function is to manage and control the execution order and dependencies of instructions in parallel vector computations, ensuring coordinated operation with the interfaces of the various functional units in the hardware accelerator. It optimizes the allocation of computing resources and improves execution efficiency by recording and analyzing instruction status and conflicts. Furthermore, the module implements a waiting and processing mechanism for outstanding instructions to ensure correctness in complex data paths.

[0096] 3) The operation executor includes a vector matrix operator, a vector mask controller, a cross-channel instruction processor, and a vector memory manager, as follows:

[0097] 3.1) The vector-matrix operator comprises a vector register, a matrix register, a first operator for vector arithmetic instruction operations, a second operator for performing floating-point and multiplication and division operations, a third operator for matrix arithmetic instruction operations, and a fourth operator for matrix multiplication and accumulation operations; wherein each channel of the vector register comprises a plurality of single ports for storing data blocks in parallel; the matrix register comprises a plurality of two-dimensional matrix registers, and the number of rows and columns of each two-dimensional matrix register is determined based on the row length of the two-dimensional matrix register;

[0098] 3.2) The vector mask controller (MASKU) performs bit-level operations on the target element corresponding to the current instruction to complete the mask operation of the current instruction; the bit-level operation is any one or more of the following operations: bitwise AND operation, bitwise OR operation, and bitwise XOR operation;

[0099] 3.3) The cross-channel instruction processor (SLDU) performs any one or more of the following operations on the current instruction: data fusion operation, data extraction operation, reduction operation, data rearrangement operation, and data shift operation;

[0100] 3.4) The Vector Memory Manager (VLSU) obtains the memory access operation type of the current instruction and performs address calculation for the current instruction based on the memory access modes supported by the scalable vector instruction set to obtain a memory access request containing the memory access address. It then merges the memory access requests that meet the preset conditions and performs a memory access operation corresponding to the memory access operation type on the data at the memory address in the target storage area through the advanced scalable interface in response to the merged memory access request. The Vector Memory Manager has only one external storage interface (i.e., the AXI interface) with a bit width of 2B / DP-FLOP. This choice is made to ensure a balance between computing power and bandwidth. The Vector Memory Manager has an Address Generation Unit (AGU) to support the memory access modes in the RVV instruction set, including unit stride mode, stride mode, and indexed addressing mode. The VLSU merges the AGU's memory access requirements into burst accesses and accesses external storage through the AXI interface.

[0101] For example Figure 7 The following is a specific hardware accelerator data flow diagram, in which the current data processing task issued by the processor is obtained through a preset standardized interface between the processor and the target hardware accelerator. The data flow is specifically as follows:

[0102] 1) The first-level instruction decoder performs initial decoding on each processing instruction in the current data processing task issued by the processor;

[0103] 2) The second instruction decoder decodes each of the target processing instructions again to obtain a micro-operation and a required operand corresponding to each of the target processing instructions, and obtains the required operand from the processor;

[0104] 3) The instruction scheduler analyzes the dependencies and conflicts between the micro-operations, determines the execution order of the micro-operations based on the dependencies, conflicts, and computing resources of the operation executors, and schedules the target executors in the operation executors according to the execution order;

[0105] 4) The target executor is any one or more executors of the vector matrix operator, the vector mask controller, the cross-channel instruction processor, and the vector memory access manager. The target executor executes each of the micro-operations using the required operands. According to the execution order, each target executor is a current executor in turn, and the current instruction is any one of the target processing instructions after being decoded again.

[0106] 4.1) If the vector mask controller is the current executor, the vector mask controller performs a bit-level operation on the target element corresponding to the current instruction to complete the mask operation on the current instruction; the bit-level operation is any one or more of a bitwise AND operation, a bitwise OR operation, and a bitwise XOR operation;

[0107] 4.2) If the cross-channel instruction processor is the current executor, the cross-channel instruction processor performs any one or more operations selected from the group consisting of a data fusion operation, a data extraction operation, a reduction operation, a data rearrangement operation, and a data shift operation on the current instruction;

[0108] 4.3) If the vector memory access manager is the current executor, the vector memory access manager obtains a memory access operation type of a current instruction, performs address calculation on the current instruction according to a memory access mode supported by the extensible vector instruction set, to obtain a memory access request including a memory access address, merges the memory access requests that meet preset conditions, and performs a memory access operation corresponding to the memory access operation type on data at the memory address in a target storage area through an advanced extensible interface to respond to the merged memory access request;

[0109] 4.4) The vector-matrix unit (VMPU) can be dynamically configured (2 to 16 units) and includes vector registers, matrix registers, a first operator, a second operator, a third operator, and a fourth operator. The first operator (VALU) is used for vector arithmetic instructions, the second operator (VMFPU) is used for floating-point and multiplication and division operations, the third operator (MALU) is used for matrix arithmetic instructions, and the fourth operator (MAC) is used for matrix multiplication and accumulation operations.

[0110] It can be seen that the present invention is that it comprehensively considers key factors such as the design, coordination, and application of hardware accelerators, and significantly improves computing efficiency and flexibility through modular division of labor and scalable architecture. Specifically, the hardware accelerator adopts an external independent design, avoiding the performance bottleneck of traditional built-in VPU solutions. Through the Dispatcher and Sequencer, efficient instruction scheduling is achieved, and vector and matrix instructions are distributed to dedicated functional modules (such as Lane, MASKU, SLDU, VLSU) for parallel processing, supporting diverse tasks such as vector operations, matrix multiplication and accumulation, data reorganization, and conditional mask control. The number of Lane modules can be dynamically configured (2 to 16), combined with flexible parameter adjustment of vector registers (VRF) and matrix registers (MRF) (such as RLEN), to adapt to computing needs of different scales, especially suitable for high-performance real-time computing in artificial intelligence scenarios. In addition, the accelerator is compatible with the RISC-V RVV instruction set through a two-level decoding mechanism to ensure ecological compatibility. At the same time, the VLSU's memory access optimization (such as AXI interface burst transmission) and the SLDU's cross-channel data processing (such as reduction and shuffle) further reduce latency and power consumption. This design can improve the CPU's performance in large-scale parallel computing and achieve high-performance, low-power real-time computing goals.

[0111] Figure 8A schematic diagram of the structure of a data processing task execution device provided in an embodiment of the present invention is applied to a target hardware accelerator independently provided outside a processor; the device comprises:

[0112] A primary decoding module 11 is configured to perform primary decoding on each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of a target type; wherein the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type;

[0113] An order determination module 12, configured to determine an execution order of the target processing instructions after being re-decoded;

[0114] The operation scheduling module 13 is used to schedule the target executors in each operation executor to execute the operations corresponding to each re-decoded target processing instruction according to the execution order, so as to complete the current data processing task.

[0115] The beneficial effects are as follows: the present invention abandons the conventional solution of integrating the hardware accelerator into the processor pipeline and is applied to a target hardware accelerator independently set outside the processor, that is, the hardware accelerator is placed outside the processor. This layout makes the hardware accelerator independent of the processor, reduces the dependence on the processor pipeline, improves the scalability of the hardware accelerator, avoids the computing power bottleneck caused by the internal resource limitation of the processor, and provides more space for improving computing power; two-level decoding is adopted for the processing instructions of the data processing task. The initial decoding simply distinguishes the processing instructions, filters out the target processing instructions of the vector type and the matrix type, and then decodes the target processing instructions again, that is, performs full decoding. The two-level decoding has a clear division of labor, improves the instruction processing efficiency, and speeds up the data processing speed; further, the execution order of the target processing instructions after re-decoding is determined, and each target executor is scheduled according to the execution order. That is, different target executors are used to reasonably execute different operations. Each executor performs its duties and works together to efficiently process various instructions and improve the overall computing power.

[0116] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 9 This is a block diagram of an electronic device according to an exemplary embodiment. The content in the diagram should not be considered as any limitation on the scope of use of this application. The electronic device may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the data processing task execution method disclosed in any of the aforementioned embodiments. In addition, the electronic device in this embodiment may specifically be an electronic computer.

[0117] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0118] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0119] The operating system 221 is used to manage and control the hardware devices on the electronic device and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs that can be used to perform the data processing task execution method performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to perform other specific tasks.

[0120] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned method for executing a data processing task is implemented. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.

[0121] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.

[0122] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0124] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0125] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for executing a data processing task, characterized in that: Applicable to a target hardware accelerator independently provided outside a processor; the method comprises: Performing initial decoding on each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of a target type; wherein the current data processing task is a task in an image recognition model built based on a neural network, and the target type includes a vector type and a matrix type; determining an execution order of the re-decoded target processing instructions; The target executors in each operation executor are scheduled to execute operations corresponding to each re-decoded target processing instruction according to the execution order to complete the current data processing task.

2. The data processing task execution method according to claim 1, characterized in that: The initial decoding of each processing instruction in the current data processing task issued by the processor includes: Acquiring a current data processing task issued by the processor through a preset standardized interface between the processor and the target hardware accelerator; Performing initial decoding on the current data processing task.

3. The data processing task execution method according to claim 1, characterized in that: The target hardware accelerator includes an instruction decoder, and the instruction decoder includes a first-level instruction decoder and a second instruction decoder; The initial decoding of each processing instruction in the current data processing task issued by the processor includes: Using the first-level instruction decoder to initially decode each processing instruction in the current data processing task issued by the processor; Accordingly, determining the execution order of the target processing instructions after re-decoding includes: The second instruction decoder is used to decode each target processing instruction again to obtain a micro-operation corresponding to each target processing instruction and determine an execution order of each micro-operation.

4. The data processing task execution method according to claim 1, characterized in that: The target hardware accelerator includes an instruction scheduler; and determining the execution order of the re-decoded target processing instructions includes: Re-decoding each of the target processing instructions to obtain a micro-operation and a required operand corresponding to each of the target processing instructions, and obtaining the required operand from the processor; Utilizing the instruction scheduler to analyze dependencies and conflicts between the micro-operations, and determining an execution order of the micro-operations based on the dependencies, conflicts, and computing resources of the operation executors; Accordingly, scheduling the target executors in each operation executor to execute operations corresponding to each re-decoded target processing instruction according to the execution order includes: The target executors in each operation executor are scheduled according to the execution order so that the target executor executes each micro-operation using the required operands.

5. The data processing task execution method according to any one of claims 1 to 4, characterized in that: Each of the operation executors includes a vector matrix operator, a vector mask controller, a cross-channel instruction processor, and a vector memory access manager, and the target executor is any one or more executors among the operation executors.

6. The data processing task execution method according to claim 5, characterized in that: The vector-matrix operator includes a vector register, a matrix register, a first operator for vector arithmetic instruction operations, a second operator for floating-point and multiplication and division operations, a third operator for matrix arithmetic instruction operations, and a fourth operator for matrix multiplication and accumulation operations; wherein each channel of the vector register includes multiple single ports for parallel storage of data blocks; the matrix register includes multiple two-dimensional matrix registers, and the number of rows and columns of each of the two-dimensional matrix registers is determined based on the row length of the two-dimensional matrix register.

7. The data processing task execution method according to claim 5, characterized in that: The step of scheduling the target executors in each operation executor to execute operations corresponding to each re-decoded target processing instruction according to the execution order includes: The target executors in each operation executor are scheduled as the current executor in turn according to the execution order; If the vector mask controller is the current executor, the vector mask controller performs a bit-level operation on the target element corresponding to the current instruction to complete the mask operation on the current instruction; the bit-level operation is any one or more of a bit-and operation, a bit-or operation, and a bit-exclusive-or operation; If the cross-channel instruction processor is the current executor, the cross-channel instruction processor performs any one or more operations of a data fusion operation, a data extraction operation, a reduction calculation operation, a data rearrangement operation, and a data shift operation on the current instruction; If the vector memory access manager is the current executor, the vector memory access manager obtains a memory access operation type of a current instruction, performs address calculation on the current instruction according to a memory access mode supported by the extensible vector instruction set, to obtain a memory access request including a memory access address, merges the memory access requests that meet a preset condition, and performs a memory access operation corresponding to the memory access operation type on data at the memory address in a target storage area through an advanced extensible interface to respond to the merged memory access request; The current instruction is any one of the target processing instructions after being decoded again.

8. A data processing task execution device, characterized in that: Applicable to a target hardware accelerator independently arranged outside a processor; the device comprises: a primary decoding module, configured to perform primary decoding on each processing instruction in the current data processing task issued by the processor to filter out target processing instructions of a target type; wherein the current data processing task is a task in an image recognition model constructed based on a neural network, and the target type includes a vector type and a matrix type; An order determination module, configured to determine an execution order of the target processing instructions after being re-decoded; The operation scheduling module is used to schedule the target executors in each operation executor to execute the operations corresponding to each re-decoded target processing instruction according to the execution order, so as to complete the current data processing task.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the data processing task execution method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the data processing task execution method according to any one of claims 1 to 7.