Matrix accelerator based on real-time operating system and operation method thereof

By introducing the hard real-time API interface and hard real-time thread of the real-time operating system into the matrix accelerator, the problems of slow speed and high power consumption in large-scale calculations are solved, and efficient and stable matrix operations are realized, which are suitable for occasions with high time accuracy requirements.

CN120066621APending Publication Date: 2025-05-30ANHUI GUOXUN CORE MICROTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411990143.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional matrix operations rely on the CPU, resulting in slow calculation speed, high power consumption when handling large-scale matrix operations, and cannot meet the high time accuracy requirements.

Method used

The matrix accelerator based on the real-time operating system is used to call the hard real-time API interface of the real-time operating system, and allocate hard real-time threads to assist the matrix accelerator in completing calculations, improving computing efficiency and stability.

Benefits of technology

It realizes efficient and stable computing of matrix accelerator, which is suitable for occasions with high time accuracy requirements, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066621A_ABST
    Figure CN120066621A_ABST
Patent Text Reader

Abstract

The invention discloses a matrix accelerator operation method based on a real-time operating system. The method comprises the following steps: configuring a hard real-time API (Application Program Interface) of the real-time operating system; the real-time operating system obtains matrix accelerator information; the matrix accelerator loads matrix data; the matrix accelerator calls the hard real-time API interface, and the real-time operating system distributes a hard real-time thread for the matrix accelerator to plan the operation time of the matrix accelerator; the matrix accelerator executes calculation, and after calculation is completed, the real-time operation system recycles the hard real-time thread. The invention further discloses the matrix accelerator based on the real-time operating system, wherein the matrix accelerator uses the running method. The matrix accelerator based on the real-time operating system is combined with the real-time operating system, a hard real-time API interface of the real-time operating system can be called to assist operation, the matrix accelerator is more efficient and stable, and the matrix accelerator is suitable for occasions with high time precision requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and specifically to a matrix accelerator based on a real-time operating system and its operation method. Background Art

[0002] With the rapid development of fields such as big data and artificial intelligence, matrix operations play an increasingly important role in scientific research, engineering calculations, etc. Traditional matrix operations mainly rely on the CPU for calculation. However, the CPU has problems such as slow calculation speed and high power consumption when dealing with large-scale matrix operations. Summary of the Invention

[0003] In order to overcome the defects in the prior art, the present invention provides a matrix accelerator based on a real-time operating system and its operation method. The matrix accelerator based on the real-time operating system is combined with the real-time operating system. The matrix accelerator can call the hard real-time API interface of the real-time operating system to assist in the operation, making the matrix accelerator more efficient and stable, and suitable for occasions with high requirements for time accuracy.

[0004] To achieve the above object, the technical solution adopted by the present invention is: A matrix accelerator operation method based on a real-time operating system, including the following steps:

[0005] Configure the hard real-time API interface of the real-time operating system;

[0006] The real-time operating system obtains matrix accelerator information;

[0007] The matrix accelerator loads matrix data;

[0008] The matrix accelerator calls the hard real-time API interface, and the real-time operating system allocates a hard real-time thread for the matrix accelerator to plan the operation time of the matrix accelerator;

[0009] The matrix accelerator performs the calculation. After the calculation is completed, the real-time operating system reclaims the hard real-time thread.

[0010] Through the above technical solution, the matrix accelerator can call the hard real-time API interface during the calculation, and the real-time operating system allocates a hard real-time thread to assist in the calculation, making the matrix accelerator more efficient and stable, and suitable for occasions with high requirements for time accuracy.

[0011] Among them, the hard real-time thread has a high priority and can adopt an exclusive method when obtaining key resources to ensure that the matrix accelerator completes the calculation within the specified time.

[0012] Further, the matrix accelerator information includes version information and a signature.

[0013] Further, the step of the matrix accelerator loading matrix data includes:

[0014] Loading the rows and columns of the matrix, loading one row of the first matrix into the first register of the matrix accelerator, and loading one column of the second matrix into the second register of the matrix accelerator.

[0015] Further, the step of the matrix accelerator calling the hard real-time API interface includes:

[0016] The real-time operating system provides a hard real-time API interface, and the hard real-time API interface detects the version information and signature of the matrix accelerator and establishes a mutual binding relationship.

[0017] Further, the step of the matrix accelerator performing calculations includes:

[0018] Using a multiplication instruction to calculate the product of corresponding elements in parallel;

[0019] Using an addition instruction to accumulate the result of the multiplication into an accumulation register;

[0020] Using a store instruction to store the calculation result from the accumulation register back to memory.

[0021] Further, the first register, the second register, and the accumulation register include 128-bit vector registers. Compared with traditional 64-bit registers, the matrix accelerator allows processing of up to 2 channels of 32-bit matrix data simultaneously; 128-bit vector registers allow the matrix accelerator to process 4 channels of 32-bit matrix data simultaneously. This design enables the matrix accelerator to process more data simultaneously and improves the data processing efficiency.

[0022] A matrix accelerator based on a real-time operating system includes:

[0023] A real-time operating system with a hard real-time API interface provided therein;

[0024] A matrix accelerator adopting a single instruction multiple data stream extension architecture and having 128-bit vector registers.

[0025] Further, based on the 128-bit vector registers, the matrix accelerator can process 4 channels of 32-bit matrix data simultaneously.

[0026] Further, based on the single instruction multiple data stream extension architecture, the matrix accelerator can perform the same operation on multiple data elements within a single instruction cycle.

[0027] By means of the above technical solutions, the beneficial effects of the present invention are as follows:

[0028] 1. This application combines a matrix accelerator with a real-time operating system. The matrix accelerator can call the hard real-time API interface of the real-time operating system, enabling the real-time operating system to allocate hard real-time threads to assist the matrix accelerator in completing calculations, making the matrix accelerator more efficient and stable, and suitable for scenarios with high requirements for time accuracy.

[0029] To make the above and other objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following briefly introduces the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0031] Figure 1 is a flowchart of the steps of the matrix accelerator operation method based on a real-time operating system in an embodiment of the present invention;

[0032] Figure 2 is a flowchart of the instruction execution of the matrix operation acceleration scheme of the matrix accelerator based on a real-time operating system in an embodiment of the present invention;

[0033] Figure 3 is a flowchart of the operation of the matrix accelerator based on a real-time operating system in an embodiment of the present invention for operating 4-way 32-bit matrix data. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0035] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and to distinguish similar objects, and there is no sequential order between them, nor can they be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0036] Embodiment: An embodiment of the present invention discloses a matrix accelerator based on a real-time operating system, including:

[0037] NECRO real-time operating system, the NECRO real-time operating system adopts a hard real-time kernel and a preemptive scheduling algorithm, which can ensure that high-priority tasks can seize resources and execute in time, ensuring the real-time performance and response speed of the system. The NECRO real-time operating system has microsecond jitter and nanosecond interrupt response speed, which enables it to accurately control the execution time of tasks and ensure efficient and accurate transmission of signals in the system. The NECRO real-time operating system is equipped with a hard real-time API interface, which is an application programming interface specially designed for hard real-time systems. The hard real-time API interface can ensure that the execution time and response time of tasks are strictly deterministic. No matter how the system load changes, it can ensure that key tasks are completed within the predetermined time, meeting the strict requirements of the hard real-time system for time accuracy.

[0038] NMA matrix accelerator, the NMA matrix accelerator adopts a single instruction multiple data stream extension architecture, allowing multiple operations to be executed simultaneously, for example, a 128-bit vector load, a 128-bit vector store, and two 128-bit vector operations can be executed in one instruction cycle, thereby improving instruction-level parallelism. The NMA matrix accelerator is designed to work with ARM's multi-core processor architecture, making parallel algorithms running on multi-core processors more efficient. The NMA matrix accelerator is designed to be compatible with the ARM Cortex-A series processors, which means that the code optimized for the NMA matrix accelerator can run on different ARM processors. The NMA matrix accelerator is designed to support non-aligned memory access mode to reduce the latency and bandwidth requirements of memory access. The programming model of the NMA matrix accelerator is designed to support high-level languages ​​combined with the compiler's automatic vectorization function to simplify the process of optimizing code. The high-level languages ​​include common computer languages ​​such as C and C++.

[0039] The NMA matrix accelerator adds a 128-bit vector register. Compared with the traditional 64-bit register, the matrix accelerator is allowed to process up to 2 channels of 32-bit matrix data at the same time; the 128-bit vector register allows the matrix accelerator to process 4 channels of 32-bit matrix data at the same time. This design allows the matrix accelerator to process more data at the same time, improving the efficiency of data processing.

[0040] The operation method of the matrix accelerator based on the real-time operating system is as follows:

[0041] Configure the hard real-time API interface of NECRO real-time operating system;

[0042] NECRO real-time operating system obtains NMA matrix accelerator information;

[0043] The NECRO real-time operating system provides a hard real-time API interface, and the hard real-time API interface detects information such as the version information and signature of the matrix accelerator;

[0044] If the hard real-time API interface detects the version information and signature of the NMA matrix accelerator, the hard real-time API interface and the NMA matrix accelerator establish a binding relationship, allowing the NMA matrix accelerator to call the hard real-time API interface. The binding process is as follows: The NMA matrix accelerator that runs for the first time sends a binding request to the system. The hard real-time API interface detects the version information and signature of the NMA matrix accelerator and sends an acknowledgement code to the NMA matrix accelerator. After receiving the acknowledgement code, the NMA matrix accelerator establishes a binding relationship with the hard real-time API interface and marks the NMA matrix acceleration signature as the activated state.

[0045] The NMA matrix accelerator loads matrix data:

[0046] Load the rows and columns of the matrix, load one row of the first matrix into the first register of the matrix accelerator, and load one column of the second matrix into the second register of the matrix accelerator;

[0047] After loading the matrix data, the NMA matrix accelerator calls the hard real-time API interface, and the NECRO real-time operating system allocates a hard real-time thread to assist in the operation for the NMA matrix accelerator to ensure that the NMA matrix accelerator completes the calculation within the specified time;

[0048] The NMA matrix accelerator executes the calculation process:

[0049] Use the multiplication instruction to calculate the product of the corresponding elements in parallel;

[0050] Use the addition instruction to accumulate the result of the multiplication into the accumulator register;

[0051] Use the store instruction to store the calculation result from the accumulator register back to the memory;

[0052] The flowchart of the matrix accelerator based on the real-time operating system for simultaneously processing 4 channels of 32-bit matrix data A, B, C, and D is as Figure 3 shown.

[0053] Among them, the hard real-time thread has a high priority and can adopt an exclusive mode when acquiring critical resources to ensure that the matrix accelerator completes the calculation within the specified time.

[0054] After completing the calculation, the real-time operating system reclaims the hard real-time thread to release the resources it occupies.

[0055] In the present invention, specific embodiments are used to illustrate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A matrix accelerator operation method based on a real-time operating system, characterized in that: The following steps are involved: Configure the hard real-time API interface of the real-time operating system; The real-time operating system obtains matrix accelerator information; The matrix accelerator loads matrix data; The matrix accelerator calls the hard real-time API interface, and the real-time operating system allocates a hard real-time thread to the matrix accelerator to plan the operation time of the matrix accelerator; The matrix accelerator performs calculations, and after the calculations are completed, the real-time operating system reclaims the hard real-time thread.

2. The matrix accelerator operation method based on a real-time operating system according to claim 1, characterized in that: The matrix accelerator information includes version information and a signature.

3. The matrix accelerator operation method based on a real-time operating system as claimed in claim 2, characterized in that: The steps to load matrix data into the matrix accelerator include: The rows and columns of matrices are loaded, a row of a first matrix is ​​loaded into a first register of the matrix accelerator, and a column of a second matrix is ​​loaded into a second register of the matrix accelerator.

4. The matrix accelerator operation method based on a real-time operating system as claimed in claim 3, characterized in that: The steps of the matrix accelerator calling the hard real-time API interface include: The real-time operating system provides a hard real-time API interface, and the hard real-time API interface detects the version information and signature of the matrix accelerator and establishes a mutual binding relationship.

5. The matrix accelerator operation method based on a real-time operating system as claimed in claim 4, characterized in that: The steps in the matrix accelerator to perform calculations include: Use multiplication instructions to compute the products of corresponding elements in parallel; Use the add instruction to accumulate the result of the multiplication into the accumulator register; Use the store instruction to store the result of the calculation from the accumulator register back to the memory.

6. The matrix accelerator operation method based on a real-time operating system as claimed in claim 5, characterized in that: The first register, the second register and the accumulation register include a 128-bit vector register.

7. A matrix accelerator based on a real-time operating system that executes the matrix accelerator operation method based on a real-time operating system according to any one of claims 1 to 6, characterized in that: include: A real-time operating system, wherein the real-time operating system is provided with a hard real-time API interface; A matrix accelerator, wherein the matrix accelerator adopts a single instruction multiple data stream extension architecture and is provided with a 128-bit vector register.

8. The matrix accelerator based on a real-time operating system as claimed in claim 7, characterized in that: Based on the 128-bit vector register, the matrix accelerator can process 4 channels of 32-bit matrix data simultaneously.

9. The matrix accelerator based on a real-time operating system as claimed in claim 8, characterized in that: Based on the SIMD extended architecture, the matrix accelerator can perform the same operation on multiple data elements in a single instruction cycle.