Cross-system NPU computing power calling system and method applied to open source gap system

By introducing the RKNPU API library and driver into the open source Hongmeng system, combining DMA and IOMMU technologies, optimizing memory access and computing pipelines, the problem of imperfect NPU support in the open source Hongmeng system is solved, and efficient AI task processing and low-energy AI computing are achieved.

CN120762890APending Publication Date: 2025-10-10GUANGDONG TELEPOWER TELECOM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510856627.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The existing open source Hongmeng system has incomplete support for NPU, resulting in low efficiency and high energy consumption in AI task processing, and increased equipment costs. In particular, it cannot fully exert its performance in scenarios that require a large amount of parallel computing, such as image recognition and speech processing.

Method used

Provides a cross-system NPU computing power calling system and method, which implements access to NPU hardware in the application layer and kernel layer of the open source Hongmeng system through the RKNPU API library and driver. Combined with DMA and IOMMU technology, it optimizes memory access and computing pipelines, reducing dependence on the CPU and GPU.

Benefits of technology

It improves the computing efficiency of AI tasks, enhances user experience, reduces hardware costs and energy consumption, and extends device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762890A_ABST
    Figure CN120762890A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-system NPU computing power calling system and a cross-system NPU computing power calling method applied to an open-source gap system, the cross-system NPU computing power calling system comprises an RKNPU API library and an RKNPU driver, and the RKNPU API library is arranged on an application layer of the open-source gap system; the RKNPU driver is arranged on a kernel layer of the open-source swan-monk system and is in butt joint with the RKNPU driver of the Linux, and when an application program of the open-source swan-monk system runs, the RKNPU driver of the kernel layer is called through the RKNPU API, and RKNPU hardware is indirectly accessed. According to the method, the AI task is processed by adopting better computing power hardware, so that the AI task processing efficiency is improved, and the user experience is improved; and the load of the CPU and the GPU is shared, the dependence on the high-performance CPU and GPU is reduced, and the energy consumption of an equipment end when the AI task is executed is reduced, so that the overall hardware cost is reduced, and the cruising ability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a cross-system NPU computing power calling system and method applied to the open source Hongmeng system. Background Art

[0002] With the emergence of large models, intelligent applications have an increasingly strong demand for artificial intelligence computing power; traditional general-purpose computing power can no longer meet such demands. Unlike traditional general-purpose computing, artificial intelligence computing requires the use of dedicated computing hardware such as GPUs and NPUs to provide artificial intelligence computing power. However, the OpenHarmony system's support for NPUs is not yet complete, and artificial intelligence computing can only use computing hardware such as CPUs and GPUs for the time being, which increases the cost of equipment manufacturing and use. In addition, CPUs and GPUs are less efficient when processing AI tasks, especially in scenarios such as image recognition and speech processing that require a large amount of parallel computing, and their performance cannot be fully utilized. Because the architecture of the CPU and GPU is not designed specifically for AI tasks, it takes a long time to process complex AI tasks, affecting the user experience. In addition, the CPU and GPU consume more energy when performing AI tasks, which is not conducive to the battery life of mobile devices and IoT devices.

[0003] In order to enhance some computing power functions, the present invention proposes a system and method for calling NPU computing power across systems, making AI task calculations more efficient, improving user experience, and sharing the load of the CPU and GPU, reducing dependence on high-performance CPUs and GPUs, thereby reducing overall hardware costs. Summary of the Invention

[0004] In order to overcome one or more of the above-mentioned technical defects, the present invention provides a cross-system NPU computing power calling system and method applied to the open source Hongmeng system, which improves the computing efficiency of AI tasks, enhances user experience, shares the load of CPU and GPU, reduces dependence on high-performance CPU and GPU, and thus reduces overall hardware costs.

[0005] In order to solve the above problems, the first aspect of the present invention discloses a cross-system NPU computing power calling system applied to the open source Hongmeng system, including:

[0006] RKNPU API library, set in the open source Hongmeng system application layer;

[0007] The RKNPU driver is set in the kernel layer of the open source Hongmeng system and connects to the RKNPU driver of Linux. When the open source Hongmeng system application is running, it calls the RKNPU driver of the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

[0008] Furthermore, it also includes:

[0009] Verification unit, used to verify the performance of RKNPU using rknn_yolov5_demo;

[0010] Configuration unit, used to configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo.

[0011] Furthermore, it also includes an RKNPU initialization unit for initializing the RKNPU driver and the RKNPU.

[0012] The second aspect of the present invention discloses a cross-system RKNPU computing power calling method applied to the open source Hongmeng system, which is applied to the above-mentioned cross-system NPU computing power calling system, including:

[0013] Start the RKNPU driver of the kernel layer;

[0014] Run the open source Hongmeng application and call the RKNPU driver at the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

[0015] Furthermore, before starting the RKNPU driver of the kernel layer, the following steps are included:

[0016] Use RKNPU_INIT function to initialize RKNPU hardware;

[0017] Use the platform driver framework to register the RKNPU driver and initialize the RKNPU driver through the probe function;

[0018] Obtain and register the interrupt handling function according to the RKNPU interrupt number;

[0019] Bind the IOMMU device in the RKNPU hardware;

[0020] Create a fast memory access channel.

[0021] Furthermore, the creation of a memory fast access channel includes:

[0022] Initialize DMA, allocate buffer area for DMA and register DMA channel;

[0023] Get the physical address of the RKNPU register and map the physical address of the register to the virtual address of the user state through the mmap function.

[0024] Furthermore, it also includes:

[0025] Configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo;

[0026] Use rknn_yolov5_demo to verify the performance of RKNPU.

[0027] Furthermore, the configuration of the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo includes:

[0028] Match ld-linux-aarch.so.1 library;

[0029] According to the GCC version of the open source Hongmeng system SDK, match the corresponding version of the C standard library;

[0030] According to the open source Hongmeng system SDK version, determine the corresponding dependent library version and GCC compilation chain to generate the so library.

[0031] A third aspect of the present invention discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program or instructions, and the processor implements the steps of the above method when executing the computer program or instructions.

[0032] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when executed by a processor.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention discloses a cross-system NPU computing power calling system and method applied to the open source Hongmeng system. The NPU computing power system includes an RKNPU API library and an RKNPU driver. The RKNPU API library is set in the application layer of the open source Hongmeng system; the RKNPU driver is set in the kernel layer of the open source Hongmeng system and docked with the RKNPU driver of Linux. When the open source Hongmeng system application is running, the RKNPU driver of the kernel layer is called through the RKNPU API to indirectly access the RKNPU hardware; by adopting better computing power hardware to process AI tasks, the efficiency of AI task processing is improved and the user experience is enhanced; and the load of the CPU and GPU is shared, the dependence on high-performance CPU and GPU is reduced, and the energy consumption of the device side when executing AI tasks is reduced, thereby reducing the overall hardware cost and improving battery life. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings, wherein:

[0036] Figure 1 This is an architectural diagram of the cross-system NPU computing power call system applied to the open source Hongmeng system described in the embodiment;

[0037] Figure 2This is a flowchart of the cross-system NPU computing power calling method applied to the open source Hongmeng system described in the embodiment;

[0038] Figure 3 The process of step S100 of the cross-system NPU computing power calling method applied to the open source Hongmeng system described in the embodiment Figure 1 ;

[0039] Figure 4 This is a flowchart of step S150 of the cross-system NPU computing power calling method applied to the open source Hongmeng system described in the embodiment;

[0040] Figure 5 The process of step S100 of the cross-system NPU computing power calling method applied to the open source Hongmeng system described in the embodiment Figure 2 . DETAILED DESCRIPTION

[0041] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0042] It should be noted that the term "comprising" and any variations thereof in the specification and claims of the present invention and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products, or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The order of execution of the named or numbered process steps can be changed according to the technical objectives to be achieved, as long as the same or similar technical effects are achieved. The division of units in this application is a logical division. In actual applications, other divisions may be used. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not performed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some ports, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present application.

[0043] RKNPU: RKNPU (Rockchip Neural Processing Unit) is a hardware unit developed by Rockchip Microelectronics specifically for accelerating neural network computing.

[0044] DMA: Direct Memory Access, direct memory access.

[0045] The embodiment of the present application provides a cross-system NPU computing power calling system applied to the open source Hongmeng system, such as Figure 1 , including RKNPU API library and RKNPU driver. The RKNPU API library is set in the application layer of the open source Hongmeng system; the RKNPU driver is set in the kernel layer of the open source Hongmeng system and connected to the RKNPU driver of Linux. When the open source Hongmeng system application is running, the RKNPU driver of the kernel layer is called through the RKNPU API, and the RKNPU driver of the kernel layer accesses the RKNPU hardware through the RKNPU driver of Linux.

[0046] In one embodiment, it also includes a verification unit and a configuration unit. The verification unit is used to verify the performance of RKNPU using the inference model rknn_yolov5_demo, and the configuration unit is used to configure the compilation environment and cross-compilation path environment variables of the inference model rknn_yolov5_demo.

[0047] In one embodiment, an RKNPU initialization unit is further included, which is used to initialize the RKNPU driver and the RKNPU.

[0048] The embodiment of the present application also provides a cross-system RKNPU computing power calling method applied to the open source Hongmeng system, which is applied to the above-mentioned cross-system NPU computing power calling system, such as Figure 2 ,include:

[0049] 100. Port the RKNPU driver to the open source Hongmeng system kernel layer, and port the RKNPU API library to the open source Hongmeng system application layer.

[0050] 200. Start the RKNPU driver of the kernel layer.

[0051] 300. Run the open source Hongmeng application, call the RKNPU driver of the kernel layer through the RKNPU API, and indirectly access the RKNPU hardware.

[0052] Specifically, before starting the RKNPU driver of the kernel layer, it includes adding the RKNPU node in the open source Hongmeng system kernel layer and configuring the RKNPU hardware:

[0053] Define the RKNPU register address, RKNPU interrupt number and working clock.

[0054] Define the power domain, dynamic voltage, and IOMMU device of the RKNPU. The IOMMU device is an input / output memory management unit (I / O memory management unit). It is used to process memory access requests from I / O devices, map the device's virtual address to a physical address, and isolate and protect the device's memory access, preventing the device from accessing unauthorized memory areas and improving system security and reliability.

[0055] In one embodiment, a single driver can be compatible with different RKNPU hardware through dynamic device tree overlay:

[0056] 1) Configure kernel support to ensure that the device tree overlay function is supported.

[0057] CONFIG_OF_OVERLAY=y

[0058] CONFIG_CONFIGFS_FS=y

[0059] 2) Write the RKNPU device tree overlay .dts file to describe the RKNPU device tree nodes and properties that need to be modified.

[0060] 3) Compile the device tree overlay file and use the device tree compiler (`dtc`) to compile the `.dts` file into a `.dtbo` file, such as dtc -I dts -O dtb -o my_rknpu.dtbo my_rknpu.dts.

[0061] 4) Load the device tree overlay file. The device tree overlay file can be loaded into the kernel through the `sysfs` interface, cpmy_rknpu.dtbo / sys / kernel / config / device-tree / overlays / .

[0062] 5) Check whether the device tree node is modified correctly.

[0063] In one embodiment, the RKNPU driver is ported to the kernel layer of the open source Hongmeng system, such as Figure 3 ,include:

[0064] 110. Use the RKNPU_INIT function to initialize the RKNPU hardware.

[0065] 120. Use the platform driver framework to register the RKNPU driver and initialize the RKNPU driver through the probe function.

[0066] 130. Obtain and register an interrupt handling function according to the RKNPU interrupt number, and initialize the working clock of the RKNPU.

[0067] 140. Bind the IOMMU device in the RKNPU hardware.

[0068] 150. Create a quick memory access channel.

[0069] In one embodiment, the creation of a memory fast access channel, such as Figure 4 ,include:

[0070] 151. Initialize DMA, allocate buffers for DMA and register DMA channels. The device can directly access memory through the DMA channel without going through the CPU.

[0071] 152. Get the physical address of the RKNPU register and map it to the user-mode virtual address using the mmap function. Use the paging mechanism to load data into memory on demand, reducing unnecessary I / O operations and improving memory utilization efficiency.

[0072] When running open-source Hongmeng applications and calling NPU computing power across systems, data is directly written to memory through DMA channels and / or mmap, facilitating NPU operations, improving computing efficiency, and reducing CPU occupancy.

[0073] In one embodiment, memory sharing between CPU / NPU / GPU is achieved through DMA_BUF+ION.

[0074] Specifically, in kernel space:

[0075] 1) Allocate shared memory: In kernel space, use dma_buf to allocate shared memory.

[0076] 2) Create a dma_buf exporter: Create a dma_buf through dma_buf_export and initialize its members, allocate an anonymous inode to obtain a file, and then bind the allocated buffer and dmabuf structure and associate them with the file to achieve sharing.

[0077] 3) Map shared memory to process address space: Use functions such as dma_buf_vmap or dma_buf_kmap to map shared memory to the kernel virtual address space so that the CPU can access and operate the data in the shared memory.

[0078] 4) Implement DMA operations: For peripherals such as GPU / NPU, the DMA engine is used to transfer data in shared memory to the peripheral's memory space, or from the peripheral's memory space to shared memory, to achieve data sharing and interaction.

[0079] On the userspace side:

[0080] 1) Open the ion device and create a client: The user space program accesses the ION subsystem by opening the / dev / ion device file and creates an ionclient for subsequent memory allocation and sharing operations.

[0081] 2) Allocate ion memory: Use ioctl to call the ION_IOC_ALLOC command to allocate ion memory, specifying parameters such as memory size, alignment, and flags.

[0082] 3) Pass the dma_buf file descriptor to other processes or modules: Pass the dma_buf file descriptor to other processes or modules that need to share memory through inter-process communication mechanisms (such as socketpair, binder, etc.).

[0083] 4) Mapping shared memory to process address space: In the receiving process or module, use the mmap function to map the dma_buf file descriptor to the virtual address space of the process, so that the data in the shared memory can be accessed like ordinary memory.

[0084] In terms of memory synchronization and management:

[0085] 1) Cache synchronization: Use functions such as dma_buf_begin_cpu_access and dma_buf_end_cpu_access to notify the system to start and end CPU access to shared memory so that the system can perform necessary cache synchronization operations.

[0086] 2) Reference count management: dma_buf has a comprehensive reference counting mechanism. When shared memory is mapped to multiple processes or used by multiple modules, the system automatically manages the reference count. When releasing shared memory, call dma_buf_put to decrement the reference count. When the reference count reaches zero, the kernel automatically releases the corresponding memory resources.

[0087] In one embodiment, sched_setaffinity is also used to improve throughput:

[0088] sched_setaffinity sets the CPU affinity mask for a process or thread, specifying the CPU cores on which it can run. This prevents frequent thread switching between different cores, reduces context switching overhead and cache invalidation, and thus improves system throughput.

[0089] For example, to set the CPU affinity mask:

[0090] cpu_set_t mas k;

[0091] CPU_ZERO(&mask); / / Clear the mask

[0092] CPU_SET(1,&mask); / / Assume it is bound to CPU core 1

[0093] Specifically, it also includes loading the power domain and dynamic voltage in the RKNPU hardware, dynamically adjusting the frequency, and cooperating with the of_devfreq_cooling_register_power interface of the Linux kernel to achieve temperature control.

[0094] Specifically, it connects to the Linux DRM (Direct Rendering Manager) subsystem to support direct output of NPU calculation results to the display pipeline, such as sharing Tensor data through GEM objects.

[0095] Specifically, user-mode interface registration can be implemented, and through subsystems such as misc, sysfs, and debugfs, operation and data query interfaces can be provided for the open source Hongmeng system application layer.

[0096] In one embodiment, the RKNPU API library is ported to the open source Hongmeng system application layer, such as Figure 5 ,include:

[0097] 160. Port RKNPU API library and dependent libraries, including rga, mpp, zlmediakit, etc.

[0098] 170. Port the inference model rknn_yolov5_demo to the open source Hongmeng system SDK.

[0099] 180. Configure the compilation environment and cross-compilation path environment variables of the inference model rknn_yolov5_demo.

[0100] 190. Use the inference model rknn_yolov5_demo to verify the performance of RKNPU.

[0101] Specifically, the CMakeLists.txt of the inference model rknn_yolov5_demo is modified, the aarch64-linux-gnu- compiler chain of the open source Hongmeng system SDK is adapted, and the dependent library path is adapted.

[0102] In one embodiment, the environment variables of the compilation environment and cross-compilation path of the rknn_yolov5_demo are configured, including:

[0103] Match the ld-linux-aarch.so.1 library, which is a dynamic linker, to parse and load the required shared library at runtime.

[0104] According to the GCC version of the open source Hongmeng system SDK, match the corresponding version of the C standard library.

[0105] According to the version of the open source Hongmeng system SDK, determine the corresponding dependent library version and GCC compilation chain, and generate a so library.

[0106] Compared with the prior art, the present application has the following advantages:

[0107] 1. Breakthrough in cross-architecture compatibility, seamless integration into Linux ecosystem:

[0108] Reuse the standard device driver model of the Linux mainline kernel (such as platform_driver, DMA-Engine, etc.), without relying on the HDF framework of the open source Hongmeng system, to achieve a 30% reduction in driver code; support adaptation to different RKNPU hardware versions (such as RK3588 / RK3568) through "Dynamic Device Tree Overlay (DTO)", avoiding code branching caused by hardware differences and improving portability and adaptability.

[0109] 2. Deep extension of performance optimization:

[0110] Zero-copy heterogeneous computing pipeline, based on DMA-BUF+ION memory framework to realize memory sharing between CPU / NPU / GPU, compared with traditional HDF memory copy scheme, the actual data throughput can be improved by 65%; in multi-core SoC (such as RK3588 Big.LITTLE architecture), bind NPU interrupt to big core processing, use sched_setaffinity to improve throughput.

[0111] 3. Reuse of ecosystem toolchain:

[0112] Interface with the Linux DRM (Direct Rendering Manager) subsystem, support NPU computing results directly output to the display pipeline, such as sharing Tensor data through GEM objects.

[0113] 4. Strengthening security mechanisms:

[0114] Adopt IOMMU protection, bind the IOMMU device in RKNPU hardware in the RKNPU driver of the open source Hongmeng system kernel layer, enable Linux IOMMU (SMMUv3) to remap the address of NPU DMA access, and prevent malicious DMA attacks.

[0115] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0116] Start the RKNPU driver of the kernel layer;

[0117] Run the open source Hongmeng application and call the RKNPU driver at the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

[0118] In one embodiment, when the processor executes the computer program, it further implements the following steps:

[0119] Use RKNPU_INIT function to initialize RKNPU hardware;

[0120] Use the platform driver framework to register the RKNPU driver and initialize the RKNPU driver through the probe function;

[0121] Obtain and register the interrupt handling function according to the RKNPU interrupt number;

[0122] Bind the IOMMU device in the RKNPU hardware;

[0123] Create a fast memory access channel.

[0124] In one embodiment, when the processor executes the computer program, it further implements the following steps:

[0125] Initialize DMA, allocate buffer area for DMA and register DMA channel;

[0126] Get the physical address of the RKNPU register and map the physical address of the register to the virtual address of the user state through the mmap function.

[0127] In one embodiment, when the processor executes the computer program, it further implements the following steps:

[0128] Configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo;

[0129] Use rknn_yolov5_demo to verify the performance of RKNPU.

[0130] In one embodiment, when the processor executes the computer program, it further implements the following steps:

[0131] Match ld-linux-aarch.so.1 library;

[0132] According to the GCC version of the open source Hongmeng system SDK, match the corresponding version of the C standard library;

[0133] According to the open source Hongmeng system SDK version, determine the corresponding dependent library version and GCC compilation chain to generate the so library.

[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0135] Start the RKNPU driver of the kernel layer;

[0136] Run the open source Hongmeng application and call the RKNPU driver at the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

[0137] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0138] Use RKNPU_INIT function to initialize RKNPU hardware;

[0139] Use the platform driver framework to register the RKNPU driver and initialize the RKNPU driver through the probe function;

[0140] Obtain and register the interrupt handling function according to the RKNPU interrupt number;

[0141] Bind the IOMMU device in the RKNPU hardware;

[0142] Create a fast memory access channel.

[0143] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0144] Initialize DMA, allocate buffer area for DMA and register DMA channel;

[0145] Get the physical address of the RKNPU register and map the physical address of the register to the virtual address of the user state through the mmap function.

[0146] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0147] Configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo;

[0148] Use rknn_yolov5_demo to verify the performance of RKNPU.

[0149] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0150] Match ld-linux-aarch.so.1 library;

[0151] According to the GCC version of the open source Hongmeng system SDK, match the corresponding version of the C standard library;

[0152] According to the open source Hongmeng system SDK version, determine the corresponding dependent library version and GCC compilation chain to generate the so library.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, or by means of dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the method described in the embodiments of the present application.

[0154] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0155] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0156] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Therefore, any modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A cross-system NPU computing power calling system applied to the open source Hongmeng system, characterized by: include: RKNPU API library, set in the open source Hongmeng system application layer; The RKNPU driver is set in the kernel layer of the open source Hongmeng system and connects to the RKNPU driver of Linux. When the open source Hongmeng system application is running, it calls the RKNPU driver of the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

2. The cross-system NPU computing power calling system according to claim 1, characterized in that: Also includes: Verification unit, used to verify the performance of RKNPU using rknn_yolov5_demo; Configuration unit, used to configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo.

3. The cross-system NPU computing power calling system according to claim 1, characterized in that: It also includes an RKNPU initialization unit for initializing the RKNPU driver and the RKNPU.

4. A cross-system RKNPU computing power calling method applied to the open source Hongmeng system, applied to the cross-system NPU computing power calling system according to any one of claims 1-3, characterized in that: include: Start the RKNPU driver of the kernel layer; Run the open source Hongmeng application and call the RKNPU driver at the kernel layer through the RKNPU API to indirectly access the RKNPU hardware.

5. The cross-system RKNPU computing power calling method according to claim 4 is characterized in that: Before starting the RKNPU driver of the kernel layer, the following steps are included: Use RKNPU_INIT function to initialize RKNPU hardware; Use the platform driver framework to register the RKNPU driver and initialize the RKNPU driver through the probe function; Obtain and register the interrupt handling function according to the RKNPU interrupt number; Bind the IOMMU device in the RKNPU hardware; Create a fast memory access channel.

6. The cross-system RKNPU computing power calling method according to claim 5, characterized in that: The step of creating a memory fast access channel includes: Initialize DMA, allocate buffer area for DMA and register DMA channel; Get the physical address of the RKNPU register and map the physical address of the register to the virtual address of the user state through the mmap function.

7. The cross-system RKNPU computing power calling method according to claim 4, characterized in that: Also includes: Configure the compilation environment and cross-compilation path environment variables of rknn_yolov5_demo; Use rknn_yolov5_demo to verify the performance of RKNPU.

8. The cross-system RKNPU computing power calling method according to claim 7, characterized in that: The environment variables for configuring the compilation environment and cross-compilation path of rknn_yolov5_demo include: Match ld-linux-aarch.so.1 library; According to the GCC version of the open source Hongmeng system SDK, match the corresponding version of the C standard library; According to the open source Hongmeng system SDK version, determine the corresponding dependent library version and GCC compilation chain to generate the so library.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program or instruction, wherein: When the processor executes the computer program or instructions, the steps of the method according to any one of claims 4 to 8 are implemented.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 4 to 8 are implemented.