Interface calling device and method for a heterogeneous computer system CPU to call a DSP

By designing the interface calling device and method for CPU to call DSP in a heterogeneous computer system, using SMC memory/GSM, shared memory and interrupts between IPC cores to interact, the programming difficulty and complex process problems of CPU calling DSP are solved, and the CPU can realize the separate and flexible calls of DSP calculation programs by the CPU, reducing the difficulty of user programming.

CN119806692BActive Publication Date: 2025-05-30HUNAN GREAT WALL GALAXY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510296101.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-30
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In existing heterogeneous computer systems, CPUs face programming difficulties and complex processes when calling DSP for multiple calculations, and cannot realize the CPU's separate and flexible calls to multiple DSP computing programs.

Method used

Design an interface calling device and method for the CPU invoking DSP in a heterogeneous computer system. Through the interaction between the CPU and DSP, SMC memory/GSM, shared memory and IPC core interrupts, the multi-category calculation of the CPU calling DSP is realized. The specific steps include: initializing the DSP heterogeneous computing call API, applying for continuous memory blocks, obtaining free DSPs, initiating a computing API request to the DSP, waiting for the calculation to be completed and releasing the memory block and exiting the computing call.

Benefits of technology

By encapsulating the main program of CPU and DSP as an API library, users do not need to care about the mechanism and underlying implementation details between heterogeneous calls. They only need to call the provided API to implement the CPU calling DSP for multi-category calculations, reducing the difficulty and complexity of user programming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806692B_ABST
    Figure CN119806692B_ABST
Patent Text Reader

Abstract

The present invention relates to an interface calling device and method for a CPU to call a DSP in a heterogeneous computer system. By registering DSP library program APIs as different computing APIs in the DSP program according to multiple categories of calculations required by the user, and then the CPU separately calls these computing APIs. In this way, in the heterogeneous calling design, the main programs of the CPU and the DSP are both encapsulated into API libraries, greatly reducing the programming difficulty for users. Users do not need to care about the mechanisms and underlying implementation details between heterogeneous calls, and only need to call the provided relevant APIs and data structures to achieve the CPU calling the DSP for multiple categories of calculations, thereby realizing the API-level call of the CPU calling the DSP in the heterogeneous computer system and solving the problems of difficult programming and complex processes faced by the CPU when calling the DSP for multiple calculations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer science and engineering, and relates to an interface calling device and method for a CPU to call a DSP in a heterogeneous computer system. Background Art

[0002] In recent years, with the increasing wide application of heterogeneous computers in the form of a central processing unit (CPU) + a digital signal processor (DSP) in the market, an efficient heterogeneous calling technology or method with moderate programming difficulty and easy development is extremely important for giving full play to the advantages of heterogeneous processors. Currently, the mature heterogeneous calling methods in the market include OpenCL and TI heterogeneous calling, etc. Among them, although using OpenCL can shield the underlying programming frameworks of each manufacturer in programming and achieve rapid transplantation of software between different platforms, the environment setup is complex, the development and transplantation are difficult, and it is also difficult to achieve the best performance of the hardware.

[0003] The technology of TI heterogeneous processor DSP core calling mainly compiles the computing program that the DSP needs to run into an.out file, and uses the method of the CPU igniting the DSP to parse the entire.out file, moves program segments, data segments, etc. to the corresponding memory spaces, and the CPU ignites the DSP to execute. This technology is not flexible enough and can only be used as a single computing task, and cannot achieve the separate and flexible calling of multiple computing programs of the CPU to the DSP. Therefore, in a heterogeneous computer system, how to solve the problems of difficult programming and complex process faced by the CPU when calling the DSP for multiple computations has important practical significance for practical applications. Summary of the Invention

[0004] Aiming at the problems existing in the above-mentioned traditional technologies, the present invention proposes an interface calling device for a CPU to call a DSP in a heterogeneous computer system and an interface calling method for a CPU to call a DSP in a heterogeneous computer system, which can solve the problems of difficult programming and complex process faced by the CPU when calling the DSP for multiple computations.

[0005] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0006] On the one hand, an interface calling device for a CPU of a heterogeneous computer system to call a DSP is provided, comprising a CPU and a DSP, wherein the CPU and the DSP interact with each other in heterogeneous computing by using SMC memory / GSM, shared memory and IPC inter-core interrupts, wherein the software stack of the CPU comprises an application, an API library and a driver, wherein the application calls an API provided by the API library to implement calling the DSP for heterogeneous computing, wherein the API library comprises a DSP heterogeneous computing calling library and an SMEM library, and the driver comprises an SMEM driver and a DSP heterogeneous computing calling driver; the DSP program comprises a DSP application program and a DSP library program, wherein the DSP heterogeneous computing kernel program is encapsulated into a DSP library program API, and is called by the DSP application program;

[0007] The CPU application calls the API provided by the API library to implement the process of calling the DSP for heterogeneous computing, including calling the DSP heterogeneous computing call library's initialization DSP heterogeneous computing call API to initialize the DSP heterogeneous computing call, calling the SMEM library's API to apply for continuous memory blocks from the SMC memory / GSM for storing calculation raw data and calculation result data, and obtaining the continuous memory block physical address, calling the DSP heterogeneous computing call library's API to obtain idle DSPs, obtain the DSP idle calculation API request structure and write the calculation data address and parameters into the calculation API request structure according to the user's calculation task, calling the SMEM library's API to write back and invalidate the cache, calling the DSP heterogeneous computing call library's API to initiate a calculation API request to the DSP, waiting for the DSP to return the calculation completion, calling the SMEM library's API to release the applied continuous memory block, and calling the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library to exit the DSP heterogeneous computing call;

[0008] Among them, the process of DSP application calling DSP library program API execution includes initializing DSP platform, creating calculation API function, running DSP heterogeneous calculation kernel program, setting the state of corresponding DSP core to idle state, judging whether CPU initiates IPC interrupt, if so, parsing IPC interrupt flag and sending IPC interrupt to CPU and executing corresponding user-defined calculation API according to the ID of calculation API request initiated by CPU, and setting the state of corresponding DSP core to idle state after DSP completes calculation.

[0009] On the other hand, a method for calling an interface of a CPU in a heterogeneous computer system to call a DSP is provided, comprising the steps of:

[0010] The application on the CPU side calls the DSP heterogeneous computing call library to initialize the DSP heterogeneous computing call API to initialize the DSP heterogeneous computing call;

[0011] Apply the API of the SMEM library through the application to apply for continuous memory blocks from the SMC memory / GSM to store the original calculation data and the calculation result data, and obtain the physical addresses of the continuous memory blocks;

[0012] Apply the API of the DSP heterogeneous computing call library through the application to obtain idle DSPs, obtain the idle calculation API request structure of the DSP, and write the calculation data address and parameters into the calculation API request structure according to the user's calculation task;

[0013] Apply the API of the SMEM library through the application to write back and invalidate the Cache, and apply the API of the DSP heterogeneous computing call library to initiate a calculation API request to the DSP;

[0014] Wait for the DSP to return the calculation completion, and then apply the API of the SMEM library through the application to release the applied continuous memory blocks;

[0015] Apply the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library through the application to exit the DSP heterogeneous computing call;

[0016] Among them, the process executed by the DSP application calling the DSP library program API includes initializing the DSP platform, creating a calculation API function, running the DSP heterogeneous computing kernel program, setting the status of the corresponding DSP core to the idle state, determining whether the CPU initiates an IPC interrupt. If so, after parsing the IPC interrupt flag, send an IPC interrupt to the CPU and execute the corresponding user-defined calculation API according to the ID of the calculation API request initiated by the CPU. After the DSP completes the calculation, set the status of the corresponding DSP core to the idle state.

[0017] One of the above technical solutions has the following advantages and beneficial effects:

[0018] The above interface calling device and method for the CPU of the heterogeneous computer system to call the DSP register different calculation APIs by calling the DSP library program API in the DSP program according to the multi-category calculations required by the user. The CPU separately calls these calculation APIs. In this way, in the heterogeneous call design, the main programs of the CPU and the DSP are both encapsulated into API libraries, which greatly reduces the programming difficulty of the user. The user does not need to care about the mechanism and underlying implementation details between the heterogeneous calls, and only needs to call the provided relevant APIs and data structures to implement the CPU to call the DSP for multi-category calculations, thus realizing the call at the calculation API level of the CPU calling the DSP in the heterogeneous computer system.

[0019] Compared with the prior art, it is more flexible in use. At the same time, the main programs for heterogeneous calls to the CPU and DSP are encapsulated into API libraries, standardizing the function interfaces, greatly reducing the programming difficulty for users. Users do not need to care about the mechanisms and underlying implementation details between heterogeneous calls. They only need to fill in the calculation request data structure and then call the relevant APIs provided to achieve the CPU to call the DSP for various types of calculations, and it does not depend on third-party uncontrollable program libraries. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a block diagram of the overall design scheme of the heterogeneous call system in an embodiment;

[0022] Figure 2 It is a schematic diagram of the technical implementation scheme of SMC memory / GSM interaction messages, calculation request data, etc. in an embodiment;

[0023] Figure 3 It is a schematic diagram of the software stack implementation scheme of the CPU in an embodiment;

[0024] Figure 4 It is a schematic diagram of the execution process of the user program calling the CPU library program API in an embodiment;

[0025] Figure 5 It is a schematic diagram of the overall overview of the technical implementation scheme module of the DSP heterogeneous computing kernel program in an embodiment;

[0026] Figure 6 It is a schematic diagram of the execution process of the DSP user program calling the DSP library program API in an embodiment;

[0027] Figure 7 It is a schematic diagram of the technical implementation scheme of the DSP core state machine in an embodiment;

[0028] Figure 8 It is a schematic diagram of the process of the interface call method for the CPU to call the DSP in the heterogeneous computer system in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0030] It should be noted that, as used herein, the mention of "embodiment" means that a particular feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase is presented at various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art can understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the description and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention.

[0032] In one embodiment, as Figure 1 shown, there is provided an interface call device for a CPU to call a DSP in a heterogeneous computer system, which may include a CPU and a DSP. Heterogeneous computing interaction is performed between the CPU and the DSP by using SMC memory / GSM, shared memory, and IPC inter-core interrupt. The software stack of the CPU includes an application program, an API library, and a driver. The application program calls the API provided by the API library to implement the call to the DSP for heterogeneous computing. The API library includes a DSP heterogeneous computing call library and an SMEM library. The driver includes an SMEM driver and a DSP heterogeneous computing call driver; the program of the DSP consists of a DSP application program and a DSP library program. Among them, the heterogeneous computing kernel program of the DSP is encapsulated into a DSP library program API and is called by the DSP application program.

[0033] The application program of the CPU calls the APIs provided by the API library to implement the process of calling the DSP for heterogeneous computing, including initializing the DSP heterogeneous computing call API of the DSP heterogeneous computing call library for DSP heterogeneous computing call initialization, calling the API of the SMEM library to apply for a continuous memory block from the SMC memory / GSM to store the original calculation data and the calculation result data, and obtaining the physical address of the continuous memory block, calling the API of the DSP heterogeneous computing call library to obtain an idle DSP, obtaining the idle calculation API request structure of the DSP and writing the calculation data address and parameters into the calculation API request structure according to the user's calculation task, calling the API of the SMEM library to write back and invalidate the Cache, calling the API of the DSP heterogeneous computing call library to send a calculation API request to the DSP, waiting for the DSP to return after the calculation is completed, calling the API of the SMEM library to release the applied continuous memory block, and calling the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library to exit the DSP heterogeneous computing call.

[0034] Among them, the process executed by the DSP application program calling the DSP library program API includes performing DSP platform initialization, creating a calculation API function, running the DSP heterogeneous computing kernel program, setting the status of the corresponding DSP core to the idle state, determining whether the CPU initiates an IPC interrupt, if so, parsing the IPC interrupt flag and sending an IPC interrupt to the CPU and executing the corresponding user-defined calculation API according to the ID of the calculation API request initiated by the CPU, and after the DSP completes the calculation, setting the status of the corresponding DSP core to the idle state.

[0035] It can be understood that in this embodiment, in the heterogeneous computer system, a call at the level of the CPU calling the DSP calculation API (Application Programming Interface) is implemented, that is: the user can implement various types of calculations to be performed as different calculation APIs in the DSP program, and the CPU can separately call these calculation APIs. In the heterogeneous call design, the main programs of the CPU and the DSP are both encapsulated into API libraries, which greatly reduces the programming difficulty of the user. The user can completely ignore the mechanisms and underlying implementation details between the heterogeneous calls, and only need to call the relevant APIs and data structures provided by this embodiment to implement the CPU calling the DSP for various types of calculations.

[0036] From the overall design, the overall design scheme block diagram of the heterogeneous call in the heterogeneous computer system can be as Figure 1As shown in the figure, in a heterogeneous computer system, the CPU calls the DSP for computing, mainly using SMC memory (Shared memory controller, i.e., shared memory controller memory) / GSM (Global Shared memory), shared memory, and IPC (Inter-Process Communication) inter-core interrupts for heterogeneous computing interaction. Among them, SMC memory / GSM is used to store the interaction message data and calculation parameter data between the CPU and the DSP, etc.; the shared memory is usually DDR and is used to store the original data of the calculation and the result data of the calculation; the IPC inter-core interrupt is used for the inter-core interrupt interaction between the CPU and the DSP. IRQ (Interrupt Request) is the interrupt request.

[0037] For the interface calling device of the above heterogeneous computer system where the CPU calls the DSP, by registering different calculation APIs by calling the DSP library program API in the DSP program according to the multi-category calculations that the user will need to perform, and then the CPU separately calls these calculation APIs. In this way, in the heterogeneous calling design, the main programs of the CPU and the DSP are both encapsulated into API libraries, which greatly reduces the programming difficulty for users. Users do not need to care about the mechanism and underlying implementation details between heterogeneous calls. They only need to call the provided relevant APIs and data structures (such as the calculation request data structure like the calculation API request structure, calculation parameter data structure, etc.) to achieve the CPU calling the DSP for multi-category calculations, thus realizing the API-level calling of the CPU calling the DSP in the heterogeneous computer system.

[0038] Compared with the prior art, it is more flexible. At the same time, the main programs of the heterogeneous calls of the CPU and the DSP are both encapsulated into API libraries, standardizing the function interfaces, greatly reducing the programming difficulty for users. Users do not need to care about the mechanism and underlying implementation details between heterogeneous calls. They only need to fill in the calculation request data structure and then call the provided relevant APIs to achieve the CPU calling the DSP for multi-category calculations, and it does not depend on third-party uncontrollable program libraries.

[0039] In one embodiment, the storage space of SMC memory / GSM is dynamically allocated by the application program by calling the API of the SMEM library to save the calculation request data structure. The last 4KB space of SMC memory / GSM is used to save the base address, size of the DSP calculation request data structure variable, and the message type of the IPC interrupt sent by the DSP to the CPU; the IPC interrupt sources sent by the CPU to the DSP include INT-0 and INT-1. INT-0 represents the handshake interrupt sent by the CPU, and INT-1 represents the API request interrupt sent by the CPU. The message types of the IPC interrupt include bit-0, bit-1, and bit-2. Bit-0 represents the DSP handshake interrupt, bit-1 represents the normal interrupt of the DSP calculation API request completion, and bit-2 represents the DSP exception interrupt.

[0040] It can be understood that the interrupt types sent by the CPU to the DSP include the handshake interrupt and the DSP calculation API request interrupt. The interrupt types sent by the DSP to the CPU include the handshake response interrupt, the calculation completion interrupt, and the DSP operation exception interrupt; when the DSP sends an IPC interrupt to the CPU, only one interrupt source is sent, and the interrupt type message is saved in the SMC memory / GSM.

[0041] As Figure 2 shown is a schematic diagram of the technical implementation solution of the interaction message data, calculation parameter data, etc. between SMC memory / GSM. Except for the last 4KB in the storage space of SMC memory / GSM, the other storage spaces are dynamically allocated by the CPU and are used to save the calculation request data structure; the last 4KB is used to save the base address, size of the DSP calculation request data structure variable, and the message type of the IPC interrupt sent by the DSP to the CPU. The detailed definition of this 4KB space can be shown in Table 1:

[0042] Table 1

[0043]

[0044] The CPU side is described through the Device Tree (a data structure for describing hardware, used to describe information such as the topological structure, attributes, and connection relationships of devices in a computer system), and the storage space of SMC memory / GSM is reserved, and the SMEM (Shared memory manager, a software component or system module for managing shared memory resources) driver is responsible for managing this continuous physical memory (that is, the previously reserved storage space of SMC memory / GSM).

[0045] Technical implementation of storing original calculation data, calculation result data, etc. in shared memory (DDR): For convenience and efficiency, the DSP needs to use continuous physical memory when obtaining original calculation data and filling in calculation result data. The CPU side reserves a section of continuous DDR physical memory through methods such as device tree description, and the SMEM driver is responsible for managing this section of continuous physical memory. Thus, it can be seen that the memory spaces involved in the functional processes of SMC memory / GSM and shared memory (DDR) are all reserved by the CPU side through device tree description for the SMC memory / GSM space and a certain section of DDR space, and the SMEM driver is uniformly responsible for managing these continuous physical memories. The SMEM library provides APIs, which are called by the CPU application layer and the DSP heterogeneous computing call library program. Users do not need to care about the underlying implementation at all, and can achieve operations such as application, release, virtual-to-physical address conversion, cache (high-speed buffer memory) refresh, and cache invalidation by calling the provided APIs for the corresponding section of memory.

[0046] Specifically, the technical implementation of the IPC inter-core interrupt interaction between the CPU and the DSP: It includes the CPU sending an IPC interrupt to the DSP, the CPU receiving an IPC interrupt sent by the DSP, the DSP sending an IPC interrupt to the CPU, and the DSP receiving an IPC interrupt sent by the CPU. The definition of the IPC interrupt source type can be as follows: The IPC interrupt sources sent by the CPU to the DSP include (1) INT-0: The handshake interrupt sent by the CPU, (2) INT-1: The API request interrupt; There is only 1 IPC interrupt source sent by the DSP to the CPU, and the interrupt type message is saved in the SMC memory / GSM area. After receiving the IPC interrupt sent by the DSP, the CPU reads the interrupt type message and performs corresponding interrupt processing. The bits of its interrupt message can be as follows: (a) bit-0: The DSP handshake interrupt, (b) bit-1: The normal interrupt indicating that the DSP calculation API request is completed, (c) bit-2: The DSP exception interrupt (such as program runaway, etc.).

[0047] In one embodiment, before the CPU initiates a calculation API request, corresponding member variables in the calculation request data structure are filled with the calculation API ID (i.e., the unique identity identifier of the API) to be called, the address of the original calculation data, the address of the calculation result data, and the calculation parameters; There is a member variable of an enumeration type in the calculation request data structure that stores the calculation API ID separately, and the API ID settings of the CPU are consistent with those of the DSP. In the DSP heterogeneous computing kernel program, the function pointers to be calculated are saved in an array by calling the DSP library program API and registered as different calculation APIs. The subscript of the calculation function array is the API ID for which the CPU calls the DSP for calculation; The subscript of the calculation function array excludes 0 to prevent invalid calls.

[0048] It can be understood that the technical implementation of CPU and DSP programming includes CPU programming and DSP programming.

[0049] On the one hand, as Figure 3 shown, the software stack implementation of the CPU in CPU programming includes three parts: the application layer, the middleware, and the driver layer. Among them, the application layer provides application programs (APPs), the middleware mainly provides API libraries including the DSP HCC lib library and the SMEM lib library, and the driver layer includes the SMEM Driver (SMEM driver) and the DSP HCC Driver (DSP HCC driver). Under the Linux operating system, the application program (APP) called by the DSP heterogeneous computing interface calls the APIs provided by the DSP HCC lib library and the SMEM lib library to implement the call to the DSP for heterogeneous computing.

[0050] In one embodiment, the DSP heterogeneous computing call library realizes sending IPC interrupts to the DSP core, setting the base address and size of the storage area of the calculation request data structure, calling the API provided by the SMEM library to apply for a section of memory in the SMC memory / GSM to store the calculation request data structure, encapsulating the DSP heterogeneous computing-related functions into standard APIs for application programs to call, realizing the maintenance of the DSP multi-core state machine, and receiving asynchronous signals sent by the DSP heterogeneous computing call driver by calling the standard ioctl interface provided by the DSP heterogeneous computing call driver.

[0051] The DSP heterogeneous computing call driver realizes obtaining IPC register resources, reserved space resources in the SMC memory / GSM, and IPC interrupt numbers from the device tree, as well as operating the underlying hardware, registering an interrupt service program, sending an asynchronous signal to the DSP heterogeneous computing call library after receiving the IPC interrupt sent by the DSP, and providing ioctl interface calls.

[0052] The SMEM library is used to provide the corresponding APIs for respectively implementing the application, release, virtual-to-physical and physical-to-virtual address conversion, cache refresh, and cache invalidation of continuous memory, and for calling the ioctl interface of the SMEM driver. The SMEM driver, as the management driver of the shared continuous memory block, is used to realize obtaining and loading driver module parameters from the device tree, realizing the management of configurable continuous physical memory areas, and providing the corresponding ioctl interfaces for the application, release, virtual-to-physical and physical-to-virtual address conversion, cache refresh, and cache invalidation of the continuous physical memory area called by the SMEM library; among them, the management of the continuous physical memory area includes the management of the SMC memory, GSM, and shared memory.

[0053] Specifically, the DSP HCC (Heterogeneous Computing Call) library program realizes the following functions by calling the standard ioctl (input / output control) interface provided by the DSP HCC driver: sending IPC interrupts to the DSP core, setting the base address and size of the storage area for the calculation request data structure, etc.; calling the APIs provided by the SMEM lib library program to apply for a section of memory in the SMC memory / GSM to store the calculation request data structure; encapsulating DSP heterogeneous computing-related functions into standard APIs for application programs (APPs) to call; implementing DSP multi-core state machine maintenance; receiving asynchronous signals sent by the DSP HCC driver.

[0054] The DSP HCC driver realizes obtaining IPC (Inter-Process Communication) register resources, reserved space resources in the SMC memory / GSM, and IPC interrupt numbers from the device tree, operating on underlying hardware such as IPC registers and SMC memory / GSM; registering an interrupt service program to receive IPC interrupts sent by the DSP and then sending asynchronous signals to the upper-layer program; providing function calls such as ioctl to the DSP HCC library program, and the provided ioctl can be as follows: (1) sending IPC interrupts to the DSP core, (2) setting the base address of the storage area for the calculation request data structure, (3) setting the size of the storage area for the calculation request data structure.

[0055] The SMEM lib library program is a library program for managing shared continuous memory blocks, providing APIs such as application, release, virtual-to-physical and physical-to-virtual address conversion, cache flushing, and cache invalidation for continuous memory, and calling the SMEM Driver ioctl to implement corresponding functions.

[0056] The SMEM driver is a driver for managing shared continuous memory blocks, realizing the management of continuously configurable physical memory areas obtained from the device tree and loaded with driver module parameters, including the management of SMC memory, GSM, and DDR, and providing ioctl such as application, release, virtual-to-physical and physical-to-virtual address conversion, cache flushing, and cache invalidation for the SMEM library program to call.

[0057] In one embodiment, the APIs provided by the API library of the software stack of the CPU include an initialization DSP heterogeneous computing call API, an offloading DSP heterogeneous computing call API, an idle acquisition API, a core status acquisition API, an idle structure acquisition API, a CPU call DSP for heterogeneous computing API, a wait for DSP heterogeneous computing API request completion API, an SMEM initialization API, an SMEM memory application API, an SMEM memory release API, an SMEM physical address acquisition API, an invalidate cache API, and a write-back and invalidate cache API.

[0058] Specifically, the API design provided by the CPU side for users (application programs) to call can be as follows:

[0059] (1) The initialization DSP heterogeneous computing call API includes the following functions: call the initialization API in the SMEM lib library program to initialize SMEM; call the memory application API in the SMEM lib library program to apply for a section of memory in the SMC memory / GSM to store the computing request data structure; call the virtual-to-physical address conversion API in the SMEM lib library program to obtain the physical address of the storage computing request data structure; initialize the computing request data structure and set the DSP status to UNKONW ("unknown status"); open the character device registered by the DSP HCC driver. When opening this character device, the DSP HCC driver will register an interrupt service program for receiving DSP IPC interrupts; register a function for receiving asynchronous signals of the DSP HCC driver; call the DSP HCC driver through ioctl to set the base address and size of the storage area of the computing request data structure; initiate a handshake with the DSP and send an IPC interrupt to the DSP through ioctl to ensure that the DSP is in a normal working state.

[0060] (2) The offloading DSP heterogeneous computing call API includes the following functions: call the memory release API in the SMEM lib library program to release the applied memory; close the character device registered by the DSP HCC driver.

[0061] (3) The idle acquisition API is used to acquire an idle DSP core. This API polls the status of all DSP cores, finds the first idle DSP core, and returns the core number of the first idle DSP core.

[0062] (4) The core status acquisition API is used to acquire the status of a specified DSP core. This API acquires the status of the DSP core through the core number of the DSP core.

[0063] (5)Idle structure acquisition API, which is used to acquire the idle computing API request structure of a specified DSP core. This API obtains an idle and unoccupied API request structure of a specified DSP core.

[0064] (6)CPU calls DSP for heterogeneous computing API, which is used for the CPU to initiate a computing API request to the DSP.

[0065] (7)Wait for DSP heterogeneous computing API request completion API, which is used to determine whether the DSP heterogeneous computing API request is completed. When this API is called, it will block until the specified DSP completes the computing task and enters the idle state.

[0066] (8)SMEM initialization API, which is used to initialize the shared memory management module.

[0067] (9)SMEM memory application API, which allocates continuous memory of a specified size from a specified memory block, supports allocating memory from DDR, SMC / GSM, supports whether the allocated memory has the Cache attribute, and supports byte alignment of the allocated memory as required.

[0068] (10)SMEM memory release API, which is used to release the memory block applied using SMEM.

[0069] (11)SMEM physical address acquisition API, which is used to convert the virtual address of the memory block start address into a physical address.

[0070] (12)Invalidate Cache API, which is used to invalidate the Cache of a memory area with a specified address and size, making it invalid.

[0071] (13)Write-back and invalidate Cache API, which is used to write back the cache of a memory area with a specified address and size to memory and make it invalid.

[0072] The user program calls the API provided by the library program in the CPU programming, and the execution process is as Figure 4 shown, where dsp_hcc_init is the DSP heterogeneous computing call initialization function, and dsp_hcc_exit is the DSP heterogeneous computing call exit function.

[0073] On the other hand, in the DSP programming, it consists of a heterogeneous computing kernel program encapsulated into a DSP library program API library (encapsulated into an API library, the user does not need to perceive the specific implementation details of the underlying code, and only needs to call the API to customize the computing program) and a DSP application program. The overall view of the module of the DSP heterogeneous computing kernel program technical implementation solution is as Figure 5As shown in the figure, the functional modules include a watchdog, a timer, an IPC interrupt module, a Cache module, a serial port module, etc.; the management modules include a calculation request data structure base address change management module, a calculation request data structure management module, a DSP multi-core management module, etc.; the calculation API modules include FFT (Fast Fourier Transform), Yolovx pre / post processing, and user-defined calculation modules, etc.).

[0074] The watchdog module is used to monitor whether the DSP program runs away and cooperate with the timer to complete the DSP program monitoring function; when the DSP program runs away and the timer fails to feed the watchdog on time, the watchdog will trigger the reset of the DSP core; after the corresponding DSP core is reset by the watchdog, the DSP core will send an IPC interrupt to the CPU to inform the CPU. The timer module is used to feed the watchdog regularly to prevent the watchdog from triggering the reset of the DSP core. When the DSP program executes abnormally, such as getting stuck, if the timer fails to feed the watchdog on time, then the watchdog will trigger the reset of the DSP core.

[0075] The IPC interrupt module is used to determine the IPC interrupt mode between the CPU and the DSP core and initiate inter-core communication through the hardware IPC module. The functional description of the DSP IPC interrupt can be as follows: (1) Each DSP core registers the following two interrupts to receive the IPC interrupt sent by the CPU: a. Interrupt INT-0: The handshake interrupt sent by the CPU; b. Interrupt INT-1: The API request interrupt sent by the CPU. (2) When the DSP sends an IPC interrupt to the CPU, there is only one interrupt source, and the interrupt type message is stored in a specific area of SMCmemory / GSM. The interrupt message can be defined as follows: a. bit-0: DSP handshake interrupt; b. bit-1: DSP API request completion normal interrupt; c. bit-2: DSP exception interrupt (such as program running away, etc.).

[0076] The Cache module is used for the maintenance of the DSP core Cache to ensure a series of operations for the effective and accurate operation of the DSP cache system. It provides operations such as opening the Cache to improve the memory access speed, Cache invalidation, Cache write-back, and closing the Cache. The serial port module is used for debugging, printing debugging, log, and other information and outputting it to the terminal. The calculation request data structure base address change management module is used to obtain the base address of the calculation request data structure. After successfully shaking hands with the CPU and receiving the IPC interrupt of the calculation task request sent by the CPU, it reads the base address of the calculation request data structure from the agreed SMC memory / GSM area. This base address is not a fixed address and is dynamically applied for by the CPU.

[0077] The calculation request data structure management module is used to read and write the corresponding member variables in the calculation request data structure. After successfully shaking hands with the CPU, it sets the status variable of the corresponding DSP core in the calculation request structure to the IDLE (idle) state; after the DSP completes the calculation, it sets the status of the corresponding DSP core to the IDLE state; it reads the calculation raw data, the storage addresses of the calculation result data, and the calculation parameters, etc. in the calculation request data structure.

[0078] The DSP multi-core management module is used for synchronization (using semaphores) of DSP multi-core calculations, etc. When the CPU calls the DSP for multi-core calculations, it uses semaphores to synchronize the multi-core, ensuring that the DSP multi-core calculations can be completed synchronously, and ensuring that before the DSP main core sends a calculation completion interrupt to the CPU, all DSP multi-cores have completed the calculations and filled in the interrupt type messages.

[0079] The calculation API module depends on the specific calculation. When the user calls the API provided by the above DSP library program, a calculation function can be added. Common calculations include FFT, LSTM (Long Short-Term Memory), matrix transpose, and Yolovx pre / post-processing, etc.

[0080] In one embodiment, the DSP library program API includes an API for initializing the DSP platform, an API for creating a calculation API, and an API for starting the DSP heterogeneous computing kernel.

[0081] Specifically, the API design provided by the DSP side for users to call can be as follows:

[0082] (1) The API for initializing the DSP platform. The functions of this API can include: a. Initializing the interrupt vector table, including the starting address of the interrupt service function for handling reset, the initialization of handling IPC inter-core interrupts, and the initialization of handling the watchdog timer feeding the dog. b. Obtaining the reset type, checking the reset type. If the reset is caused by the watchdog, then an interrupt will be sent to the CPU to notify that the DSP program has run away, and then it will be idle. c. Initializing the Cache, including initializing the L1 Cache (level 1 cache) and the L2 Cache (level 2 cache). d. Initializing the memory management, initializing the heap management in the space reserved by the CPU. e. Initializing the DSP status, initializing the DSP core status, including the core number (ID), the program entry address, the timer period, the reset source RSTTYPE_SRC, the DSP request pointer, and the interrupt flag, etc. f. Initializing the hardware timer. g. Initializing the watchdog. h. Initializing the semaphore synchronization / lock. The semaphore is used for multi-core DSP synchronization and as a lock. i. Cache write-back, that is, writing all DSP caches back to memory to ensure Cache consistency.

[0083] (2)Calculate the API to create an API, bind the starting address of the user's calculation API function and the callback function array members of the heterogeneous calculation kernel of the DSP, and use it for the requests sent by the CPU to call the corresponding API. The order of the function pointers needs to be consistent with the API call ID (identity) on the CPU side.

[0084] (3)Start the DSP heterogeneous calculation kernel API. Wait for the IPC interrupt sent by the CPU. When the IPC interrupt sent by the CPU is received, enter the IPC interrupt service program, check the IPC interrupt flag bit, set the flag bit in the DSP status management, and the IPC interrupt service program returns to the main function. In the main function, the DSP performs corresponding operations according to the flag bit, such as handshake (sending an IPC interrupt handshake to the CPU) and request (calling the corresponding API).

[0085] The user program calls the above DSP library program API, and the execution process is as Figure 6 shown. Among them, cl_init_platform is a function call, and its purpose is to perform initialization settings on the DSP side to ensure that both the hardware and software environments of the DSP platform are in a ready state. The called cl_create_program function is used to create the calculation API, and the cl_run_kernel function is used to run the DSP heterogeneous calculation kernel program.

[0086] The above interface call device for the CPU of the heterogeneous computer system to call the DSP encapsulates the calculation data, calculation parameters, etc. between the CPU and the DSP cores into a data structure that the user can modify by itself through a standardized function call interface and encapsulating the driver and underlying code details, and through structuring the calculation data and calculation parameters. This enables the user to simply use API calls on the CPU side and the DSP side and fill in the data structure of the calculation-related parameters for the DSP call according to the agreement, so as to achieve the CPU to call the DSP for heterogeneous calculation without having to understand the complex code and hardware details of the CPU and the DSP at the bottom layer, greatly reducing the difficulty of inter-core heterogeneous call development in the heterogeneous computer system and solving the problem of difficult and complex programming for the CPU to call the DSP core for multi-category calculations and calculation acceleration in the heterogeneous computer system. At the same time, it supports the CPU to separately call multiple calculation programs of the DSP. The DSP program only needs to be loaded once, and the user can flexibly add and delete calculation functions, greatly improving the reusability and flexibility of the CPU to call the DSP calculation program in the heterogeneous computer system.

[0087] It should be noted that the above CPU can be a single-core or multi-core CPU, and the DSP can be a single-core or multi-core DSP. In a heterogeneous computer system, the above design does not rely on third-party uncontrollable program libraries, realizes the CPU's call to the DSP calculation API level, calls the DSP library program API in the DSP program, registers it into different calculation APIs, and the CPU separately calls these calculation APIs. Compared with the prior art, it is more flexible; at the same time, the main programs for heterogeneous calls of the CPU and DSP are both encapsulated into API libraries, standardizing the function interfaces, greatly reducing the user programming difficulty. Users can completely ignore the mechanisms and underlying implementation details between heterogeneous calls, and only need to fill in the calculation request data structure and call the provided relevant APIs to realize the CPU's call to the DSP for multiple types of calculations.

[0088] In one embodiment, the data structure of the DSP multi-core state machine exists in the DSP function request structure, which can be read and written by both the CPU and the DSP. The state descriptions of the DSP multi-core state machine include: when the CPU calls the initialization function of the DSP heterogeneous calculation call library program, the CPU will initialize each DSP core state machine and set it to the UNKOWN state; when the CPU and the DSP perform a handshake, after the DSP receives the handshake interrupt, the DSP core will set the corresponding state machine to the idle state; after the CPU initiates a calculation API request, the CPU will change the DSP core state machine from the IDLE state to the running state; after the DSP completes the calculation, it will set the corresponding DSP core state machine to the idle state; when the DSP watchdog triggers a reset and the program runs again, the DSP will set the DSP core state machine to the reset state; after the CPU receives the DSP exception interrupt and initiates a handshake process, after the handshake is successful, the DSP core state machine becomes the idle state; if the three handshakes fail, then the CPU will set the DSP core state machine to the damaged state.

[0089] Specifically, the technical implementation of the DSP core state machine: Since there is no status register on the hardware DSP that can be directly read by the CPU, software needs to establish and maintain the state machine. The implementation of the DSP core state machine is jointly participated by the CPU and the DSP. The running state of each DSP core is saved in the calculation request data structure, as Figure 7 shown, both the CPU and the DSP can read and write, and the state descriptions are as follows:

[0090] When the CPU calls the dsp_hcc_init (DSP heterogeneous computing call initialization) function, the CPU will initialize (Init) each DSP core state machine and set it to the UNKOWN state. The CPU and the DSP perform a handshake. After the DSP receives the handshake interrupt, the DSP core sets the corresponding state machine to the IDLE state. After the CPU initiates a computing API request, the CPU changes the DSP core state machine from the IDLE state to the RUNNING state. After the DSP completes the computation, the DSP sets the corresponding DSP core state machine to the IDLE state. When the watchdog triggers a DSP reset and the program runs again, the DSP sets the DSP core state machine to the RESET state. After the CPU receives a DSP exception interrupt, it initiates a handshake process. After the handshake is successful, the DSP core state machine becomes the IDLE state; if the handshake fails three times, then the CPU sets the DSP core state machine to the BAD state.

[0091] In one embodiment, as Figure 8 shown, a method for an interface call of a CPU of a heterogeneous computer system to call a DSP may include the following processing steps S10 to S18:

[0092] S10, initialize the DSP heterogeneous computing call through the application program on the CPU side by calling the initialization DSP heterogeneous computing call API of the DSP heterogeneous computing call library;

[0093] S12, apply for a continuous memory block from the SMC memory / GSM through the API of the SMEM library by the application program to store the original computation data and the computation result data, and obtain the physical address of the continuous memory block;

[0094] S14, obtain an idle DSP through the API of the DSP heterogeneous computing call library by the application program, obtain an idle computing API request structure of the DSP, and write the computation data address and parameters into the computation API request structure according to the user's computation task;

[0095] S16, write back and invalidate the Cache through the API of the SMEM library by the application program, and initiate a computing API request to the DSP by calling the API of the DSP heterogeneous computing call library;

[0096] S18, wait for the DSP to return the computation completion, and then release the applied continuous memory block through the API of the SMEM library by the application program;

[0097] S20, exit the DSP heterogeneous computing call through the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library by the application program;

[0098] Among them, the process of the DSP application program of the DSP calling the API of the DSP library program includes initializing the DSP platform, creating a computing API function, running the DSP heterogeneous computing kernel program, setting the status of the corresponding DSP core to the idle state, determining whether the CPU initiates an IPC interrupt. If so, after parsing the IPC interrupt flag, sending an IPC interrupt to the CPU and executing the corresponding user-defined computing API according to the ID of the computing API request initiated by the CPU. After the DSP completes the calculation, the status of the corresponding DSP core is set to the idle state.

[0099] The above interface calling method for the CPU of the heterogeneous computer system to call the DSP registers different computing APIs by calling the API of the DSP library program in the DSP program according to the multi-category calculations that the user needs to perform. The CPU calls these computing APIs separately. In this way, in the heterogeneous call design, the main programs of the CPU and the DSP are both encapsulated into API libraries, greatly reducing the programming difficulty for users. Users do not need to care about the mechanism and underlying implementation details between heterogeneous calls. They only need to call the provided relevant APIs and data structures to implement the CPU calling the DSP for multi-category calculations, thus realizing the CPU calling the DSP at the computing API level in the heterogeneous computer system.

[0100] Compared with the prior art, it is more flexible. At the same time, the main programs of the heterogeneous call CPU and DSP are both encapsulated into API libraries, standardizing the function interfaces, greatly reducing the programming difficulty for users. Users do not need to care about the mechanism and underlying implementation details between heterogeneous calls. They only need to fill in the computing request data structure and then call the provided relevant APIs to implement the CPU calling the DSP for multi-category calculations, and it does not depend on third-party uncontrollable program libraries.

[0101] For the specific limitations and explanatory notes on the interface calling method for the CPU of the heterogeneous computer system to call the DSP, reference can be made to the corresponding limitations and explanatory notes of the interface calling device for the CPU of the heterogeneous computer system in the above text, which will not be elaborated here.

[0102] It should be understood that although Figure 8 the steps in Figure 8 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover Figure 8 at least a part of the steps in Figure 8 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus dynamic random access memory (Rambus DRAM, abbreviated as RDRAM), and interface dynamic random access memory (DRDRAM), etc.

[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0105] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the protection scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, which all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. An interface calling device for a heterogeneous computer system CPU to call a DSP, characterized in that: The invention comprises a CPU and a DSP, wherein the CPU and the DSP interact with each other for heterogeneous computing by using SMC memory / GSM, shared memory and IPC inter-core interrupts, the software stack of the CPU comprises an application, an API library and a driver, the application calls an API provided by the API library to implement calling the DSP for heterogeneous computing, the API library comprises a DSP heterogeneous computing call library and a SMEM library, the driver comprises an SMEM driver and a DSP heterogeneous computing call driver, and the heterogeneous computing kernel program of the DSP is encapsulated into a DSP library program API; The application of the CPU calls the API provided by the API library to implement the process of calling the DSP to perform heterogeneous computing, including calling the initialization DSP heterogeneous computing call API of the DSP heterogeneous computing call library to initialize the DSP heterogeneous computing call, calling the API of the SMEM library to apply for a continuous memory block from the SMC memory / GSM for storing calculation original data and calculation result data, and obtaining the physical address of the continuous memory block, calling the API of the DSP heterogeneous computing call library to obtain an idle DSP, obtain the idle DSP calculation API request structure and write the calculation data address and parameters into the calculation API request structure according to the user calculation task, calling the API of the SMEM library to write back and invalidate the cache, calling the API of the DSP heterogeneous computing call library to initiate a calculation API request to the DSP, waiting for the DSP to return the calculation completion, calling the API of the SMEM library to release the applied continuous memory block, and calling the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library to exit the DSP heterogeneous computing call; Among them, the process of the DSP application of the DSP calling the DSP library program API execution includes initializing the DSP platform, creating a calculation API function, running the DSP heterogeneous computing kernel program, setting the state of the corresponding DSP core to an idle state, judging whether the CPU initiates an IPC interrupt, and if so, parsing the IPC interrupt flag and sending an IPC interrupt to the CPU and executing the corresponding user-defined calculation API according to the ID of the CPU-initiated calculation API request, and after the DSP completes the calculation, setting the state of the corresponding DSP core to an idle state.

2. The interface calling device for a heterogeneous computer system CPU to call a DSP according to claim 1, characterized in that: The application program calls the API dynamic application of the SMEM library to obtain a section of space in the SMC memory / GSM storage space and use it to store the calculation request data structure. The last 4KB space of the SMC memory / GSM is used to store the base address and size of the DSP calculation request data structure variable and the message type of the IPC interrupt sent by the DSP to the CPU. The IPC interrupt sources sent by the CPU to the DSP include INT-0 and INT-1, INT-0 indicates a handshake interrupt sent by the CPU, and INT-1 indicates an API request interrupt sent by the CPU. The message type of the IPC interrupt includes bit-0, bit-1 and bit-2, bit-0 indicates a DSP handshake interrupt, bit-1 indicates a normal interrupt of the DSP calculation API request completion, and bit-2 indicates an abnormal DSP interrupt.

3. The interface calling device for a heterogeneous computer system CPU to call a DSP according to claim 2, characterized in that: The API library of the CPU software stack provides APIs including initializing DSP heterogeneous computing call API, unloading DSP heterogeneous computing call API, idle acquisition API, core status acquisition API, idle structure acquisition API, CPU calling DSP to perform heterogeneous computing API, waiting for DSP heterogeneous computing API request completion API, SMEM initialization API, SMEM memory application API, SMEM memory release API, SMEM physical address acquisition API, invalidation Cache API, and write-back and invalidation Cache API; The DSP library program API includes initializing the DSP platform API, creating a computing API, and starting a DSP heterogeneous computing kernel API.

4. The interface calling device for a heterogeneous computer system CPU to call a DSP according to any one of claims 1 to 3, characterized in that: Before the CPU initiates a calculation API request, the calculation API ID to be called, the address of the original calculation data, the address of the calculation result data and the calculation parameters are filled in the corresponding member variables in the calculation request data structure; the calculation request data structure is provided with a member variable of an enumeration type for storing the calculation API ID separately, and the API ID setting of the CPU is consistent with the API ID setting of the DSP; In the heterogeneous computing kernel program of the DSP, the DSP library program API is called to save the function pointers that need to be calculated in an array and registered as different computing APIs. The subscript of the computing function array is the API ID that the CPU calls the DSP for calculation; 0 is excluded in the subscript of the computing function array to prevent invalid calls.

5. The interface calling device for a heterogeneous computer system CPU to call a DSP according to claim 4, characterized in that: The DSP heterogeneous computing call library implements sending an IPC interrupt to the DSP core, setting the base address and size of the storage area of ​​the calculation request data structure, calling the API provided by the SMEM library to apply for a section of memory in the SMC memory / GSM for storing the calculation request data structure, encapsulating DSP heterogeneous computing related functions into a standard API for the application program to call, implementing DSP multi-core state machine maintenance, and receiving asynchronous signals sent by the DSP heterogeneous computing call driver by calling the standard ioctl interface provided by the DSP heterogeneous computing call driver; The DSP heterogeneous computing call driver obtains IPC register resources, SMC memory / GSM reserved space resources and IPC interrupt numbers from the device tree, operates the underlying hardware, registers interrupt service routines, receives IPC interrupts sent by the DSP, sends asynchronous signals to the DSP heterogeneous computing call library, and provides ioctl interface calls; The SMEM library is used to provide corresponding APIs for implementing application, release, virtual-to-physical address conversion, cache refresh and cache invalidation of continuous memory, and an ioctl interface for calling the SMEM driver; The SMEM driver serves as a management driver for shared continuous memory blocks, and is used to obtain and load driver module parameters from a device tree, implement configurable continuous physical memory area management, and provide corresponding ioctl interfaces for application, release, virtual-to-real physical address conversion, cache refresh, and cache invalidation of continuous physical memory areas for the SMEM library to call; wherein, continuous physical memory area management includes management of SMC memory, GSM, and shared memory.

6. The interface calling device for a heterogeneous computer system CPU to call a DSP according to claim 5, characterized in that: The data structure of the DSP multi-core state machine is stored in the DSP function request structure, which is readable and writable by both the CPU and the DSP. The description of each state of the DSP multi-core state machine includes: When the CPU calls the DSP heterogeneous computing call library program initialization function, the CPU will initialize each DSP core state machine and set it to the UNKOWN state; The CPU and the DSP perform handshake, and after the DSP receives the handshake interrupt, the DSP core sets the corresponding state machine to an idle state; After the CPU initiates a calculation API request, the CPU changes the DSP core state machine from an IDLE state to a running state; After the DSP completes the calculation, the corresponding DSP core state machine is set to an idle state; The DSP watchdog triggers a reset, and after re-running the program, the DSP sets the DSP core state machine to a reset state; After receiving the DSP abnormal interrupt, the CPU initiates a handshake process. After the handshake succeeds, the DSP core state machine changes to an idle state; if the handshake fails three times, the CPU sets the DSP core state machine to a damaged state.

7. A method for calling a DSP interface by a CPU in a heterogeneous computer system, characterized in that: Includes steps: The application on the CPU side calls the DSP heterogeneous computing call library to initialize the DSP heterogeneous computing call API to initialize the DSP heterogeneous computing call; The application program calls the API of the SMEM library to apply for a continuous memory block from the SMC memory / GSM for storing the original calculation data and the calculation result data, and obtains the physical address of the continuous memory block; The application program calls the API of the DSP heterogeneous computing call library to obtain an idle DSP, obtain an idle DSP computing API request structure, and write the computing data address and parameters into the computing API request structure according to the user computing task; The application program calls the API of the SMEM library to write back and invalidate the cache, and calls the API of the DSP heterogeneous computing call library to initiate a computing API request to the DSP; After waiting for DSP to return the calculation, the application program calls the API of the SMEM library to release the applied continuous memory block; Exit the DSP heterogeneous computing call by calling the DSP heterogeneous computing call exit API of the DSP heterogeneous computing call library through the application program; Among them, the process of the DSP application of the DSP calling the DSP library program API execution includes initializing the DSP platform, creating a calculation API function, running the DSP heterogeneous computing kernel program, setting the state of the corresponding DSP core to an idle state, judging whether the CPU initiates an IPC interrupt, and if so, parsing the IPC interrupt flag and sending an IPC interrupt to the CPU and executing the corresponding user-defined calculation API according to the ID of the CPU-initiated calculation API request, and after the DSP completes the calculation, setting the state of the corresponding DSP core to an idle state.

8. The interface calling method for a heterogeneous computer system CPU to call a DSP according to claim 7, characterized in that: The application program calls the API dynamic application of the SMEM library to obtain a section of space in the SMC memory / GSM storage space and use it to store the calculation request data structure. The last 4KB space of the SMC memory / GSM is used to store the base address and size of the DSP calculation request data structure variable and the message type of the IPC interrupt sent by the DSP to the CPU. The IPC interrupt sources sent by the CPU to the DSP include INT-0 and INT-1, INT-0 indicates a handshake interrupt sent by the CPU, and INT-1 indicates an API request interrupt sent by the CPU. The message type of the IPC interrupt includes bit-0, bit-1 and bit-2, bit-0 indicates a DSP handshake interrupt, bit-1 indicates a normal interrupt of the DSP calculation API request completion, and bit-2 indicates an abnormal DSP interrupt.

9. The interface calling method for a heterogeneous computer system CPU to call a DSP according to claim 8, characterized in that: The API library of the CPU software stack provides APIs including initializing DSP heterogeneous computing call API, unloading DSP heterogeneous computing call API, idle acquisition API, core status acquisition API, idle structure acquisition API, CPU calling DSP to perform heterogeneous computing API, waiting for DSP heterogeneous computing API request completion API, SMEM initialization API, SMEM memory application API, SMEM memory release API, SMEM physical address acquisition API, invalidation Cache API, and write-back and invalidation Cache API; The DSP library program API includes initializing the DSP platform API, creating a computing API, and starting a DSP heterogeneous computing kernel API.

10. The interface calling method for a heterogeneous computer system CPU to call a DSP according to claim 7, characterized in that: The data structure of the DSP multi-core state machine is stored in the DSP function request structure, which is readable and writable by both the CPU and the DSP. The description of each state of the DSP multi-core state machine includes: When the CPU calls the DSP heterogeneous computing call library program initialization function, the CPU will initialize each DSP core state machine and set it to the UNKOWN state; The CPU and the DSP perform handshake, and after the DSP receives the handshake interrupt, the DSP core sets the corresponding state machine to an idle state; After the CPU initiates a calculation API request, the CPU changes the DSP core state machine from an IDLE state to a running state; After the DSP completes the calculation, the corresponding DSP core state machine is set to an idle state; The DSP watchdog triggers a reset, and after re-running the program, the DSP sets the DSP core state machine to a reset state; After receiving the abnormal interrupt from DSP, the CPU initiates the handshake process. After the handshake succeeds, the DSP core state machine changes to the idle state. If the handshake fails three times, the CPU sets the DSP core state machine to the damaged state.

Citation Information

Patent Citations

  • Novel communication method between multi-core processors

    CN112069124A

  • Processing method and device based on heterogeneous computing framework, equipment and medium

    CN116360971A