Parallel computing framework based on software definition and implementation method thereof

By adopting a software-defined parallel computing architecture on the embedded heterogeneous system platform, dynamic self-organization reconstruction of computing resources and cross-platform deployment of neural network computing are realized, solving the problems of low parallel computing efficiency and poor computing stability.

CN120104306APending Publication Date: 2025-06-06EAST CHINA INST OF COMPUTING TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510040752.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The parallel computing efficiency under the embedded heterogeneous system platform is low, the computing framework lacks standardization, the deployment of neural network frameworks is difficult, and the computing stability of different neural network models is poor.

Method used

Using a software-defined parallel computing architecture, through the "data-control-capability" modular design, computing resources have the ability to dynamic self-organize and reconstruct, and relevant operator functions are optimized.

Benefits of technology

It improves the parallel computing efficiency of the embedded heterogeneous system platform, realizes dynamic self-organization reconstruction of computing resources, supports cross-platform deployment of neural network computing, and improves the computing stability of different neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104306A_ABST
    Figure CN120104306A_ABST
Patent Text Reader

Abstract

The invention relates to a parallel computing framework based on software definition and an implementation method thereof, the parallel computing framework comprises an operating system domain, a system service domain and a unified interface domain, and the operating system domain comprises equipment drive, task scheduling, signal management, memory management and various protocol stacks; the system service domain comprises a data plane, a control plane and a capability plane, and the data plane is responsible for loading, forwarding, mapping and exporting all data; the control plane is responsible for calculation task preprocessing, load separation, scheduling distribution, dynamic reconstruction, calculation scheduling and task estimation functions and supports expansion customization; the capability plane is mainly responsible for computing node capability registration, capability reconstruction, operator management, capability management and failure detection functions in the system; the unified interface domain has the functions of distributed calculation, model training, model reasoning and resource management. The problem of parallel computing efficiency under a resource-limited embedded heterogeneous system platform is solved, so that computing resources have dynamic self-organizing reconstruction capability according to computing requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embedded heterogeneous platforms, and in particular to a software-defined parallel computing framework and an implementation method thereof, especially to an embedded heterogeneous platform with limited resources. Background Art

[0002] In the past few years, with the continuous innovation and evolution of artificial intelligence technology, China has become a leading country in the world in AI technology innovation and scenario application. Related technologies have been widely used in autonomous driving, smart medical care, portrait recognition and other fields. The promotion and application of artificial intelligence has promoted the explosive growth of computing power demand. AI computing demand is doubling every four months, which is more efficient than Moore's Law. In the embedded field, embedded processors as computing power infrastructure have been difficult to meet the computing power requirements of artificial intelligence on their own. Heterogeneous parallel computing systems based on central processing units (CPU) + field programmable gate arrays (FPGA) + extensible processor units (XPU) are expected to become the key to solving this problem. In embedded systems, there is currently no computing framework that can dynamically self-organize and reconstruct computing power to support neural network computing; the technical problems to be solved by the present invention are reflected in the following points:

[0003] 1) Parallel computing efficiency issues on embedded heterogeneous system platforms;

[0004] 2) Standardization of parallel computing frameworks on embedded platforms;

[0005] 3) Issues with deploying neural network frameworks on embedded platforms;

[0006] 4) Computational stability issues of different neural network models. Summary of the invention

[0007] Aiming at the problem of parallel computing efficiency in resource-constrained embedded heterogeneous system platforms, a software-defined parallel computing framework and its implementation method are proposed. The innovative "data-control-capability" modular design enables computing resources to have the ability to dynamically self-organize and reconstruct according to computing needs, while optimizing related operator functions, which has certain practicality and reference value.

[0008] The technical solution of the present invention is:

[0009] A parallel computing architecture based on an embedded platform. The parallel computing architecture SDPCA includes an operating system domain, a system service domain and a unified interface domain. The operating system domain is the basic software support for the entire architecture and is also the software-based interface for the underlying hardware accelerator. The hardware capabilities are registered to the upper-level system service domain through device drivers, including device drivers, task scheduling, signal management, memory management and various protocol stacks; the system service domain is the core of SDPCA, including three planes: the data plane, the control plane and the capability plane. The data plane is responsible for loading, forwarding, mapping and exporting all data, and is the main body of rapid data exchange; the control plane is responsible for computing task preprocessing, load separation, scheduling allocation, dynamic reconstruction, computing scheduling, and task estimation functions, and supports extended customization; the capability plane is mainly responsible for computing node capability registration, capability reconstruction, operator management, capability management, and failure detection functions within the system; the unified interface domain includes distributed computing, model training, model reasoning, and resource management functions, and provides a cross-platform programming interface for upper-level applications.

[0010] Furthermore, the underlying architecture includes a series of general-purpose processors CPU, graphics processor GPU, signal processor DSP, programmable logic device FPGA, and neural network processor NPU.

[0011] Furthermore, the unified interface domain is compatible with and supports NCNN and MNN frameworks.

[0012] A method for implementing a parallel computing architecture based on an embedded platform, based on the implementation of the parallel computing architecture based on the embedded platform as described above, the parallel computing architecture SDPCA software defines computing resources based on plane features, maintains a dynamically reconfigurable virtual computing resource pool, and different application components can be transplanted and reused between architectures of the same standard;

[0013] Control plane: This method builds the control plane of the SDPCA framework based on the software definition concept, realizes the dynamic self-organizing computing function of the host and device based on computing needs, and reduces the impact of the short board factor on the system computing power; sets the task estimation module, capability reconstruction module and capability database in the control plane. The task estimation module includes three parts: computing feature acquisition, accelerator feature extraction and task evaluation, namely task analysis and estimation; workload cutting, simulation sampling operation and cache consistency processing modules are introduced in this control plane; the execution of threads in the hardware accelerator is dynamically analyzed and scheduled in real time through the precise timing model to meet the requirements of the computing time limit; when the hardware accelerator performs parallel acceleration, the runtime state dynamically analyzes the load state of each accelerator according to the workload of each thread in the hardware accelerator and the built-in scheduling detection point, and then dynamically allocates the work to the relatively idle thread for execution. The computing scheduler integrates static scheduling and dynamic scheduling execution. In a program containing Fourier transform, image processing and filtering processing, the task data is deployed to different virtual memory spaces through static scheduling, and the dynamic scheduling inserts scheduling detection points during the program execution process to control the thread execution process of the computing task in the accelerator in real time;

[0014] Data plane: This data plane is equipped with a minimalist protocol distributed soft bus, which virtualizes a series of computing memory blocks on the host side through the virtual memory management unit on the host side. These memory blocks reside in the physical memory of the host side. The virtual memory management unit is responsible for the unified addressing between the host side and the device side, and directly calls the high-speed interface at the physical layer for memory data transmission, forming a shared memory space between the host side and the device side, and always maintains the consistency of the host memory, accelerator device memory and cache data during the entire system operation;

[0015] Capability plane: The capability plane realizes the dynamic construction of device capabilities and operator capabilities, and abstractly describes resource computing capabilities from the system level. Device capabilities are a software expression of hardware resources and are registered to the system through the interface provided by the framework. This capability plane divides devices into high-speed general-purpose processors, digital signal processors, and field programmable gate arrays. Operators are the basis for building the capability plane. In neural network computing, they are represented as kernel functions and are implemented in specific programming languages ​​in heterogeneous programming environments. When kernel functions run on heterogeneous devices, the resources of heterogeneous devices are processed in blocks so that the calculations can be distributed among multiple computing units for parallel execution.

[0016] Furthermore, in the control plane, the maximum computing capacity of the parallel computing architecture should be represented by a set of related factors. The computing power of the i-th device is D[i]cap, the scheduling strategy capacity Scap, and the memory exchange capacity Mcap. The system computing capacity Ccap is expressed as shown in Formula 1:

[0017]

[0018] Combine the above features with parallel computing to maximize the computing power of the system.

[0019] Furthermore, the computing characteristics are obtained by inputting the static code of the computing task, which is determined by the front-end compiler through static code analysis. The computing task characteristic data set collection and conversion of the computing task is quickly completed when the task is loaded. The task analysis and estimation module divides the characterized computing task into accelerator device subtasks, obtains the currently available accelerator capability by searching the device characteristics in the capability database, and matches it with the computing characteristics of this time, and evaluates the running time of the computing task on the corresponding device. When the accelerator device capability in the capability database does not meet the operator function or execution time requirements of the task, the capability reconstruction module is notified to reconstruct the capability of the accelerator device and update the capability description information to the capability database. The force reconstruction module dynamically deploys operators to the corresponding accelerator devices based on the characteristics of the computing tasks using the host-side and device-side interaction interfaces. The host-side and device-side interaction interface definitions include SDPCA_GetPlatformInfo, SDPCA_GetDeviceCap, SDPCA_RegDeviceCap, SDPCA_RebuildDeviceCap, SDPCA_CreateCmdQue, SDPCA_CreateContext, SDPCA_CreateKernel, SDPCA_SetKernelArg, and SDPCA_TaskEmqueue.

[0020] Furthermore, a memory management model is set up in the data plane. The concept of OpenCL memory variables is used in the host-side memory model. There are three memory types: array cache, image object, and pipeline object. The data in the array cache is continuous in the address space. The adjacent data in the image array of the image object is not guaranteed to be stored continuously in the memory. The pipeline object is a data element queue that follows the first-in-first-out (FIFO) principle. The device side defines global memory, constant memory, local memory, and private memory. The global memory is visible to each work item in the execution kernel and is the memory used for transmission between the host and device sides. The constant memory is accessed by all work items and is used to pass in configuration parameters. The local memory realizes the sharing of work items within the same work group and usually uses the on-chip memory of the device side. The private memory is only accessed within the work item.

[0021] Furthermore, in the capability plane, the GPP resources in the system are defined to consist of N GPP devices interconnected by a specific topology, which are modeled as a GPP device list of length N. Each GPP device dynamically maintains M matching characteristics and A allocation characteristics, which are represented by P M and PA Expression,the GPP device modeling in the system is shown in Equation 2;

[0022]

[0023] For the nth GPP device list, n is one of 1, 2, ..., N, R M (GPP n ) represents GPP n The matching capability of the device refers to the device attribute characteristics in the computing environment that will constrain the computing process but cannot be quantitatively calculated, including processor model, operating system, and system clock; A (GPP n ) represents GPP n The allocation capacity of the device is the resource capacity attribute in the device that has a quantitative impact on the computing process, including the number of cores, main frequency, memory size, operator capacity description table, GPP n The theoretical value of the equipment capacity is given by R A (GPP n ) decides that when a new computing device GPP n After registering to the system, the capability plane automatically completes the software description of the device, builds typical matching characteristics and allocation characteristics for it, and dynamically maintains the R of the device throughout its life cycle. A (GPP n );

[0024] This method adapts the relevant kernel functions in the SDPCA framework and optimizes performance based on SIMD technology.

[0025] The beneficial effects of the present invention are:

[0026] 1) Solve the problem of parallel computing efficiency in resource-constrained embedded heterogeneous system platforms;

[0027] 2) Based on the software-defined distributed computing framework, computing resources have the ability to dynamically self-organize and reconfigure according to computing needs;

[0028] 3) Cross-platform deployment of neural network reasoning framework on embedded real-time system platforms;

[0029] 4) The computational stability issues of different neural network models under embedded platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 The overall design diagram of parallel computing defined by the SDPCA software of the present invention;

[0031] Figure 2 To estimate the tasks and reconstruct the capabilities of the invention;

[0032] Figure 3 It is a schematic diagram of the execution of the computing scheduler of the present invention;

[0033] Figure 4 This is a data plane memory management model diagram of the present invention;

[0034] Figure 5 This is a diagram describing the system capabilities of the present invention. DETAILED DESCRIPTION

[0035] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0036] A parallel computing architecture based on an embedded platform (named: SDPCA (Software Defined Parallel Computing Architecture)), involving hardware platforms, operating systems, parallel computing frameworks, etc. This computing architecture is based on the ideas of SCA and SDN, integrating the ideas of modularization, componentization, and resource virtualization. The computing architecture is divided into the operating system domain, system service domain, and unified interface domain. In the system service domain at the core of the architecture, three planes of data, control, and capability are constructed based on the modular idea. Computing resources are software-defined based on plane characteristics. The architecture maintains a dynamically reconfigurable virtual computing resource pool. Different application components can be transplanted and reused between architectures of the same standard. The overall architecture is as follows: Figure 1 shown.

[0037] The bottom layer of the SDPCA architecture includes a series of general-purpose processors (CPUs), graphics processors (GPUs), signal processors (DSPs), programmable logic devices (FPGAs), neural network processors (NPUs), etc. These processors are the basic physical units that constitute high-performance computing. The operating system domain is mainly the basic software support for the entire architecture, and is also the software interface of the underlying hardware accelerator. Through the device driver, the hardware capabilities are registered to the upper-layer system service domain. This domain mainly includes device drivers, task scheduling, signal management, memory management, and various protocol stacks. The system service domain is the core of SDPCA and is divided into three planes: "data-control-capability". The data plane is responsible for loading, forwarding, mapping, and exporting all data, and is the main body of fast data exchange. The control plane is responsible for computing task preprocessing, load separation, scheduling and allocation, dynamic reconstruction, computing scheduling, task estimation, and other functions, and supports expansion and customization. The capability plane is mainly responsible for computing node capability registration, capability reconstruction, operator management, capability management, failure detection, and other functions within the system. The unified interface domain includes distributed computing, model training, model reasoning, resource management and other functions. It provides a cross-platform programming interface for upper-level applications, and is compatible with frameworks such as NCNN and MNN to achieve cross-platform compatibility of SDPCA.

[0038] 1. Control Plane

[0039] Logical control is the brain for building the software-defined computing capabilities of the entire system. An excellent computing model should have good scalability, reliability, and flexibility, be applicable to applications in different scenarios, and have software-defined capabilities. In traditional parallel computing systems, the platform model is divided into the host side and the device side. The host side is undertaken by a general-purpose processor CPU, which sends commands to the device side and receives feedback information after processing by the device. It is responsible for managing the task queue and arranging new tasks to enter the queue. The device side is composed of a series of hardware accelerators such as GPU, NPU, FPGA, etc. It is the actual executor of high-performance computing, completes the calculation and feeds the results back to the host system.

[0040] The main factors affecting the overall computing power of a parallel computing system include host scheduling capability, device computing power, number of devices, and high-speed bus data communication capability. The influencing factors of host scheduling capability include the number of tasks, scheduling strategy, data block size, and data regularity. Theoretically, when none of the influencing factors is a bottleneck, the overall computing power of the system is positively correlated with the influencing factors.

[0041] The computing power of a parallel system is related to multiple related factors. Increasing the number of computing devices can significantly improve the computing power of the system. When the number of devices is constant, increasing the number of parallel tasks on the host will bring about the overhead of task scheduling, which will reduce the computing power of the system. Therefore, the maximum computing power value of a parallel computing system should be represented by a set of related factors. The computing power of the i-th device is D[i]cap, the scheduling strategy capacity Scap, and the memory exchange capacity Mcap. The expression of the system computing power Ccap is shown in Formula 1:

[0042]

[0043] In combination with the above characteristics of parallel computing, in order to maximize the computing power of the system, this solution builds the control plane of the SDPCA framework based on the software-defined concept, realizes the dynamic self-organizing computing function of the host and device based on computing needs, and reduces the impact of short board factors on the computing power of the system. In the control plane, a task estimation module, a capability reconstruction module and a capability database are set. The task estimation module includes three parts: computing feature acquisition, accelerator feature extraction and task evaluation (that is, task analysis and estimation). Figure 2 Restructure module relationships for task estimation modules and capabilities.

[0044] The computing features are mainly obtained by inputting the static code of the computing task and are determined by the front-end compiler through static code analysis. For example, for a neural network computing task, the convolution, pooling, normalization, sampling and other major categories in the computing task can be quickly completed when the task is loaded, and the input data block attributes and the number of calls of related operators in the major categories are instantiated as computing task feature data sets. The task analysis and estimation module divides the characterized computing task into accelerator device subtasks, obtains the currently available accelerator capabilities by searching the device features in the capability database, matches them with the computing features of this time, and evaluates the running time of the computing task on the corresponding device. When the accelerator device capabilities in the capability database do not meet the operator functions or execution time requirements of the task, the capability reconstruction module is notified to reconstruct the capabilities of the accelerator device and update the capability description information to the capability database. The capability reconstruction module uses the host-side and device-side interaction interfaces to dynamically deploy operators to the corresponding accelerator devices according to the characteristics of the computing tasks. The host-side and device-side interaction interfaces are defined as shown in Table 1, including SDPCA_GetPlatformInfo, SDPCA_GetDeviceCap, SDPCA_RegDeviceCap, SDPCA_RebuildDeviceCap, SDPCA_CreateCmdQue, SDPCA_CreateContext, SDPCA_CreateKernel, SDPCA_SetKernelArg, and SDPCA_TaskEmqueue.

[0045] Table 1 SDPCA interface function definition

[0046] Table 1SDPCA Interface functions definition

[0047]

[0048] This control plane introduces modules such as workload cutting, simulation sampling operation, and cache consistency processing. The execution of threads in the hardware accelerator is dynamically analyzed and scheduled in real time through an accurate timing model to meet the requirements of the computing time limit. When the hardware accelerator performs parallel acceleration, the runtime state dynamically analyzes the load state of each accelerator according to the workload of each thread in the hardware accelerator and the built-in scheduling detection point, and then dynamically allocates the work to relatively idle threads for execution. Figure 3 A schematic diagram of the execution of a computing scheduler that integrates static scheduling and dynamic scheduling is given. In a program that includes Fourier transform, image processing, and filtering processing, task data is deployed to different virtual memory spaces through static scheduling, while dynamic scheduling inserts scheduling checkpoints during program execution to control the thread execution process of computing tasks in the accelerator in real time.

[0049] 2. Data plane

[0050] The data plane mainly solves the problem of data transmission between the host and the accelerator device. The current computing framework still widely has the overhead of redundant data replication, which is mainly reflected in the data transmission and framework data format conversion process. This framework sets up a minimalist protocol distributed soft bus, which virtualizes a series of computing memory blocks on the host side through the virtual memory management unit on the host side. These memory blocks reside in the physical memory of the host side. The virtual memory management unit is responsible for the unified addressing between the host side and the device side, abandoning the complex protocol interaction processes such as the presentation layer, session layer, and transport layer, making the performance infinitely close to the hard bus capability, and directly calling high-speed interfaces such as RDMA and RapidIO at the physical layer for memory data transmission, forming a shared memory space between the host side and the device side, and always maintaining the consistency of the host memory, accelerator device memory and cache data during the operation of the entire system. Since the CPU is bypassed for data movement, the throughput and transmission delay of data interaction are improved. Figure 4The data platform memory management model is described. In order to maintain compatibility with other computing frameworks, the OpenCL memory variable concept is used in the host-side memory model. There are three memory types: array buffer, image object, and pipe object. The array buffer is similar to the space allocated by malloc in C language, and the data is continuous in the address space. Image objects are different from ordinary arrays. The adjacent data of image arrays are not guaranteed to be stored continuously in memory. Its design purpose is to give full play to the advantages of hardware in spatial locality and use the device hardware acceleration capabilities to achieve efficient processing of image data. Pipe objects are a data element queue that follows the first-in-first-out (FIFO) principle to support specific types of parallel computing tasks. The device side defines global memory, constant memory, local memory, and private memory. Global memory is visible to each work item in the execution kernel and is the memory used for transmission between the host and device sides. Constant memory is accessed by all work items and is used to pass in configuration parameters. Local memory is shared by work items in the same work group and usually uses the on-chip memory of the device side. Private memory is only accessed within the work item.

[0051] 3. Capability plane

[0052] The capability plane mainly realizes the dynamic construction of device capabilities and operator capabilities, and abstractly describes resource computing capabilities from the system level. Device capabilities are a software expression of hardware resources and are registered to the system through the interface provided by the framework. This capability plane follows the European Secure Software Radio (ESSOR) specification to divide devices into high-speed general-purpose processors (General-Purpose Processors, GPP), digital signal processors (Digital Signal Processors, DSP) and field programmable gate arrays (Field Programmable Gate Array, FPGA). The GPP resources in the system are defined to consist of N GPP devices interconnected by a specific topology, modeled as a GPP device linked list of length N. Each GPP device dynamically maintains M matching characteristics and A allocation characteristics, respectively represented by P M and P A Expression, the GPP device modeling in the system is shown in Equation 2.

[0053]

[0054] Among them, for the nth GPP device list GPP n , n is one of 1, 2, ..., N, R M (GPP n ) represents GPP nThe matching capability of the device refers to the device attribute characteristics in the computing environment that will constrain the computing process but cannot be quantitatively calculated, mainly including processor model, operating system, system clock, etc. A (GPP n ) represents GPP n The allocation capacity of a device is the resource capacity attribute in the device that has a quantitative impact on the computing process, mainly including the number of cores, main frequency, memory size, operator capacity description table, etc. GPP n The theoretical value of the equipment capacity is given by R A (GPP n ) decides that when a new computing device GPP n After registering to the system, the capability plane automatically completes the software description of the device, builds typical matching characteristics and allocation characteristics for it, and dynamically maintains the R of the device throughout its life cycle. A (GPP n ), P of computing devices in the system M and P A Description Figure 5 shown.

[0055] Operators are the basis for building the capability plane. In neural network computing, they are represented as kernel functions and implemented by specific programming languages ​​in heterogeneous programming environments. When kernel functions run on heterogeneous devices, the resources of heterogeneous devices are processed in blocks so that the calculations can be distributed to multiple computing units for parallel execution.

[0056] This solution adapts the relevant kernel functions in the SDPCA framework and optimizes performance based on SIMD technology. SIMD can realize single instruction multi-group data processing. Since multiple parallel processor units share instructions and decoding logic, the SIMD structure can achieve a higher performance-to-power ratio. The solution finally completed the implementation of nearly 100 commonly used kernel functions, including string segmentation, string concatenation, Fourier transform, matrix inversion, etc. Table 2 shows the performance test data of typical kernel functions before and after SIMD optimization. The computing performance of most operators has increased by more than 2 times after SIMD optimization, and the performance of some operators has increased by 17 times.

[0057] Table 1 Typical operator SIMD optimization performance Unit: us

[0058] Table 1SIMD Performance of Typical Operators

[0059]

[0060] The present invention relates to a parallel computing architecture based on software definition. The computing architecture innovatively adopts the "data-control-capability" modular design, so that the computing resources have the ability to dynamically self-organize and reconstruct according to computing needs. At the same time, relevant operator functions are designed to enhance the GPU parallel computing capability, which is suitable for resource-constrained embedded platforms. The C language-based implementation ensures that the design method has the ability to cross multiple operating system platforms, which is convenient for rapid porting and use on various hardware platforms, and has strong practical and promotion value.

[0061] Implementation example: Development based on embedded operating system

[0062] This design method is mainly based on the development of embedded operating systems, and the content involves operating systems, IDE development environments, SDPCA control modules, SDPCA data modules, SDPCA capability modules, operator modules, configuration resource components, etc. The IDE development environment mainly performs operating system mirroring, module library compilation, user software compilation, resource component development and management, and ultimately realizes a high-performance parallel computing method that can effectively implement resource-constrained embedded platforms.

[0063] This technology takes software-defined thinking as the main line, provides a lightweight modular parallel computing solution for embedded heterogeneous platforms, and proposes a parallel computing architecture based on software definition.

[0064] The solution has the following features:

[0065] 1. Can be widely applied to most embedded system hardware platforms;

[0066] 2. It can provide the underlying computing base for neural network reasoning operations;

[0067] 3. Based on software-defined lightweight module design, computing resources have the ability to dynamically self-organize and reconfigure according to computing needs;

[0068] 4. High computational efficiency and strong operational stability across neural network models;

[0069] 5. The development process is simple, easy to use and easy to operate.

[0070] The above-mentioned embodiment only expresses one implementation mode of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the attached claims.

Claims

1. A software-defined parallel computing framework, characterized in that: The parallel computing architecture SDPCA includes an operating system domain, a system service domain, and a unified interface domain. The operating system domain is the basic software support for the entire architecture and is also the software interface for the underlying hardware accelerator. It registers hardware capabilities to the upper-level system service domain through device drivers, including device drivers, task scheduling, signal management, memory management, and various protocol stacks; the system service domain is the core of SDPCA, including three planes: the data plane, the control plane, and the capability plane. The data plane is responsible for loading, forwarding, mapping, and exporting all data, and is the main body of rapid data exchange; the control plane is responsible for computing task preprocessing, load separation, scheduling and allocation, dynamic reconstruction, computing scheduling, and task estimation functions, and supports extended customization; the capability plane is mainly responsible for computing node capability registration, capability reconstruction, operator management, capability management, and failure detection functions within the system; the unified interface domain includes distributed computing, model training, model reasoning, and resource management functions, and provides a cross-platform programming interface to upper-level applications.

2. The software-defined parallel computing framework according to claim 1, characterized in that: The underlying architecture includes a series of general-purpose processors CPU, graphics processor GPU, signal processor DSP, programmable logic device FPGA, and neural network processor NPU.

3. The software-defined parallel computing framework according to claim 1, characterized in that: The unified interface domain is compatible with and supports NCNN and MNN frameworks.

4. A method for implementing a parallel computing framework based on software definition, characterized in that: Based on the software-defined parallel computing framework implementation as described in any one of claims 1 to 3, the parallel computing architecture SDPCA software-defined computing resources based on plane features, maintained a dynamically reconfigurable virtual computing resource pool, and different application components can be transplanted and reused between architectures of the same standard; Control plane: This method builds the control plane of the SDPCA framework based on the software definition concept, realizes the dynamic self-organizing computing function of the host and device based on computing needs, and reduces the impact of the short board factor on the system computing power; sets the task estimation module, capability reconstruction module and capability database in the control plane. The task estimation module includes three parts: computing feature acquisition, accelerator feature extraction and task evaluation, namely task analysis and estimation; workload cutting, simulation sampling operation and cache consistency processing modules are introduced in this control plane; the execution of threads in the hardware accelerator is dynamically analyzed and scheduled in real time through the precise timing model to meet the requirements of the computing time limit; when the hardware accelerator performs parallel acceleration, the runtime state dynamically analyzes the load state of each accelerator according to the workload of each thread in the hardware accelerator and the built-in scheduling detection point, and then dynamically allocates the work to the relatively idle thread for execution. The computing scheduler integrates static scheduling and dynamic scheduling execution. In a program containing Fourier transform, image processing and filtering processing, the task data is deployed to different virtual memory spaces through static scheduling, and the dynamic scheduling inserts scheduling detection points during the program execution process to control the thread execution process of the computing task in the accelerator in real time; Data plane: This data plane is equipped with a minimalist protocol distributed soft bus, which virtualizes a series of computing memory blocks on the host side through the virtual memory management unit on the host side. These memory blocks reside in the physical memory of the host side. The virtual memory management unit is responsible for the unified addressing between the host side and the device side, and directly calls the high-speed interface at the physical layer for memory data transmission, forming a shared memory space between the host side and the device side, and always maintains the consistency of the host memory, accelerator device memory and cache data during the entire system operation; Capability plane: The capability plane realizes the dynamic construction of device capabilities and operator capabilities, and abstractly describes resource computing capabilities from the system level. Device capabilities are a software expression of hardware resources and are registered to the system through the interface provided by the framework. This capability plane divides devices into high-speed general-purpose processors, digital signal processors, and field programmable gate arrays. Operators are the basis for building the capability plane. In neural network computing, they are represented as kernel functions and are implemented in specific programming languages ​​in heterogeneous programming environments. When kernel functions run on heterogeneous devices, the resources of heterogeneous devices are processed in blocks so that the calculations can be distributed among multiple computing units for parallel execution.

5. The method for implementing a software-defined parallel computing framework according to claim 4, characterized in that: In the control plane, the maximum computing capacity of the parallel computing architecture should be represented by a set of related factors. The computing power of the i-th device is D[i]cap, the scheduling strategy capacity Scap, and the memory exchange capacity Mcap. The system computing capacity Ccap is expressed as shown in Formula 1: Combine the above features with parallel computing to maximize the computing power of the system.

6. The method for implementing a software-defined parallel computing framework according to claim 4, characterized in that: The computing characteristics are obtained by inputting the static code of the computing task, and are determined by the front-end compiler through static code analysis. When the task is loaded, the computing task characteristic data set collection and conversion are quickly completed. The task analysis and estimation module divides the characterized computing task into accelerator device subtasks, obtains the currently available accelerator capability by searching the device characteristics in the capability database, and matches it with the computing characteristics of this time, and evaluates the running time of the computing task on the corresponding device. When the accelerator device capability in the capability database does not meet the operator function or execution time requirements of the task, the capability reconstruction module is notified to reconstruct the capability of the accelerator device and update the capability description information to the capability database; capability reconstruction The construction module uses the host-side and device-side interaction interface to dynamically deploy operators to the corresponding accelerator devices according to the characteristics of the computing tasks. The host-side and device-side interaction interface definitions include SDPCA_GetPlatformInfo, SDPCA_GetDeviceCap, SDPCA_RegDeviceCap, SDPCA_RebuildDeviceCap, SDPCA_CreateCmdQue, SDPCA_CreateContext, SDPCA_CreateKernel, SDPCA_SetKernelArg, and SDPCA_TaskEmqueue.

7. The method for implementing a software-defined parallel computing framework according to claim 4, characterized in that: The data plane is equipped with a memory management model. The OpenCL memory variable concept is used in the host-side memory model. There are three memory types: array cache, image object, and pipeline object. The data in the array cache is continuous in the address space. The adjacent data in the image object image array is not guaranteed to be stored continuously in the memory. The pipeline object is a data element queue that follows the FIFO principle. The device side defines global memory, constant memory, local memory and private memory. Global memory is visible to every work item in the execution kernel and is the memory used for transmission between the host and device sides. Constant memory is accessed by all work items and is used to pass in configuration parameters. Local memory enables work items in the same work group to be shared and usually uses the on-chip memory of the device side. Private memory is only accessed within a work item.

8. The method for implementing a software-defined parallel computing framework according to claim 4, characterized in that: In the capability plane, the GPP resources in the system are defined to consist of N GPP devices interconnected by a specific topology, which are modeled as a GPP device linked list of length N. Each GPP device dynamically maintains M matching characteristics and A allocation characteristics, which are represented by P M and P A Expression,the GPP device modeling in the system is shown in Equation 2; For the nth GPP device list, n is one of 1, 2, ..., N, R M (GPP n ) represents GPP n The matching capability of the device refers to the device attribute characteristics in the computing environment that will constrain the computing process but cannot be quantitatively calculated, including processor model, operating system, and system clock; A (GPP n ) represents GPP n The allocation capacity of the device is the resource capacity attribute in the device that has a quantitative impact on the computing process, including the number of cores, main frequency, memory size, operator capacity description table, GPP n The theoretical value of the equipment capacity is given by R A (GPP n ) decides that when a new computing device GPP n After registering to the system, the capability plane automatically completes the software description of the device, builds typical matching characteristics and allocation characteristics for it, and dynamically maintains the R of the device throughout its life cycle. A (GPP n ); This method adapts the relevant kernel functions in the SDPCA framework and optimizes performance based on SIMD technology.

Citation Information

Cited By

  • FPGA (Field Programmable Gate Array) parallel computing unit scheduling method and device for image processing

    CN120744980A

  • FPGA parallel computing unit scheduling method and device for image processing

    CN120744980B

  • Distributed soft bus access method, device interconnection method and distributed soft bus architecture

    CN121567498A