Cross-platform programming optimization method, system and device for numerical simulation computing

Through technical means such as presetting common data types and creating work queues, the complex problem of heterogeneous computing platform programming model is solved, cross-platform unified programming and rapid migration are realized, and development efficiency and computing performance are significantly improved.

CN119759328BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510278341.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-16
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The diversity of heterogeneous computing platforms leads to complex programming models, requiring manual porting and optimization of algorithms, increasing programming complexity and development time, and improper resource management may lead to waste of computing resources and reduced performance.

Method used

By presetting common data types, determine the numerical simulation function and cyclic calculation mode, create a work queue to match the calculation tasks corresponding to the input data with the computing devices of heterogeneous platforms, and use the data accessor to transmit data and read results, realizing unified cross-platform programming and rapid migration.

Benefits of technology

It reduces the workload of manual porting and optimization, improves development efficiency and computing performance, and ensures unified and efficient data interaction on different heterogeneous platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759328B_ABST
    Figure CN119759328B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-platform programming optimization method, system and device for numerical simulation calculation. The method comprises: determining a numerical simulation function and a cyclic calculation mode according to a preset general data type. Obtaining input data of a host side by parsing the numerical simulation function. Creating a work queue according to the cyclic calculation mode and the computing task group of the device side, matching the computing task corresponding to the input data with the computing device of the heterogeneous platform according to the work queue, so that the computing device performs the computing task, obtains the computing result, and outputs the computing result to the host side through the data access device of the device side. According to the completion status of the computing task of parallel processing on the device side, the host side reads the computing result through the data access device to complete the numerical simulation calculation of the heterogeneous platform. The method can achieve unified programming and rapid migration between different heterogeneous platforms in the field of numerical simulation, thereby improving development efficiency and computing performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of numerical simulation technology, and in particular to a cross-platform programming optimization method, system and device for numerical simulation calculation. Background Art

[0002] Numerical simulation is an important branch of computer-aided mathematical problem solving. It discretizes and digitizes continuous mathematical problems in a digital way. With the continuous development of scientific research and engineering applications, the demand for computing performance in numerical simulation is also increasing. Traditional homogeneous computing devices are difficult to meet the growing computing needs due to their architectural limitations. In order to overcome the performance limitations of homogeneous devices, heterogeneous computing came into being. Heterogeneous computing improves computing performance and energy efficiency by integrating different types of computing units, and can dynamically allocate computing tasks to the most suitable computing units according to the characteristics of the computing tasks, thereby achieving performance optimization.

[0003] The efficient operation of heterogeneous computing systems depends not only on hardware, but also on corresponding software support. Since each hardware has unique architecture and performance characteristics, different hardware usually uses its own programming environment, tool chain, and instruction system. How to efficiently allocate tasks among multiple hardware and maximize the coordination of each hardware requires complex scheduling and optimization strategies. In order to fully utilize the computing power of heterogeneous systems, developers usually need to implement a new programming language or expand and transform existing languages.

[0004] As the importance of numerical simulations in scientific research and engineering applications becomes increasingly prominent, their scale and complexity are also growing, and the demand for computing performance has become more urgent. In order to meet these challenges, heterogeneous computing models have emerged, which improve computing performance and energy efficiency by integrating different types of computing units. However, the diversity of heterogeneous computing platforms also brings the complexity of programming models, resulting in the need for manual porting and optimization when deploying algorithms on different platforms. This is not only time-consuming, but also prone to errors.

[0005] Heterogeneous computing platforms integrate multiple computing units, including CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-programmable Gate Array), DSP (Digital Signal Processor), etc. Each computing unit has a unique programming model and interface. Therefore, when migrating an algorithm from one platform to another, manual transplantation and optimization are required to adapt to the programming interface and architecture of the new platform. However, this process increases the complexity of programming, making it challenging for the algorithm to fully utilize the computing power of multiple platforms. In addition, in a heterogeneous computing environment, how to effectively manage and schedule different computing resources for the architecture of each heterogeneous device to achieve optimal performance remains a huge challenge. Improper resource management may lead to waste of computing resources, increase energy consumption, and reduce energy efficiency. Especially in large-scale parallel computing tasks such as numerical simulation calculations, the efficiency of performance optimization and resource management directly affects the completion time and cost of computing tasks. Summary of the invention

[0006] Based on this, it is necessary to provide a cross-platform programming optimization method, system and device for numerical simulation calculations that can efficiently run numerical simulation calculations across platforms to address the above technical problems.

[0007] A cross-platform programming optimization method for numerical simulation computing, the method comprising:

[0008] The numerical simulation function and the loop calculation mode are determined according to the preset general data type.

[0009] The input data of the host side is obtained by parsing the numerical simulation function.

[0010] A work queue is created according to the cyclic computing mode and the computing task group on the device side. The computing tasks corresponding to the input data are matched with the computing devices of the heterogeneous platform according to the work queue, so that the computing devices execute the computing tasks and obtain the computing results. The computing results are output to the host side through the data access device on the device side.

[0011] According to the completion status of the parallel processing computing tasks on the device side, the host side reads the calculation results through the data access device to complete the numerical simulation calculation of the heterogeneous platform.

[0012] A cross-platform programming optimization system for numerical simulation computing, the system comprising:

[0013] The calculation mode and function configuration module is used to determine the numerical simulation function and the cyclic calculation mode according to the preset general data type.

[0014] The data acquisition module is used to obtain input data from the host by analyzing the numerical simulation function.

[0015] The heterogeneous platform computing task execution module is used to create a work queue according to the cyclic computing mode and the computing task group on the device side, and match the computing tasks corresponding to the input data with the computing devices on the heterogeneous platform according to the work queue, so that the computing devices execute the computing tasks, obtain the computing results, and output the computing results to the host side through the data accessor on the device side.

[0016] The numerical simulation calculation module is used to complete the numerical simulation calculation of the heterogeneous platform by reading the calculation results through the data access device on the host side according to the completion status of the parallel processing calculation tasks on the device side.

[0017] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0018] The numerical simulation function and the loop calculation mode are determined according to the preset general data type.

[0019] The input data of the host side is obtained by parsing the numerical simulation function.

[0020] A work queue is created according to the cyclic computing mode and the computing task group on the device side. The computing tasks corresponding to the input data are matched with the computing devices of the heterogeneous platform according to the work queue, so that the computing devices execute the computing tasks and obtain the computing results. The computing results are output to the host side through the data access device on the device side.

[0021] According to the completion status of the parallel processing computing tasks on the device side, the host side reads the calculation results through the data access device to complete the numerical simulation calculation of the heterogeneous platform.

[0022] The above cross-platform programming optimization method, system and device for numerical simulation calculation, firstly, determine the numerical simulation function and the cyclic calculation mode by presetting the general data type, which builds a unified data processing foundation for different platforms. Regardless of the heterogeneous platform, the numerical simulation function can be parsed according to this general standard to obtain the input data of the host side, which makes the data reading and preparation stage achieve cross-platform consistency. In the calculation task allocation link, a work queue is created according to the cyclic calculation mode and the calculation task group on the device side. This mechanism cleverly shields the differences in the hardware architecture of different platforms. Through the work queue, the calculation tasks corresponding to the input data can be accurately matched with the computing devices of the heterogeneous platform, so that each computing device can efficiently perform the calculation tasks it is good at. In this process, there is no need to manually write complex task allocation codes for different platforms, which greatly reduces the workload of manual transplantation and optimization. The data accessor on the device side plays a key bridge role in the whole process. It is not only responsible for outputting the calculation results to the host side, but also facilitates the host side to read the calculation results through the data accessor according to the completion of the parallel processing calculation tasks on the device side. This standardized data transmission and access method further ensures the uniformity and efficiency of data interaction on different heterogeneous platforms. In summary, this technical solution successfully solves the problem of complex programming models of heterogeneous computing platforms through a series of technical means such as presetting common data types, creating work queues, and using data accessors, and realizes unified programming and rapid migration between different heterogeneous platforms in the field of numerical simulation, significantly improving development efficiency and computing performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A schematic diagram of a flow chart of a cross-platform programming optimization method for numerical simulation calculations in one embodiment;

[0024] Figure 2 A schematic flow chart of cross-platform programming optimization steps for numerical simulation calculations in one embodiment;

[0025] Figure 3 A schematic diagram of a cross-platform programming architecture for numerical computing tasks in one embodiment;

[0026] Figure 4 An abstract schematic diagram of the application layer, framework layer and hardware layer of a cross-platform programming framework for the field of numerical simulation in one embodiment;

[0027] Figure 5 A structural block diagram of a cross-platform programming optimization system for numerical simulation calculations in one embodiment;

[0028] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0030] In one embodiment, Figure 1 As shown, a cross-platform programming optimization method for numerical simulation calculation is provided, comprising the following steps:

[0031] Step 102, determining a numerical simulation function and a cyclic calculation mode according to a preset general data type.

[0032] Specifically, write source code, reference necessary header files, and define common data types for numerical simulation applicable to heterogeneous devices.

[0033] Furthermore, the numerical simulation function is analyzed, the input data of the function is clarified, and the input data file is read to determine the output data of the numerical simulation function and the loop calculation mode.

[0034] Step 104, obtaining input data of the host side by analyzing the numerical simulation function.

[0035] Step 106, create a work queue according to the cyclic computing mode and the computing task group on the device side, match the computing tasks corresponding to the input data with the computing devices of the heterogeneous platform according to the work queue, so that the computing devices execute the computing tasks, obtain the computing results, and output the computing results to the host side through the data access device on the device side.

[0036] Specifically, use sycl::queue to create a work queue , submit computing tasks to the queue, and manage task execution on the device side through asynchronous operations. Use device selection functions such as default_selector(), host_selector(), cpu_selector(), or gpu_selector() to select the specified computing device through the device selector. Use sycl::buffer to create a buffer , encapsulate the host input data and device output data obtained in step 2.1 and step 2.2, buffer A reference to data used for data transfer between the host and the device. Set the data type, dimension, and size to match the data. Based on the input data requirements determined above, use sycl::accessor to declare the data accessor inside the lambda function. , use the get_access() method to initialize the accessor , bind it to a specific buffer On the top, set the accessor according to actual business needs. Read and write access to the corresponding buffer.

[0037] Furthermore, a lambda function is created through the submit() method and submitted to the work queue. The lambda function captures all variables in the external scope of the function as references and obtains the buffer Direct reference to the buffer inside the lambda function body Create a handler object. , used to manage parallel tasks executed on the device side.

[0038] Furthermore, according to the determined loop calculation mode, the number of working groups and the size of work items are set, which usually depends on the nature of the problem. Use sycl::range to create an N-dimensional object , where N represents the number of working groups, Specifies the size of the work-item grid for parallel computing tasks.

[0039] Furthermore, the parallel_for() function is used to start the parallel computing task, divide the task into multiple workgroups according to the cyclic computing mode, and allocate computing resources to each workgroup. On the corresponding computing device, a type identifier is used to assign a unique identifier to each computing task to be executed, so as to facilitate debugging and performance analysis. , specifies the size of the work item grid, that is, the number of tasks that need to be executed in parallel.

[0040] Further, create a sycl::id object , Each component in corresponds to the index of the current work item in the N-dimensional work item grid, from the device's data accessor Get input data from Calculating global index , used to access or modify data and start parallel computing tasks, which will be executed on the specified heterogeneous devices.

[0041] Step 108, according to the completion status of the parallel processing computing tasks on the device side, the host side reads the computing results through the data access device to complete the numerical simulation calculation of the heterogeneous platform.

[0042] Specifically, block the calling thread in the main program until the queue All submitted commands in the process are completed, ensuring that the execution commands of the parallel computing tasks have completed the data modification and can be used for subsequent access. Then, according to the determined output data requirements, use sycl::accessor to declare the accessor , use the get_access() method to initialize the accessor , set the accessor Buffer Access rights to data in.

[0043] Further, from the output accessor Read data, obtain numerical calculation results, and execute subsequent calculation tasks until all calculation tasks have been completed, completing the numerical simulation calculation of the heterogeneous platform.

[0044] In the above cross-platform programming optimization method for numerical simulation calculation, first, the numerical simulation function and the cyclic calculation mode are determined by presetting the general data type. This measure builds a unified data processing foundation for different platforms. Regardless of the heterogeneous platform, the numerical simulation function can be parsed according to this general standard to obtain the input data on the host side, which makes the data reading and preparation stages achieve cross-platform consistency. In the calculation task allocation link, a work queue is created according to the cyclic calculation mode and the calculation task group on the device side. This mechanism cleverly shields the differences in the hardware architecture of different platforms. Through the work queue, the calculation tasks corresponding to the input data can be accurately matched with the computing devices of the heterogeneous platform, so that each computing device can efficiently perform the calculation tasks it is good at. In this process, there is no need to manually write complex task allocation codes for different platforms, which greatly reduces the workload of manual transplantation and optimization. The data accessor on the device side plays a key bridge role in the whole process. It is not only responsible for outputting the calculation results to the host side, but also facilitates the host side to read the calculation results through the data accessor according to the completion of the parallel processing calculation tasks on the device side. This standardized data transmission and access method further ensures the uniformity and efficiency of data interaction on different heterogeneous platforms. In summary, this technical solution successfully solves the problem of complex programming models of heterogeneous computing platforms through a series of technical means such as presetting common data types, creating work queues, and utilizing data accessors. It realizes unified programming and rapid migration between different heterogeneous platforms in the field of numerical simulation, and significantly improves development efficiency and computing performance.

[0045] In one embodiment, the preset universal data type includes: defining a compilation header file based on the compilation type of heterogeneous devices in numerical simulation, selecting a numerical simulation function corresponding to the numerical simulation calculation of the computing device through the host side according to the type identifier of the header file, and determining a cyclic calculation mode according to the numerical simulation function.

[0046] In one of the embodiments, sycl::queue is used to create a work queue according to the cyclic computing mode and the computing tasks on the device side, and the computing tasks are matched with the computing devices of the heterogeneous platform through the device selector using device selection functions such as default_selector(), host_selector(), cpu_selector() or gpu_selector() and type identifiers.

[0047] In one of the embodiments, after creating a buffer to encapsulate input data and calculation results, a lambda function is created through the submit function, the calculation task is submitted to the work queue, all variables in the external scope of the function are captured in reference form, and a reference to the data pointed to by the buffer is obtained to complete the data transmission between the host and the device.

[0048] In one embodiment, the computing device uses a lambda function to retrieve input data of a data accessor, sets read and write access rights of a buffer corresponding to the data accessor according to the computing task, binds the input data to the corresponding buffer according to the read and write access rights, and sets the execution mode of each work group in the computing task and the execution range corresponding to the execution mode according to the cyclic computing mode. The computing task is executed according to the execution mode and the execution range to obtain a computing result, and the computing result is output to the host through the data accessor on the device side.

[0049] In one of the embodiments, attribute information of each work group in the computing task is set according to the execution mode and execution scope, a computing grid of the work items in the work group is created according to the attribute information, and the computing device executes the computing task according to the computing grid to obtain the computing result.

[0050] In one embodiment, after marking each computing task to be executed with a type identifier on computing devices at multiple device ends, a global index is calculated for the index object in the computing network according to the attribute information, and a parallel computing task is started to perform numerical simulation calculations on the computing devices until all computing devices at the device end are completed. The host end uses a synchronization function in the main program to block the calling thread until all commands in the work queue have been submitted, and then obtains the calculation results stored in the buffer through the data accessor to complete the numerical simulation calculation of the heterogeneous platform.

[0051] In one embodiment, if Figure 2 As shown, a cross-platform programming optimization step for numerical simulation calculation is provided, the specific contents are as follows:

[0052] Step 1: Define common data types applicable to heterogeneous devices.

[0053] Step 2: Determine the input and output data and loop calculation mode of the numerical simulation function.

[0054] Step 3: Create work queues, buffers and data accessors.

[0055] Specifically, the work queue is used to identify and specify heterogeneous computing devices; the buffer is used for data transmission between the host and the device; and the data accessor on the device side is used to access the buffer data object.

[0056] Step 4: Submit the work task to the queue.

[0057] Step 5: Set the execution mode of the workgroup and start the parallel computing task;

[0058] Specifically, the working group is configured according to the cyclic calculation mode of step 2, including defining the scope of parallel execution.

[0059] Step 6: Wait for the kernel to complete execution, read the output accessor, and extract the result;

[0060] Specifically, a parallel computing task is started, which will be executed on the heterogeneous device specified in step 3, and the host waits for the kernel execution on the device to complete. Then, the output accessor of the buffer object is obtained, the data in the buffer is read, and subsequent computing tasks are performed.

[0061] In one embodiment, if Figure 3 As shown, a cross-platform programming architecture for numerical computing tasks is provided. The numerical computing tasks are parsed to obtain the Node 0 character, which is used as the unique identifier for subsequent cross-platform programming of numerical computing tasks. The platform may include core chips such as multi-core CPUs, multi-core GPUs, FPGAs, and DSPs.

[0062] In one embodiment, if Figure 4 As shown, a cross-platform programming architecture for the field of numerical simulation is provided, which includes: an application layer, a framework layer and a hardware layer, wherein the application layer is used to obtain input data, output data and determine the cyclic calculation mode, execute multiple computing tasks in parallel in the performance optimization layer, divide the work items and work groups of each computing task, and perform data storage, call and calculation processing with the hardware layer in a task grid format. In addition, work queues, buffers and accessors are created for multiple computing tasks. The hardware layer includes: CPU, GPU and FPGA, which are used to start and execute parallel computing tasks according to the cyclic calculation mode of the application layer and the work queue and task grid format of the framework layer.

[0063] It should be understood that although Figure 1-Figure 2The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0064] In one embodiment, Figure 5 As shown, a cross-platform programming optimization system for numerical simulation calculation is provided, including: a calculation mode and function configuration module 502, a data acquisition module 504, a heterogeneous platform computing task execution module 506 and a numerical simulation calculation module 508, wherein:

[0065] The calculation mode and function configuration module 502 is used to determine the numerical simulation function and the cyclic calculation mode according to the preset general data type.

[0066] The data acquisition module 504 is used to acquire input data of the host side by analyzing the numerical simulation function.

[0067] The heterogeneous platform computing task execution module 506 is used to create a work queue according to the cyclic computing mode and the computing task group on the device side, and match the computing task corresponding to the input data with the computing device of the heterogeneous platform according to the work queue, so that the computing device executes the computing task, obtains the computing result, and outputs the computing result to the host side through the data access device on the device side.

[0068] The numerical simulation calculation module 508 is used to complete the numerical simulation calculation of the heterogeneous platform by reading the calculation results through the data access device at the host end according to the completion status of the parallel processing calculation tasks at the device end.

[0069] For the specific definition of the cross-platform programming optimization system for numerical simulation calculations, please refer to the definition of the cross-platform programming optimization method for numerical simulation calculations in the above text, which will not be repeated here. Each module in the above-mentioned cross-platform programming optimization system for numerical simulation calculations can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0070] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input system connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a cross-platform programming optimization method for numerical simulation calculations is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input system of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0071] Those skilled in the art will understand that Figure 5-Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0072] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0073] The numerical simulation function and the loop calculation mode are determined according to the preset general data type.

[0074] The input data of the host side is obtained by parsing the numerical simulation function.

[0075] A work queue is created according to the cyclic computing mode and the computing task group on the device side. The computing tasks corresponding to the input data are matched with the computing devices of the heterogeneous platform according to the work queue, so that the computing devices execute the computing tasks and obtain the computing results. The computing results are output to the host side through the data access device on the device side.

[0076] According to the completion status of the parallel processing computing tasks on the device side, the host side reads the calculation results through the data access device to complete the numerical simulation calculation of the heterogeneous platform.

[0077] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0078] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0079] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A cross-platform programming optimization method for numerical simulation computing, characterized in that: The method comprises: Determine the numerical simulation function and the cyclic calculation mode according to the preset general data type; Acquire input data of the host side by analyzing the numerical simulation function; Creating a work queue according to the cyclic computing mode and the computing task group on the device side, matching the computing task corresponding to the input data with the computing device of the heterogeneous platform according to the work queue, so that the computing device executes the computing task, obtains the computing result, and outputs the computing result to the host side through the data access device on the device side; According to the completion status of the parallel processing computing tasks on the device side, the host side reads the computing results through the data access device to complete the numerical simulation calculation of the heterogeneous platform.

2. The method according to claim 1, characterized in that The preset common data types include: compilation header files defined based on the compilation types of heterogeneous devices in numerical simulation; Determine the numerical simulation function and loop calculation mode according to the preset general data type, including: According to the type identifier of the header file, a numerical simulation function corresponding to the numerical simulation calculation of the computing device is selected through the host side; and a cyclic calculation mode is determined according to the numerical simulation function.

3. The method according to claim 2, characterized in that Creating a work queue according to the cyclic computing mode and the computing task group on the device side, and matching the computing task corresponding to the input data with the computing device of the heterogeneous platform according to the work queue, including: A work queue is created according to the cyclic computing mode and the computing tasks on the device side, and the computing tasks are matched with the computing devices of the heterogeneous platform through the device selection function of the device selector and the type identifier.

4. The method according to claim 3, characterized in that: Before the step of executing the computing task on the computing device to obtain a computing result and outputting the computing result to the host through the data access device on the device, the step further includes: After creating a buffer to encapsulate the input data and the calculation result, create a lambda function, submit the calculation task to the work queue, capture all variables in the external scope of the function in reference form, obtain a reference to the data pointed to by the buffer, and complete the data transmission between the host end and the device end.

5. The method according to claim 4, characterized in that The computing device executes the computing task, obtains a computing result, and outputs the computing result to the host through the data access device on the device side, including: The computing device uses a lambda function to retrieve input data of the data accessor, and sets read and write access permissions of a buffer corresponding to the data accessor according to the computing task, binds the input data to the corresponding buffer according to the read and write access permissions, and sets an execution mode of each work group in the computing task and an execution range corresponding to the execution mode according to the cyclic computing mode; The computing task is executed according to the execution mode and the execution range to obtain a computing result, and the computing result is output to the host side through the data access device on the device side.

6. The method according to claim 5, characterized in that Executing the computing task according to the execution mode and the execution range to obtain a computing result includes: The attribute information of each work group in the computing task is set according to the execution mode and the execution scope, a computing grid of the work items in the work group is created according to the attribute information, and the computing device executes the computing task according to the computing grid to obtain a computing result.

7. The method according to claim 6, characterized in that According to the completion status of the computing task processed in parallel on the device side, the host side reads the computing result through the data access device, and the numerical simulation computing of the heterogeneous platform is completed, including: After marking each of the computing tasks to be executed on the computing devices at the multiple device ends with a type identifier, a global index is calculated for the index objects in the computing network according to the attribute information, and a parallel computing task is started to perform numerical simulation calculations on the computing devices until all the computing devices at the device ends have completed the execution; The host side uses a synchronization function in the main program to block the calling thread until all commands in the work queue have been submitted, and then obtains the calculation results stored in the buffer through the data accessor to complete the numerical simulation calculation of the heterogeneous platform.

8. A cross-platform programming optimization system for numerical simulation computing, characterized in that: The system comprises: A calculation mode and function configuration module, used to determine the numerical simulation function and the cyclic calculation mode according to the preset general data type; A data acquisition module, used for acquiring input data of a host terminal by analyzing the numerical simulation function; A heterogeneous platform computing task execution module is used to create a work queue according to the cyclic computing mode and the computing task group on the device side, and match the computing task corresponding to the input data with the computing device of the heterogeneous platform according to the work queue, so that the computing device executes the computing task, obtains the computing result, and outputs the computing result to the host side through the data access device on the device side; The numerical simulation calculation module is used to complete the numerical simulation calculation of the heterogeneous platform by reading the calculation results through the data access device according to the completion status of the parallel processing calculation tasks on the device side.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Heterogeneous distributed task processing system and processing method in cloud computing platform

    CN105022670A

  • Grid-based universe online simulation system for computing resource integration

    CN115543628A