Self-adaptive acceleration calculation method and system and computer readable medium

Through the adaptive acceleration calculation method, the calculation paths in the numerical simulation software are automatically analyzed and determined, which solves the problems of insufficient utilization of computing resources and poor efficiency of manual selection of acceleration methods in the prior art, and realizes the optimization configuration of computing resources and the improvement of computing efficiency.

CN120179406APending Publication Date: 2025-06-20SHANGHAI NUCLEAR ENGINEERING RESEARCH & DESIGN INSTITUTE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339746.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When existing numerical simulation software deals with complex physical problems, the calculation resources are insufficient, resulting in too long calculation time, and manual acceleration method is not efficient, making it difficult to find the optimal calculation path.

Method used

An adaptive acceleration calculation method is proposed. By analyzing the geometric structure and physical characteristics of the case model, efficient calculation paths are automatically analyzed and determined, including grid cutting and division, preset parameter calculation, calculation component filtering and parallel acceleration library selection, and the final calculation path is generated and filtered.

Benefits of technology

The optimized configuration of computing resources is realized, the computing efficiency and accuracy of numerical simulation tasks are improved, and the complex analysis process that originally required users to manually perform is automated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179406A_ABST
    Figure CN120179406A_ABST
Patent Text Reader

Abstract

The invention relates to a self-adaptive acceleration calculation method and system and a computer readable medium. The method comprises the steps that S1, the geometric structure and physical characteristics of an example model in numerical simulation software are analyzed; s2, performing grid cutting division on the geometric structure to obtain the number of grids; s3, corresponding preset parameters are obtained according to the physical characteristics, and the number of calculation cores preliminarily adopted by the example model is calculated according to the preset parameters and the number of the grids; s4, screening out a calculation component of an algebraic matrix solving process according to the physical characteristics; s5, screening out a parallel acceleration library according to the physical characteristics; s6, generating a plurality of preliminary calculation paths according to the preliminarily adopted calculation core number, the calculation component and the parallel acceleration library; and S7, screening out at least one final calculation path from the plurality of preliminary calculation paths. According to the method, the efficient calculation path can be automatically analyzed and determined, and the calculation efficiency can be improved by adopting the calculation path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application mainly relates to the field of software acceleration computing technology, and particularly relates to an adaptive acceleration computing method, system and computer-readable medium. Background Art

[0002] The application of Computational Fluid Dynamics (CFD) in the engineering field is becoming more and more extensive, and numerical simulation and analysis of example models can be carried out through open-source CFD software. When users use numerical simulation software to solve physical problems, they mainly focus on the accuracy and precision of the problems rather than the consumption of computing resources. This results in a longer three-dimensional calculation time for complex physical problems under the same hardware conditions, extending the cycle of scientific research computing and engineering design.

[0003] Some numerical simulation software provides various means of accelerating computing, but when dealing with complex problems, these acceleration methods are not always the optimal computing strategies. When dealing with large-scale computing problems, such as when the number of parallel computing cores required exceeds 400 cores, the built-in solver of numerical simulation software may not be able to provide sufficient parallel expansion capabilities and cannot meet the needs of high-performance computing. At the same time, GPU (Graphics Processing Unit) acceleration technology is becoming more and more popular due to its powerful parallel processing capabilities. However, numerical simulation software is still in the stage of development and improvement in terms of utilizing GPU resources to achieve improved computing performance.

[0004] In engineering design and scientific research computing, computing efficiency directly affects the project cycle and cost. Currently, numerical simulation software provides various acceleration technologies for users to choose from, but the efficiency of manually selecting acceleration methods is not good. Moreover, how to select the most suitable computing path during the computing process and make full use of the parallel processing capabilities of the computer to achieve optimal performance is a complex and challenging problem. Summary of the Invention

[0005] The technical problem to be solved by this application is to provide an adaptive acceleration computing method, system and computer-readable medium, which can automatically analyze and determine an efficient computing path according to the factors affecting the computing time during the computing process. Adopting this computing path can improve computing efficiency and effectively reduce computing costs.

[0006] The technical solution adopted by this application to solve the above technical problems is an adaptive acceleration calculation method for numerical simulation software, including: Step S1: Analyze the geometric structure and physical properties of the example model in the numerical simulation software; Step S2: Perform mesh cutting and division on the geometric structure to obtain the number of meshes; Step S3: Obtain the corresponding preset parameters according to the physical properties, and calculate the initially adopted number of computing cores of the example model according to the preset parameters and the number of meshes; Step S4: Screen out the computing components in the algebraic matrix solution process according to the physical properties; Step S5: Screen out the parallel acceleration library according to the physical properties; Step S6: Generate multiple initial calculation paths according to the initially adopted number of computing cores, computing components, and parallel acceleration library; Step S7: Screen out at least one final calculation path from the multiple initial calculation paths.

[0007] In an embodiment of this application, the geometric structure has numerical values in the X-axis direction, Y-axis direction, and Z-axis direction under the three-axis coordinate system OXYZ, and the X-axis, Y-axis, and Z-axis of the three-axis coordinate system OXYZ are perpendicular to each other; in Step S2, performing mesh cutting and division on the geometric structure includes: Step S2a: Construct a cutting relational expression using the following formula:

[0008]

[0009] a*b*c=n

[0010] where X represents the numerical value in the X-axis direction; Y represents the numerical value in the Y-axis direction; Z represents the numerical value in the Z-axis direction; a represents the corresponding number of computing cores in the X-axis direction; b represents the corresponding number of computing cores in the Y-axis direction; c represents the corresponding number of computing cores in the Z-axis direction; n represents the total number of computing cores; Step S2b: Perform equal division on the numerical values in the X-axis direction, Y-axis direction, and Z-axis direction according to the cutting relational expression.

[0011] In an embodiment of this application, the preset parameters include: the performance coefficient of the computing core and the physical coefficient of the physical property; in Step S3, calculating the initially adopted number of computing cores of the example model according to the preset parameters and the number of meshes includes: calculating the initially adopted number of computing cores using the following formula:

[0012] core=NG / (β*p / α)

[0013] where core represents the initially adopted number of computing cores; NG represents the number of meshes; β represents the physical coefficient of the physical property; α represents the performance coefficient of the computing core; p represents the frequency of the computing core.

[0014] In an embodiment of the present application, in step S7, screening at least one final calculation path from multiple preliminary calculation paths includes: step S7a: calculating the optimal resource consumption value according to the multiple preliminary calculation paths; step S7b: taking the preliminary calculation path corresponding to the optimal resource consumption value as the final calculation path.

[0015] In an embodiment of the present application, the following formula is used to calculate the optimal resource consumption value:

[0016] v = min(n i *t i )

[0017] where v represents the optimal resource consumption value; min(·) represents the minimum value function; i represents the number of preliminary calculation paths; n i represents the number of computing cores of the i-th preliminary calculation path; t i represents the calculation time of the i-th preliminary calculation path.

[0018] In an embodiment of the present application, after step S7a and before step S7b, it further includes: judging whether the optimal resource consumption value is greater than a preset threshold; if the judgment is yes, then turn to execute step S4; if the judgment is no, then turn to execute step S7b.

[0019] In an embodiment of the present application, the numerical simulation software is such as CFD software; the computing components include: a pre-processor, a solver, a smoother, an external library solver. The pre-processor includes one or any combination of Diagonal, DIC, FDIC, GAMG, DILU. The solver includes one or any combination of Diagonal, PCG, PBiCG, GAMG, CG. The smoother includes Dic GaussSeidel or Diabetesgonal. The external library solver includes one or any combination of GMRES, Richardson, CG, CGS; the parallel acceleration library includes one or any combination of PETSc, AmgX, SUNDIALS, Hypre, SuperLU, MUMPS.

[0020] In an embodiment of the present application, the physical properties include: one or any combination of compressible, incompressible, multiphase flow, heat transfer, mass transfer, chemical reaction.

[0021] The present application also proposes an adaptive acceleration computing system to solve the above technical problems, including: a memory for storing instructions executable by a processor; a processor for executing the instructions to implement the above adaptive acceleration computing method.

[0022] The present application also provides a computer-readable medium storing computer program code, which implements the above adaptive acceleration calculation method when executed by a processor.

[0023] The technical solution of the present application can automatically analyze the characteristics of a calculation task. By parsing the geometric structure and physical properties of the example model, and comprehensively considering several influencing factors such as mesh generation, physical model, numerical solver, and parallel acceleration library, it can adaptively find multiple reasonable preliminary calculation paths according to the specific requirements of the task, computer hardware configuration, available computing components, and available parallel acceleration libraries, and finally select the optimal calculation path, thereby realizing the optimal allocation of computing resources and improving the calculation efficiency and accuracy of numerical simulation tasks. The present application automates the complex analysis process that originally needed to be manually performed by users. Compared with traditional manual optimization, the present application improves the degree of automation of the calculation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application will be given with reference to the accompanying drawings, where:

[0025] Figure 1 is an exemplary flowchart of the adaptive acceleration calculation method according to an embodiment of the present application;

[0026] Figure 2 is a schematic diagram of the solution process of CFD software and the accelerable modules in an embodiment of the present application;

[0027] Figure 3 is a schematic diagram of different calculation modules of CFD software in an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of the selectable matrix acceleration library of CFD software in an embodiment of the present application;

[0029] Figure 5 is a flowchart of adaptively selecting a solver and solving an example in an embodiment of the present application;

[0030] Figure 6 is a flowchart of the adaptive acceleration calculation method based on CFD software in an embodiment of the present application;

[0031] Figure 7 is a system block diagram of the adaptive acceleration calculation system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application will be given with reference to the accompanying drawings.

[0033] In the following description, numerous specific details are set forth to provide a thorough understanding of the present application. However, the present application may be implemented in other ways different from those described herein. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0034] As used in the present application and the claims, unless the context clearly dictates otherwise, the words "a," "an," "one," and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" merely indicate the inclusion of the steps and elements that have been specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0035] Flowcharts are used in the present application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, they can be executed in reverse order or simultaneously. Also, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0036] First, an analysis of the application process of the numerical simulation software in the present application is introduced herein.

[0037] In practical applications, the core technical challenge faced by adaptive acceleration computing lies in identifying the influencing factors and selecting the optimal computing path to optimize the computing time. Figure 2 It is a schematic diagram of the solution process of the CFD software and the accelerable modules. As Figure 2 shown, there are various factors affecting the computing duration during the computing process of the numerical simulation software. For example, it is necessary to define the momentum equation, solve the momentum prediction, calculate the coefficients of pressure and velocity, etc. The present application extracts the modules that affect the computing rate, analyzes and arranges the computing process, and determines that there are four categories of key influencing factors for parallel acceleration: (1) the grid cutting and division in the three directions of the X, Y, and Z axes; (2) the diversity of the types of physical problems to solve the physical model; (3) the types of numerical solvers; (4) the available parallel acceleration libraries.

[0038] In order to determine the optimal acceleration method among these four categories of key influencing factors for computing speed, it is difficult to achieve the optimal acceleration effect solely based on experience. To achieve the optimal acceleration effect, it is necessary to conduct a detailed analysis and distributed processing of the entire computing process, construct a comprehensive acceleration scheme that can simultaneously consider the interactions between these factors to achieve the acceleration of software operation.

[0039] Based on the above analysis, this application proposes an adaptive acceleration calculation method, which is suitable for being applied to numerical simulation software (such as CFD software), that is, this application is equivalent to an adaptive acceleration calculation method based on CFD software. The adaptive acceleration calculation method of this application can run locally on a computer or in a cloud platform. When the adaptive acceleration calculation method runs on a cloud platform, the data on the local computer and the cloud platform data are interacted through a wireless network. Exemplarily, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an interconnected cloud, a multi-cloud, etc. or any combination thereof. This application places no restrictions on the operating environment of the adaptive acceleration calculation method.

[0040] Here, algorithms 1, 2, 3, criterion 1, and criterion 2 designed in this application are introduced uniformly first, and these algorithms and criteria will be used in the adaptive acceleration calculation method introduced later.

[0041] Algorithm 1: The best resource consumption formula is defined as: the number of computing cores (n) * computing time (t), denoted as min(n*t).

[0042] Algorithm 2: To ensure that each process can be allocated sufficient computing power, this application divides the numerical values of the X-axis, Y-axis, and Z-axis respectively according to the geometric dimensions of the input example model to obtain corresponding cutting schemes. For the sake of simplifying the decomposition, the length, width, and height of the computational domain of a complex example model can be represented as X, Y, and Z respectively, and the number of computer cores used can be represented as n, which is decomposed into the form of a*b*c = n, where a, b, and c respectively represent the integer decomposition of the number of computer cores n in the length, width, and height directions. This application makes the three values of X / a, Y / b, and Z / c approximate.

[0043] Algorithm 3: According to the different physical models of the example, reasonable grid numbers, performance coefficients of computing cores, and physical coefficients of the physical characteristics of the example model can be preset. For example: Based on the CPU (Central Processing Unit) of Intel Xeon Gold 6258R, in the X86 environment, the grid numbers are set as follows:

[0044] a) For an incompressible example, it is more reasonable to calculate 200,000 grids with a single core, and for a compressible example, it is more reasonable to calculate 100,000 grids with a single core.

[0045] b) For a multiphase flow case, it is more reasonable to calculate 50,000 grids with a single core.

[0046] c) For a heat transfer case, it is more reasonable to calculate 30,000 grids with a single core.

[0047] d) For a mass transfer case, it is more reasonable to calculate 30,000 grids with a single core.

[0048] e) For a complex reaction case, it is more reasonable to calculate 30,000 grids with a single core.

[0049] If the 2.7HZ performance coefficient α of the Intel Xeon Gold 6258R CPU is set to 1, the performance coefficients of the other CPUs are divided by 2.7GHZ to obtain the corresponding coefficients.

[0050] The performance coefficient of the computing core is set as follows:

[0051] a) The incompressible algorithm coefficient β is set to 20.

[0052] b) The compressible flow coefficient β is set to 10.

[0053] d) The coefficient β of the multiphase flow is set to 5.

[0054] e) The heat transfer coefficient β is set to 5.

[0055] f) The mass transfer coefficient β is set to 1.

[0056] g) The coefficient β of reaction and combustion is set to 0.5.

[0057] This application uses α to represent the performance coefficient of the computing core CPU; β represents the coefficient of the physical problem adopted, that is, the physical coefficient of the physical characteristics of the example model; NG represents the number of grids in the computing problem. Assuming that the number of grids of the given example model is NG, this application determines NG / (β*2.7 / α) as the number of computing cores that can be initially adopted.

[0058] Criteria 1: The pre-processing methods used in this application include Diagonal, DIC, FDIC, GAMG, DILU, etc.; the solvers include Diagonal, PCG, PBiCG, GAMG, etc.; the smoothers include Dic GaussSeidel, Diabetesgonal, etc.; the external library solvers include GMRES, Richardson, CG, CGS, etc. This application can adaptively preliminarily screen these algorithms based on the algorithm comparison of the official case of CFD software.

[0059] Criterion 2: This application provides multiple parallel acceleration libraries for selection, and the appropriate parallel module can be selected according to the characteristics of the example. For example, for large-scale supercomputers with more than 400 cores, an external PETSc solver accelerator can be selected, and if the cluster has a GUDA accelerator card, the AmgX solver can be used. According to the hardware characteristics of the computer, the solver, PETSc and AmgX accelerator card that come with the CFD software can be selected.

[0060] This application addresses the problem of the computational efficiency of numerical simulation software in a large-scale parallel computing environment, especially high-performance systems with more than 400 cores. This application integrates the PETSc solver module and the GPU acceleration module into the CFD software, enhancing the performance of the adaptive acceleration system and providing diverse high-performance computing solutions to meet the needs of different computing tasks and hardware environments.

[0061] The adaptive acceleration calculation method of this application will be introduced later.

[0062] Figure 1 is an exemplary flowchart of the adaptive acceleration calculation method of an embodiment of this application. Refer to Figure 1 As shown, the adaptive acceleration calculation method of this embodiment includes the following steps:

[0063] Step S1: Analyze the geometric structure and physical properties of the case model in the numerical simulation software.

[0064] Step S2: Perform mesh cutting and division on the geometric structure to obtain the number of meshes.

[0065] Step S3: Obtain the corresponding preset parameters according to the physical properties, and calculate the number of computing cores initially adopted by the case model according to the preset parameters and the number of meshes.

[0066] Step S4: Screen out the computing components in the algebraic matrix solution process according to the physical properties.

[0067] Step S5: Screen out the parallel acceleration library according to the physical properties.

[0068] Step S6: Generate multiple preliminary calculation paths according to the initially adopted number of computing cores, computing components, and parallel acceleration library.

[0069] Step S7: Screen out at least one final calculation path from the multiple preliminary calculation paths.

[0070] The above steps S1 to S7 will be described in detail below:

[0071] In step S1, analyze the geometric structure and physical properties of the case model in the numerical simulation software.

[0072] Exemplarily, the numerical simulation software includes CFD software. This application analyzes the physical properties of the case model to facilitate full consideration of load balancing. The physical properties or calculation contents of the case model include compressible, incompressible, multiphase flow, heat transfer, mass transfer, chemical reaction, etc. In practical applications, a case model may have multiple physical properties.

[0073] Figure 3 is a schematic diagram of different calculation modules of CFD software in an embodiment of this application. Exemplarily, asFigure 3 As shown, the CFD software provides different algorithms for different physical properties, including incompressible, compressible flow, multiphase flow, conjugate heat transfer, mass transfer, reaction, combustion, etc. Each physical model can correspond to multiple algorithm paths, and the optimal amount of calculation for single CPU and GPU calculations of the grid can be set according to the complexity of the example.

[0074] In step S2, the geometric structure is cut and divided into grids to obtain the number of grids. Exemplarily, such a setting in the present application can provide data support for subsequent generation of the calculation path and improve the reliability of the calculation path.

[0075] In some embodiments, the geometric structure has numerical values in the X-axis direction, Y-axis direction, and Z-axis direction in the three-axis coordinate system OXYZ, and the X-axis, Y-axis, and Z-axis of the three-axis coordinate system OXYZ are perpendicular to each other; in step S2, cutting and dividing the geometric structure into grids includes:

[0076] Step S2a: Construct a cutting relation formula using the following formulas (1) and (2):

[0077]

[0078] a*b*c = n (2)

[0079] Where X represents the numerical value in the X-axis direction; Y represents the numerical value in the Y-axis direction; Z represents the numerical value in the Z-axis direction; a represents the number of computing cores corresponding in the X-axis direction; b represents the number of computing cores corresponding in the Y-axis direction; c represents the number of computing cores corresponding in the Z-axis direction; n represents the total number of available computing cores;

[0080] Step S2b: According to the above cutting relation formula, equally divide the numerical values in the X-axis direction, Y-axis direction, and Z-axis direction.

[0081] Exemplarily, in the present application, through a specific cutting relation formula, the grid cutting and division of the geometric structure are realized. By equally dividing the numerical values in the X, Y, and Z-axis directions, the reasonable matching of the grid division and the number of computing cores is ensured, thereby improving the reliability of the computing resources in the calculation path.

[0082] In step S3, obtain the corresponding preset parameters according to the physical properties, and calculate the number of computing cores initially adopted by the example model according to the preset parameters and the number of grids.

[0083] In some embodiments, the preset parameters include: the performance coefficient of the computing core and the physical coefficient of the physical property; in step S3, calculating the number of computing cores initially adopted by the example model according to the preset parameters and the number of grids includes: calculating the initially adopted number of computing cores using the following formula (3):

[0084] core = NG / (β * p / α) (3)

[0085] Among them, core represents the number of computing cores initially adopted; NG represents the number of grids; β represents the physical coefficient of physical properties; α represents the performance coefficient of the computing core; p represents the frequency of the computing core, and p can be set as a constant. In this application, p is set to 2.7. Exemplarily, by calculating the number of computing cores initially adopted in this application, it is convenient to estimate the utilization of the computing cores, and subsequently, appropriate computing paths can be selected in combination with different numbers of computing cores.

[0086] In step S4, the computing components for the algebraic matrix solution process are screened according to the physical properties.

[0087] Figure 4 It is a schematic diagram of the matrix acceleration library that can be selected by the CFD software in an embodiment of this application. Exemplarily, referring to Figure 4 As shown, the CFD software provides a variety of computing components, including a preprocessor, a solver, a smoother, and an external library solver, and these components all have a significant impact on the computing speed. This application will intelligently select the best matrix preprocessing technology, solution algorithm, and smoothing strategy based on the characteristics of the example to optimize the overall computing performance.

[0088] In some embodiments, the computing components include: a preprocessor, a solver, a smoother, and an external library solver. The preprocessor includes one or any combination of Diagonal, DIC, FDIC, GAMG, and DILU. The solver includes one or any combination of Diagonal, PCG, PBiCG, GAMG, and CG. The smoother includes Dic GaussSeidel or Diabetesgonal. The external library solver includes one or any combination of GMRES, Richardson, CG, and CGS.

[0089] In step S5, the parallel acceleration library is screened according to the physical properties. Exemplarily, the parallel acceleration library can improve the computing efficiency in the numerical simulation process, ensure that the computing tasks are executed quickly and accurately in a high-performance computing environment, thereby optimizing resource utilization and shortening the computing cycle.

[0090] In some embodiments, the parallel acceleration library includes one or any combination of PETSc, AmgX, SUNDIALS, Hypre, SuperLU, and MUMPS. Exemplarily, these parallel acceleration libraries can basically adapt to clusters with more than 400 cores and are suitable for using the latest GPU resources.

[0091] Exemplarily, in practical applications, assume that the number of cores of the computer MPI (Message Passing Interface) used is N, and the number of graphics cards is M. For CFD software, it has a certain parallel computing ability. If the computer to be used does not have the large-scale parallel computing ability of more than 400 cores, the PETSc solver can be introduced for calculation. If the computer system to be used has GPU graphics cards that can be accelerated, and the computational fluid dynamics software adopted for calculation does not have a GPU acceleration module, the AmgX solver can be introduced.

[0092] Figure 5 It is a flowchart of adaptively selecting a solver and solving an example in an embodiment of the present application. Exemplarily, referring to Figure 5 As shown, at step S510, select the compilation version during compilation. After compilation is completed, input the example and select a solver to solve it; at step S520, use CPU parallel solution; at step S530, use the original solver to solve; at step S540, use the PETSc solver to solve; at step S550, use GPU parallel solution; at step S560, use the AmgX solver to solve; at step S570, the solution ends and the result is output.

[0093] In step S6, multiple preliminary calculation paths are generated according to the initially adopted number of computing cores, computing components, and parallel acceleration libraries. Exemplarily, during the process of generating multiple preliminary calculation paths, the initially calculated number of computing cores can be adopted, or the number of computing cores can be appropriately increased or decreased, which is not limited in the present application. Such a setting in the present application can flexibly generate multiple preliminary calculation paths, providing diversified execution schemes for subsequent calculation tasks, thereby improving the adaptability and execution efficiency of the calculation tasks.

[0094] In step S7, at least one final calculation path is screened out from the multiple preliminary calculation paths. In some embodiments, it includes:

[0095] Step S7a: Calculate the optimal resource consumption value according to the multiple preliminary calculation paths.

[0096] Step S7b: Take the preliminary calculation path corresponding to the optimal resource consumption value as the final calculation path.

[0097] Exemplarily, by calculating and comparing the resource consumption values of multiple preliminary calculation paths, the path with the optimal resource consumption can be accurately identified and determined as the final calculation path, thereby realizing the efficient utilization of computing resources and the effective control of costs while ensuring the calculation accuracy.

[0098] In some embodiments, the following formula (4) is used to calculate the optimal resource consumption value:

[0099] v = min(n i *t i ) (4)

[0100] where v represents the optimal resource consumption value; min(·) represents the minimum value function; i represents the number of preliminary calculation paths; n i represents the number of computing cores of the i-th preliminary calculation path; t i represents the computing time of the i-th preliminary calculation path.

[0101] In some embodiments, after step S7a and before step S7b, it further includes: determining whether the optimal resource consumption value is greater than a preset threshold; if the determination is yes, then proceed to execute step S4; if the determination is no, then proceed to execute step S7b. Exemplarily, by calculating the resource consumption values of each preliminary calculation path and selecting the path with the smallest product of the number of computing cores and the computing time as the optimal calculation path, the calculation scheme with the optimal resource consumption can be accurately determined, thereby realizing the efficient utilization of computing resources while ensuring the calculation accuracy.

[0102] Figure 6 is a flowchart of an adaptive acceleration calculation method based on CFD software in an embodiment of the present application. Exemplarily, as shown in Figure 6 the present application first analyzes and classifies the characteristics of the calculation example and the hardware characteristics; then obtains the physical characteristics (such as single-phase flow, etc.), the number of grids, the geometric structure, the number of CPU cores and other hardware characteristics, and the solver type; after comprehensively considering various influencing factors, solves and screens out the optimal number of cores, the optimal grid division, the optimal solver, and the optimal acceleration library; finally, screens the solution line with the best performance for calculation.

[0103] Exemplarily, the overall process of the adaptive acceleration calculation method of the present application is as follows:

[0104] (1) For any CFD software calculation example, obtain the geometric structure, physical model, numerical solver, and acceleration library of the calculation example model, and analyze and determine the calculation paths attempted by this case according to the algorithms 2, 3, criterion 1, and criterion 2 mentioned above, and preliminarily screen out better numbers of computing cores, better grid divisions, better solvers, and better acceleration libraries.

[0105] (2) Perform preliminary calculations, iterate a small number of steps, and obtain calculation paths with faster calculations.

[0106] (3) Further optimize the calculation paths, and evaluate the efficiency of the calculation example according to algorithm 1.

[0107] (4) Based on the results of the performance evaluation, continue to automatically adjust the calculation paths according to criterion 1 and criterion 2 to find the optimal number of computing cores, the best grid division, the best solver, and the best acceleration library.

[0108] This application has been verified on three implementation cases, which are introduced below.

[0109] Implementation Case 1:

[0110] In a wind environment simulation example of a certain city, the number of grids in the example is 1.98 million, the solution algorithm is SIMPLE, and the turbulence model is k-Omega SST. This application encapsulates the adaptive acceleration system of the adaptive acceleration calculation method and successfully identifies the calculation strategy with the lowest resource consumption on a 56-core computing platform. The identification process is as follows:

[0111] This case involves single-phase flow. The optimal number of grids allocated to each CPU is about 100,000 to 200,000. The system preliminarily determines that the grid cutting process for 8-core computing is 2-2-2, and the iteration time of the solver provided by the CFD program is 131 steps.

[0112] (1) Using the solver algorithm provided by the CFD program, the preprocessor, smoother and solver are GaussSeidel+none+GAMG respectively, and the calculation time is 182 seconds.

[0113] (2) The PETSc algorithm added to the acceleration system is GS+AMG+CG, and the calculation time is 622 seconds.

[0114] (3) The calculation time of GS+AMG+Richardson is 902 seconds.

[0115] (4) GS+AMG+biCG does not converge.

[0116] (5)GS+jacobi(icc)+CG does not converge.

[0117] (6) The calculation time of GS+jacobi(ilu)+CG is 534 seconds.

[0118] (7)GS+ilu(ilu)+CG does not converge.

[0119] (8) The calculation time of GS+GAMG+CG is 392 seconds.

[0120] (9) The calculation time of GS+GAMG+pipeprCG is 511 seconds.

[0121] (10) The calculation time for none+GAMG+CG is 398 seconds.

[0122] (11) The calculation time of DICGaussSeidel+GAMG+CG is 390 seconds.

[0123] (12) The calculation time of DICGaussSeidel + GAMG + biCGs is 452 seconds.

[0124] (13) The calculation time of DICGaussSeidel + GAMG + groppCG is 416 seconds.

[0125] (14) The calculation time of DICGaussSeidel + GAMG + CG after removing the KSP_CG_single_reduction function is 400 seconds.

[0126] (15) The calculation time of DICGaussSeidel + GAMG + richardson is 452 seconds.

[0127] According to the preliminary calculation times in (1)-(15), the solver times of the CFD software's built-in solver GaussSeidel + none + GAMG and the DICGaussSeidel + GAMG + CG of PETSc added to the system are the best. Subsequently, these two initially screened solvers were used, and according to Criterion 1 and Criterion 2, calculations were performed using different numbers of cores. The total number of calculation steps was 2000 steps, and the solution strategies are summarized in Table 1 below:

[0128] Table 1 Solution strategies for Implementation Case 1

[0129]

[0130] By multiplying the number of cores by the calculation time, the optimal calculation path selected through the screening method is the CFD program's built-in solver GaussSeidel + none + GAMG, with 4 cores, divided into 2*2*1, which consumes the least computing resources. Subsequently, subsequent calculations were performed using this calculation path, which can save computing resources.

[0131] Implementation Case 2:

[0132] In this case, a numerical simulation of turbulent flow in a three-dimensional straight pipe was carried out, with a total of 1.76 million grid points. The server model used was A800, equipped with 56 processor cores. The SIMPLE algorithm was selected as the solution method, and the hybrid k-Omega SST turbulence model was used to simulate single-phase flow. During the initial screening process, the following three solver configurations were considered in this application:

[0133] (1) The system's built-in solver, combined with the Gauss-Seidel iteration method, no preconditioning (none), and geometric algebraic multigrid (GAMG) technology.

[0134] (2) Utilize the large-scale parallel computing library PETSc, select the Gauss-Seidel iterative solver, also adopt no preconditioning, and combine the conjugate gradient (CG) solver.

[0135] (3) For GPU-accelerated computing, consider the AmgX solver, adopt the Gauss-Seidel iterative method, with no preconditioning, and combine the GAMG technology.

[0136] Based on these configurations, perform strong scalability tests on the GAMG, PETSc, and AmgX solvers on an A800 server. The tests cover different numbers of cores, namely 1, 2, 8, 16, 32, and 56 cores, to evaluate their performance in different parallel environments. The solution strategies are summarized in Table 2 below:

[0137] Table 2 Solution Strategies for Embodiment 2

[0138]

[0139]

[0140] The best calculation path selected in this application is the 4-core 2*2*1 calculation path of GAMG+none+CG of AmgX, and further calculations can be performed using this calculation path.

[0141] Embodiment 3:

[0142] When performing numerical simulations of three-dimensional cavity flows, use 10 million grid points and execute the calculation tasks with 1000 cores on a large-scale cluster. To achieve optimal computational efficiency, compare two different solver configurations:

[0143] (1) The solver provided by the system, adopt the Gauss-Seidel iterative method, cooperate with no preconditioning (none) and geometric algebraic multigrid (GAMG) technology.

[0144] (2) The Gauss-Seidel iterative solver provided by the large-scale parallel computing library PETSc, also adopts no preconditioning, but combines the conjugate gradient (CG) solver. The solution strategies are summarized in Table 3 below:

[0145] Table 3 Solution Strategies for Embodiment 3

[0146]

[0147] The best calculation path selected in this application is: the 448-core 8*8*7 mode of the GAMG+none+CG solver with an external PETSc, and further calculations can be performed using this calculation path.

[0148] The beneficial effects that this application can produce are as follows.

[0149] (1) Intelligent pre-selection of calculation paths. By comprehensively considering the geometric structure complexity of the calculation task, the characteristics of the target physical model, and the available hardware resources, this application can intelligently pre-select a variety of potential efficient calculation paths.

[0150] (2) Resource optimization decision-making. This application can further calculate the pre-selected calculation paths to evaluate the resource consumption of each path. Based on these data, the path with the least resource consumption can be selected for actual calculation to ensure the maximization of calculation efficiency.

[0151] (3) Automated calculation path screening. The innovation of this application lies in the automated calculation path screening mechanism. Compared with manual selection, this application can automatically judge and select the optimal acceleration path according to the example requirements of different calculations. This automated process significantly improves the speed and accuracy of calculation path selection, reduces human errors, and ensures high performance in various calculation environments.

[0152] (4) Dynamic adaptive calculation strategy. This application has dynamic adaptability, can monitor the calculation process in real time, and dynamically adjust the calculation strategy according to the feedback data. When facing changing calculation conditions, this application can immediately optimize the calculation process and maintain the optimization of calculation efficiency.

[0153] (5) Advanced comprehensive evaluation algorithm. By comprehensively considering various factors such as calculation efficiency, resource consumption, and time cost, this application comprehensively evaluates different calculation paths, providing a scientific basis for subsequent calculation path decisions.

[0154] This application also includes an adaptive acceleration calculation system, including a memory and a processor. Among them, the memory is used to store instructions that can be executed by the processor; the processor is used to execute the instructions to implement the adaptive acceleration calculation method described above.

[0155] Exemplarily, the adaptive acceleration calculation system of this application adopts a modular architecture design, which can improve the maintainability and scalability of the system, ensuring that the system can flexibly adapt to the changes in technological development and user needs. Through the modular architecture, new functional modules and optimization algorithms can be easily integrated, thereby continuously improving the system performance and user experience. This application can not only significantly reduce the calculation time but also lower the usage threshold for users.

[0156] Figure 7 is the system block diagram of the adaptive acceleration calculation system of an embodiment of this application. Refer to Figure 7As shown, the adaptive acceleration computing system 700 may include an internal communication bus 701, a processor 702, a read-only memory (ROM) 703, a random access memory (RAM) 704, and a communication port 705. The adaptive acceleration computing system 700 may further include a hard disk 706. The internal communication bus 701 may enable data communication among the components of the adaptive acceleration computing system 700. The processor 702 may make judgments and issue prompts. In some embodiments, the processor 702 may be composed of one or more processors. The communication port 705 may enable data communication between the adaptive acceleration computing system 700 and the outside. In some embodiments, the adaptive acceleration computing system 700 may send and receive information and data from a network through the communication port 705. The adaptive acceleration computing system 700 may further include different forms of program storage units and data storage units, such as the hard disk 706, the read-only memory (ROM) 703, and the random access memory (RAM) 704, which can store various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor 702. The processor executes these instructions to implement the main part of the method. The results processed by the processor are transmitted to the user device through the communication port and displayed on the user interface.

[0157] The above-mentioned adaptive acceleration computing method may be implemented as a computer program, stored in the hard disk 706, and loaded into the processor 702 for execution to implement the adaptive acceleration computing method of the present application.

[0158] The present application further includes a computer-readable medium storing computer program code, which implements the adaptive acceleration computing method described above when executed by a processor.

[0159] When the adaptive acceleration computing method is implemented as a computer program, it may also be stored in a computer-readable storage medium as an article of manufacture. For example, the computer-readable storage medium may include, but is not limited to, magnetic storage devices (such as hard disks, floppy disks, magnetic strips), optical disks (such as compact discs (CDs), digital versatile discs (DVDs)), smart cards, and flash memory devices (such as electrically erasable programmable read-only memories (EEPROMs), cards, sticks, key drives). In addition, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) that can store, contain, and / or carry code and / or instructions and / or data.

[0160] It should be understood that the embodiments described above are merely illustrative. The embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For a hardware implementation, the processor can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and / or other electronic units designed to perform the functions described herein, or a combination thereof.

[0161] Some aspects of the present application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components", or "systems". The processor can be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or a combination thereof. In addition, aspects of the present application may be embodied as a computer product located on one or more computer readable media, which includes computer readable program code. For example, the computer readable media may include, but is not limited to, magnetic storage devices (such as hard disks, floppy disks, magnetic tapes...), optical discs (such as compact discs CD, digital versatile discs DVD...), smart cards, and flash memory devices (such as cards, sticks, key drives...).

[0162] The computer readable media may contain a propagated data signal having computer program code embodied therein, for example, on a baseband or as part of a carrier wave. The propagated signal may have various forms of manifestation, including electromagnetic form, optical form, etc., or a suitable combination thereof. The computer readable media can be any computer readable media other than a computer readable storage media, which can be connected to an instruction execution system, apparatus, or device to effect communication, propagation, or transmission for use of the program. The program code located on the computer readable media can be propagated through any suitable media, including radio, cable, fiber optic cable, radio frequency signal, or similar media, or any combination of the above media.

[0163] The basic concepts have been described above. Obviously, for those skilled in the art, the above application disclosure is merely an example and does not constitute a limitation to the present application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to the present application. Such modifications, improvements, and corrections are proposed in the present application, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of the present application.

[0164] Meanwhile, the present application uses specific terms to describe the embodiments of the present application. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present application can be appropriately combined.

[0165] In some embodiments, numbers are used to describe the components and the quantity of attributes. It should be understood that such numbers used to describe the embodiments are modified by the modifiers "about", "approximately", or "substantially" in some examples. Unless otherwise stated, "about", "approximately", or "substantially" indicate that the said numbers allow a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and such approximate values can be changed according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used to confirm the breadth of the scope in some embodiments of the present application are approximate values, in specific embodiments, such numerical settings are made as precise as possible within the feasible range.

Claims

1. An adaptive accelerated computing method, characterized in that: Software for numerical simulation, including: Step S1: analyzing the geometric structure and physical properties of the example model in the numerical simulation software; Step S2: performing mesh cutting and division on the geometric structure to obtain the number of meshes; Step S3: acquiring corresponding preset parameters according to the physical characteristics, and calculating the number of computing cores initially adopted by the example model according to the preset parameters and the number of grids; Step S4: selecting computing components of the algebraic matrix solving process according to the physical characteristics; Step S5: screening out a parallel acceleration library according to the physical characteristics; Step S6: generating a plurality of preliminary computing paths according to the number of computing cores, the computing components, and the parallel acceleration library initially taken; Step S7: selecting at least one final calculation path from the plurality of preliminary calculation paths.

2. The adaptive accelerated computing method according to claim 1, wherein: The geometric structure has a value in the X-axis direction, a value in the Y-axis direction, and a value in the Z-axis direction in the three-axis coordinate system OXYZ, and the X-axis, Y-axis, and Z-axis of the three-axis coordinate system OXYZ are perpendicular to each other; in the step S2, the geometric structure is meshed and divided, including: Step S2a: Construct a cutting relationship using the following formula: a*b*c=n Wherein, X represents the value in the X-axis direction; Y represents the value in the Y-axis direction; Z represents the value in the Z-axis direction; a represents the number of computing cores corresponding to the X-axis direction; b represents the number of computing cores corresponding to the Y-axis direction; c represents the number of computing cores corresponding to the Z-axis direction; n represents the number of all computing cores; Step S2b: According to the cutting relationship, the values ​​in the X-axis direction, the values ​​in the Y-axis direction, and the values ​​in the Z-axis direction are equally divided.

3. The adaptive accelerated computing method according to claim 1, wherein: The preset parameters include: a performance coefficient of the computing core and a physical coefficient of the physical property; in the step S3, the number of computing cores initially adopted by the example model is calculated according to the preset parameters and the number of grids, including: using the following formula to calculate the number of computing cores initially adopted: core=NG / (β*p / α) Among them, core represents the number of computing cores initially adopted; NG represents the number of grids; β represents the physical coefficient of the physical property; α represents the performance coefficient of the computing core; and p represents the frequency of the computing core.

4. The adaptive accelerated computing method according to claim 1, wherein: In step S7, selecting at least one final calculation path from the plurality of preliminary calculation paths includes: Step S7a: Calculating the optimal resource consumption value according to the multiple preliminary calculation paths; Step S7b: taking the preliminary calculation path corresponding to the optimal resource consumption value as the final calculation path.

5. The adaptive accelerated computing method according to claim 4, characterized in that: The optimal resource consumption value is calculated using the following formula: v=min(n i *t i ) Wherein, v represents the optimal resource consumption value; min(·) represents the minimum function; i represents the number of preliminary calculation paths; n i represents the number of computing cores of the i-th preliminary computing path; t i Represents the computation time of the i-th preliminary computation path.

6. The adaptive accelerated computing method according to claim 4 or 5, characterized in that: After step S7a and before step S7b, the method further includes: determining whether the optimal resource consumption value is greater than a preset threshold; if so, executing step S4; if not, executing step S7b.

7. The adaptive accelerated computing method according to claim 1, wherein: The numerical simulation software includes CFD software; the computing components include: a preprocessor, a solver, a smoother, and an external library solver, the preprocessor includes one or any combination of Diagonal, DIC, FDIC, GAMG, and DILU, the solver includes one or any combination of Diagonal, PCG, PBiCG, GAMG, and CG, the smoother includes Dic GaussSeidel or Diabetesgonal, and the external library solver includes one or any combination of GMRES, Richardson, CG, and CGS; the parallel acceleration library includes one or any combination of PETSc, AmgX, SUNDIALS, Hypre, SuperLU, and MUMPS.

8. The adaptive accelerated computing method according to claim 1, wherein: The physical properties include: compressibility, incompressibility, multiphase flow, heat transfer, mass transfer, chemical reaction, or any combination thereof.

9. An adaptive accelerated computing system, characterized in that: include: a memory for storing instructions executable by a processor; A processor, configured to execute the instructions to implement the adaptive accelerated computing method as described in any one of claims 1-8.

10. A computer readable medium storing computer program code, characterized in that: When executed by a processor, the computer program code implements the adaptive accelerated computing method as described in any one of claims 1 to 8.