Isomerism accelerated high-efficiency analysis method for ultimate bearing of rocket body thin-wall structure
By using GPU heterogeneous parallel technology and asynchronous transmission and double buffering methods in the ultimate bearing analysis of the thin-wall structure of arrow body, the problems of large computing resources and long computing time in the existing technology are solved, and efficient dynamic calculations and fast design adaptability are achieved.
Patent Information
- Application Number
- CN202411740662.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The existing explicit dynamic analysis methods use large computing resources and long calculation time when dealing with the ultimate bearing of thin-walled structures of arrow bodies, making it difficult to meet the complex and changeable design needs.
Using GPU heterogeneous parallel technology, the image model of the thin-wall structure of the arrow body is meshed and searched, dynamic calculation is performed using the GPU calculation unit, and the calculation results are output through asynchronous transmission and double buffering methods.
It significantly accelerates the explicit dynamic analysis process, improves computational efficiency, and can adapt to complex and variable design needs more quickly, while maintaining the accuracy of numerical simulations.
Smart Images

Figure CN119962279A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a high-efficiency analysis method for extreme load-bearing heterogeneous acceleration of a thin-walled structure of a rocket body, and belongs to the technical field of high-performance numerical simulation analysis. Background Art
[0002] As a key part of the main load-bearing structure of the launch vehicle, the thin-walled structure of the rocket body often suffers from structural instability due to the axial compression load caused by the launch overload. When solving the ultimate bearing capacity of such structures, explicit dynamics simulation of quasi-static processes is often used to solve the non-convergence problems caused by strong nonlinearities such as large deformation, material nonlinearity, and complex contact. However, the existing explicit dynamics analysis methods face challenges such as high consumption of computing resources and long calculation time, and are difficult to adapt to complex and changing design requirements. The commercial software Abaqus explicit dynamics is limited by computational efficiency and is not efficient enough when processing large-scale models. It cannot meet the needs of rapid product iteration and urgently needs to use high-performance computing resources such as GPU and FPGA to improve the analysis speed. Summary of the invention
[0003] The technical problem solved by the present invention is: to overcome the deficiencies of the prior art, to provide an efficient analysis method for the extreme load heterogeneous acceleration of a thin-walled structure of a rocket body, and to achieve acceleration of the explicit dynamics analysis process.
[0004] The technical solution of the present invention is:
[0005] The present invention discloses a method for efficient analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure, comprising:
[0006] Mesh the image model of the rocket body thin-wall structure to obtain several mesh units and nodes;
[0007] According to a number of grid units, the image model is searched and divided to obtain a number of unit combinations;
[0008] Allocating the unit combination to GPU computing units according to a GPU load balancing strategy;
[0009] According to the central difference method, a GPU computing unit is used to perform dynamic calculations on the nodes, units and unit combinations to obtain calculation results;
[0010] The calculation result is output to the external CPU by using asynchronous transmission and double buffering method.
[0011] Furthermore, in the above method, the image model is searched and divided, and the specific method is:
[0012] The combinations of units are divided into combination a, combination b, combination c and combination d; wherein combination a is composed of 2*2*2 units; combination b is composed of 2*2 units; combination c is composed of 2*1 units; combination d is composed of one unit;
[0013] According to combination a, combination b, combination c and combination d, the priorities are A, B, C and D respectively, where A>B>C>D;
[0014] The image model is searched and divided according to the priority from large to small to obtain several unit combinations.
[0015] Furthermore, in the above method, the GPU load balancing strategy is specifically:
[0016] The unit combinations of type combination a, combination b, combination c and combination d are distributed evenly to several GPUs in turn.
[0017] Furthermore, in the above method, the GPU computing unit performs dynamics calculation on the nodes, units and unit combinations, and the specific method is:
[0018] S40, initializing time step and calculation time;
[0019] S41. Calculation The speed and t n+1 Time displacement;
[0020] S42, calculate t n+1 Internal force at all times external force and the next time step;
[0021] S43, according to t n+1 Internal force at all times external force and the next time step, calculate t n+1 Acceleration at time t n+1 Time speed;
[0022] S44, accumulating the calculation time; judging whether the calculation time exceeds the threshold, if so, outputting the displacement, velocity, acceleration and internal force of the node and exiting the calculation; if not, entering step S45;
[0023] S45. Repeat steps S41 to S44 according to the next time step.
[0024] Furthermore, in the above method, the specific method for calculating the next time step is as follows:
[0025] t n+1 =t n +Δt n+1
[0026]
[0027] Among them, Δt n+1 is the time step, t n and t n+1 is the time corresponding to the execution of step n, It is the moment between n and n+1.
[0028] Furthermore, in the above method, the calculation The speed and t n+1 The time displacement is:
[0029]
[0030] in, for The speed of time, t n The speed of time, t n The acceleration of the moment; is the middle moment of the time step interval; t n Displacement of time.
[0031] Furthermore, in the above method, the calculation t n+1 Acceleration at time t n+1 The speed at the moment is:
[0032]
[0033] in, are the velocity and acceleration at time n+1, c i For damping, are the external and internal force values, respectively. for Time speed.
[0034] The beneficial effects of the present invention and the prior art are:
[0035] (1) The present invention utilizes GPU heterogeneous parallel technology to achieve efficient acceleration of the explicit dynamics analysis process, thereby solving the problem that the explicit dynamics algorithm implemented on a single processor in the prior art is difficult to meet the requirements of the ultimate bearing capacity of thin-walled structures on the computational efficiency;
[0036] (2) The present invention proposes a heterogeneous accelerated and efficient analysis method for the ultimate load of a rocket body thin-wall structure. The method fully utilizes the multi-core characteristics of the GPU to calculate such independent tasks with large amounts of data using GPU threads. The efficiency of explicit dynamics calculations is significantly improved while maintaining the accuracy of numerical simulations. This method solves the problems faced by existing explicit dynamics methods in analyzing rocket body thin-wall structures, such as large consumption of computational resources and long calculation time. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is the overall acceleration optimization framework diagram of the structural explicit simulation parallel algorithm of the present invention;
[0038] Figure 2 It is an execution flow chart of the structural explicit simulation algorithm of the present invention;
[0039] Figure 3 1 is a preprocessing partition diagram of the present invention; (a) is a combination of eight units, (b) is a combination of four units, (c) is a combination of two units, and (d) is a combination of one unit;
[0040] Figure 4 It is a multi-GPU load balancing strategy diagram of the present invention;
[0041] Figure 5 It is a schematic diagram of asynchronous transmission and double buffering of the present invention. DETAILED DESCRIPTION
[0042] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0043] The present invention discloses a method for efficient analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure, comprising:
[0044] Mesh the image model of the rocket body thin-wall structure to obtain several mesh units and nodes;
[0045] According to a number of grid units, the image model is searched and divided to obtain a number of unit combinations;
[0046] According to the GPU load balancing strategy, the unit combination is allocated to the GPU computing unit;
[0047] According to the central difference method, the GPU computing unit is used to perform dynamic calculations on nodes, units and unit combinations to obtain calculation results;
[0048] Asynchronous transmission and double buffering methods are used to output the calculation results to the external CPU.
[0049] Preferably, the image model is searched and divided, and the specific method is:
[0050] The combinations of units are divided into combination a, combination b, combination c and combination d; wherein combination a is composed of 2*2*2 units; combination b is composed of 2*2 units; combination c is composed of 2*1 units; combination d is composed of one unit;
[0051] According to combination a, combination b, combination c and combination d, the priorities are A, B, C and D respectively, where A>B>C>D;
[0052] According to the priority from large to small, the image model is searched and divided to obtain several unit combinations. Preferably, the GPU load balancing strategy is specifically as follows:
[0053] The unit combinations of type combination a, combination b, combination c and combination d are distributed evenly to several GPUs in turn.
[0054] Preferably, if Figure 2 As shown, the GPU computing unit performs dynamic calculations on nodes, units, and unit combinations. The specific method is as follows:
[0055] S40, initializing time step and calculation time;
[0056] S41. Calculation The speed and t n+1 Time displacement;
[0057] S42, calculate t n+1 Internal force at all times external force and the next time step;
[0058] S43, according to t n+1 Internal force at all times external force and the next time step, calculate t n+1 Acceleration at time t n+1 Time speed;
[0059] S44, accumulating the calculation time; judging whether the calculation time exceeds the threshold, if so, outputting the displacement, velocity, acceleration and internal force of the node and exiting the calculation; if not, entering step S45;
[0060] S45. Repeat steps S41 to S44 according to the next time step.
[0061] Preferably, the next time step is calculated by:
[0062] t n+1 =t n +Δt n+1
[0063]
[0064] Among them, Δt n+1 is the time step, t n and t n+1 is the time corresponding to the execution of step n, It is the moment between n and n+1.
[0065] Preferably, the calculation The speed and t n+1 The time displacement is:
[0066]
[0067] in, for The speed of time, t n The speed of time, t n The acceleration of the moment; is the middle moment of the time step interval; t n Displacement of time.
[0068] Preferably, calculate t n+1 Acceleration at time t n+1 The speed at the moment is:
[0069]
[0070] in, are the velocity and acceleration at time n+1, c i For damping, are the external and internal force values, respectively. for Time speed.
[0071] Example
[0072] The present invention provides a GPU-based explicit dynamics heterogeneous computing framework, such as Figure 1 As shown in the figure, under the CUDA architecture, a single CUDA thread is used to correspond to a single computing block, and the memory access method and computing process are optimized based on the GPU storage structure to complete the solution process of the structural explicit simulation algorithm.
[0073] A unit preprocessing algorithm in parallel computing is designed to improve reuse and reduce memory access overhead by optimizing unit combination and allocation. The unit preprocessing algorithm allocates unit combinations into four basic situations according to node reuse, such as Figure 3As shown, the unit combination with high node reuse is preferentially assigned to the same parallel thread block. Taking the first-order hexahedral mesh division of three-dimensional space as an example, when calculating the internal force of each unit and the next step, it is necessary to read the data of the eight nodes on the unit. However, part of the data required by adjacent units is repeated, and two adjacent units will share the data of four nodes on the same face. In the CUDA storage structure, each thread in the same parallel thread block can share data through the shared memory storage structure. Therefore, the grid units processed in a block can be adjusted to be placed as close to each other as possible, share more nodes, reduce the number of nodes required to calculate these grid units, and thus reduce the memory access overhead. In order to assign appropriate grid units to each parallel thread block and make the units calculated in a parallel thread block share as many nodes as possible, it is necessary to preprocess the data before parallel calculation. Since the relative position between units does not change during the deformation of the entire simulated object, only one preprocessing is required before all parallel calculations, and subsequent parallel operations can be calculated according to the partitioning scheme obtained by the preprocessing. For this preprocessing algorithm, in order to reduce its algorithm complexity, it can also better complete the partitioning task. Design a greedy unit partitioning preprocessing algorithm. Specifically, the preprocessing algorithm divides the unit combination into four basic combinations according to the reuse of nodes, such as Figure 3 As shown. Combination a is composed of 2*2*2 units, and each unit only requires 27 / 8=3.375 nodes on average, which is the combination with the highest node reuse; combination b is composed of 2*2 units, and each unit requires 18 / 4=4.5 nodes on average, which is the second highest node reuse combination; combination c is composed of 2*1 units, and each unit requires 12 / 2=6 nodes on average; combination d has only one unit, and each unit requires 8 nodes on average.
[0074] Based on the multi-GPU load balancing strategy, for the multi-GPU environment, in the preprocessing scheme, since the units processed by each thread block are as adjacent as possible, and the units processed by adjacent thread blocks share more units than other thread blocks, the threads in the block can be evenly scheduled to different GPU devices in the order of allocation, so that the finite element units that each GPU needs to process are as adjacent as possible and concentrated in a certain area, while the finite element units that need to be processed by different GPUs are separated as much as possible, so as to realize the multi-GPU task load balancing scheduling strategy, such as Figure 4 shown.
[0075] The data transmission strategy based on asynchronous transmission and double buffering technology, asynchronous transmission allows other computing tasks to be performed while data is transmitted; using double buffering technology, the strategy of alternating the use of two buffers is used. During the data transmission process, one buffer is used for computing and the other is used for data transmission. When the calculation is completed, the roles of the two buffers are swapped to continue the next round of computing and data transmission. Double buffering technology can not only make the computing and data transmission tasks better concurrently executed, but also effectively solve the competition problem between computing and data transmission. While the previous round of calculation is still in progress, by using another buffer for data transmission, the problem of data consistency is avoided, and the correctness of asynchronously transmitted data is further guaranteed. In addition, double buffering technology can also improve the robustness of the system and reduce computing failures caused by data transmission errors. In summary, through the comprehensive application of data compression, asynchronous transmission and double buffering technology, the data transmission overhead between GPU and CPU can be significantly reduced, computing efficiency can be improved, and data consistency and correctness can be ensured, thereby achieving more efficient distributed computing in large-scale explicit finite element calculations, such as Figure 5 shown.
[0076] For the existing explicit dynamics analysis method based on the central difference method, a parallel strategy based on GPU (Graphics Processing Unit) can be constructed to solve the technical problem of low efficiency of explicit dynamics simulation, as follows:
[0077] See also Figure 1 , Figure 1 The GPU-based parallel explicit dynamics calculation framework provided for this application example has the following specific steps:
[0078] S101, preprocess the input model and prepare the data.
[0079] S102, construction of explicit dynamics parallel computing framework based on central difference method.
[0080] Step 1: Preprocess the input model:
[0081] Taking a three-dimensional first-order hexahedral grid cell as an example, in parallel computing with cells as the granularity, each thread needs to obtain the physical quantities of the eight nodes of the cell, calculate the internal force of the cell and the next time step, but some nodes of adjacent cells are repeated. In the CUDA storage structure of this application example, each thread in the same parallel thread block shares data through shared memory. The cell preprocessing algorithm divides adjacent cells into the same parallel thread block so that the cells it processes share more nodes, thereby reducing memory access overhead and improving execution efficiency. The execution flow of the preprocessing algorithm is described as follows:
[0082] The model data is preprocessed before parallel computing, and the unit combinations are allocated according to the node reuse situation. Figure 3 Four basic cases, for each unit combination, the division obtained Figure 3 (a) The more, the higher the reuse degree of the unit combination node; Figure 3 (d) The more, the lower the node reuse of the unit combination. In order to maximize the node reuse, when the preprocessing algorithm divides the unit for the current parallel block, it first checks whether there are adjacent units that have not been divided yet at the boundary of the unit that has been divided to the parallel thread block. Figure 3 If there is a unit combination of (a), add it to the parallel thread block calculation range and repeat this step. If not, find out whether there is a unit combination of (a). Figure 3 (b), and so on, until the unit division of the parallel thread block is completed.
[0083] Based on the above unit preprocessing algorithm, the preprocessing stage implements a multi-GPU load balancing strategy; based on the multi-GPU load balancing strategy, for the multi-GPU environment, in the preprocessing solution, since the units processed by each thread block are as adjacent as possible, and the units processed by adjacent thread blocks share more units than other thread blocks. Therefore, the threads in the block can be evenly scheduled to different GPU devices in the order of allocation, so that the finite element units that each GPU needs to process are as adjacent as possible and concentrated in a certain area, while the finite element units that need to be processed by different GPUs are separated as much as possible, thereby realizing a multi-GPU task load balancing scheduling strategy, such as Figure 5 shown.
[0084] Step 2: Construction of explicit dynamics heterogeneous computing framework based on central difference method:
[0085] When performing explicit dynamics calculations based on the central difference method, in one time step calculation iteration, the velocity at the next half moment, the displacement at the next moment, the force at the current moment, the velocity at the next moment, and the acceleration will be updated in sequence. Therefore, the node-based kernel function is designed to calculate the node's displacement, velocity, and acceleration; the unit-based kernel function is designed to calculate the unit's overall force (internal force) and the next moment step.
[0086] 1) Initialization: Divide the solution area into finite elements and calculate according to the initial boundary conditions and
[0087] 2) Update time:
[0088] t n+1 =t n +Δt n+1
[0089]
[0090] 3) Calculation The speed and t n+1 Time displacement
[0091]
[0092] 4) Calculate t n+1 Internal force at all times external force and the next time step
[0093] 5) Calculate t n+1 Acceleration at time t n+1 Time speed:
[0094]
[0095] 6) If t n+1 =T, that is, the current moment is the end of the solution time interval, then the loop ends, otherwise jumps back to step 2) to continue iterative calculation.
[0096] In the calculation iteration of a time step, it is necessary to first perform the calculation at the node granularity (calculate displacement and velocity), then perform the calculation at the unit granularity (calculate internal force and the next time step), and finally perform the calculation at the node granularity (calculate velocity and acceleration). From this perspective, if the kernel function is designed with nodes and units as parallel granularities respectively, the kernel function needs to be called three times for each calculation iteration of the time step. However, when the calculation iterations of two time steps do not involve data transmission from the device to the host, the node granularity calculation of the previous time step and the node granularity calculation of the next time step can be combined in one kernel function. Therefore, the calculation iteration of each time step can reduce the overhead of starting the kernel function once.
[0097] In the kernel function with unit as parallel granularity, each parallel thread block reads the node data corresponding to the unit partitioning scheme obtained by the preprocessing algorithm into the shared memory unique to the parallel block. Then each thread in the parallel block obtains the node data corresponding to the unit from the shared memory according to the sequence number for calculation.
[0098] After completing a specific round of computation iterations, the intermediate process data needs to be output, so the data in the GPU memory is unloaded to the CPU host memory. Figure 5 Use asynchronous transmission and double buffering technology. Using asynchronous transmission while data transmission allows other tasks to be performed without waiting for the transmission to complete, and double buffering technology is used to alternately use two buffers, one for calculation and the other for data transmission during the data transmission process; when the current round of calculation is completed, the two buffers swap roles and continue with the next round of calculation and data transmission.
[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. The above description is only a preferred embodiment of the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0100] Although the content of the present invention has been described in detail through the above preferred embodiments, it should be appreciated that the above description should not be considered as a limitation of the present invention. After reading the above content, it will be apparent to those skilled in the art that various modifications and substitutions of the present invention will occur. Therefore, the protection scope of the present invention should be limited by the appended claims.
[0101] The contents not described in detail in the specification of the present invention belong to the common knowledge of the professionals in this field.
Claims
1. A highly efficient analysis method for the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure, characterized in that: include: Mesh the image model of the rocket body thin-wall structure to obtain several mesh units and nodes; According to a number of grid units, the image model is searched and divided to obtain a number of unit combinations; Allocating the unit combination to GPU computing units according to a GPU load balancing strategy; According to the central difference method, a GPU computing unit is used to perform dynamic calculations on the nodes, units and unit combinations to obtain calculation results; The calculation result is output to the external CPU by using asynchronous transmission and double buffering method.
2. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 1, characterized in that: The specific method for searching and dividing the image model is as follows: The combinations of units are divided into combination a, combination b, combination c and combination d; wherein combination a is composed of 2*2*2 units; combination b is composed of 2*2 units; combination c is composed of 2*1 units; combination d is composed of one unit; According to combination a, combination b, combination c and combination d, the priorities are A, B, C and D respectively, where A>B>C>D; The image model is searched and divided according to the priority from large to small to obtain several unit combinations.
3. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 2, characterized in that: The GPU load balancing strategy is specifically as follows: The unit combinations of type combination a, combination b, combination c and combination d are distributed evenly to several GPUs in turn.
4. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 1, characterized in that: The GPU computing unit performs dynamics calculation on the nodes, units and unit combinations, and the specific method is as follows: S40, initializing time step and calculation time; S41. Calculation The speed and t n+1 Time displacement; S42, calculate t n+1 Internal force at all times external force and the next time step; S43, according to t n+1 Internal force at all times external force and the next time step, calculate t n+1 Acceleration and t n+1 Time speed; S44, accumulating the calculation time; judging whether the calculation time exceeds the threshold, if so, outputting the displacement, velocity, acceleration and internal force of the node and exiting the calculation; if not, entering step S45; S45. Repeat steps S41 to S44 according to the next time step.
5. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 4, characterized in that: The specific method for calculating the next time step is: t n+1 =t n +Δt n+1 Among them, Δt n+1 is the time step, t n and t n+1 is the time corresponding to the execution of step n, It is the moment between n and n+1.
6. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 4, characterized in that: The calculation The speed and t n+1 The time displacement is: in, for The speed of time, t n The speed of time, t n The acceleration of the moment; is the middle moment of the time step interval; t n Displacement of time.
7. The method for high-efficiency analysis of the ultimate load-bearing heterogeneous acceleration of a rocket thin-wall structure according to claim 4, characterized in that: The calculation t n+1 Acceleration and t n+1 The speed at the moment is: in, are the velocity and acceleration at time n+1, c i For damping, are the external and internal force values, respectively. for Time speed.
Citation Information
Patent Citations
Light tracing parallel optimization method based on Intel many-core framework
CN104700447A
Structural-dynamic-analysis explicit-different-step-length parallel computing method
CN108228970A
High-performance encryption and decryption method based on PCI-e channel
CN112035388A
Dynamic task scheduling algorithm for real-time requirements of cloud computing system
CN113220428A
Fluid-solid coupling method for flexible multi-body thin-wall structure
CN115392148A