Sea clutter task-level simulation optimization method, medium and program product
By dividing sea clutter computing tasks to different levels and combining CUDA multi-stream processing technology, the problem of task-level data dependence limitation in the existing technology is solved, the efficiency of sea clutter simulation is improved, and the high-performance computing needs of modern radar systems are met.
Patent Information
- Application Number
- CN202510438805.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
The existing sea clutter simulation methods are difficult to meet the high-performance requirements of modern radar systems for computing efficiency due to the data dependence between task levels on GPUs, resulting in serial tasks in the computing process and the GPU's parallel processing capabilities cannot be fully utilized.
By dividing sea clutter computing tasks to different levels, using CUDA multi-stream processing technology to realize task-level parallel execution, and optimizing task division using topological sorting algorithms. There is no data dependence between computing tasks at the same level, and task calculations are performed in combination with CUDA parallel computing framework.
It improves the efficiency of sea clutter simulation, realizes rapid sea clutter simulation, provides theoretical and technical support for sea clutter simulation, and improves computing efficiency.
Smart Images

Figure CN120337554A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sea clutter mission simulation, and particularly relates to a sea clutter mission-level simulation optimization method, medium and program product. Background Art
[0002] During the process of a radar system transmitting electromagnetic waves to the sea surface, the echo signal obtained by the receiving end has multi-source composite characteristics. Specifically, this echo signal not only contains the effective information reflected by the target, but also is mixed with noise, random interference signals introduced by the external environment, and sea clutter signals generated by sea surface reflections from different directions. These sea clutter signals will seriously interfere with the radar's target detection and tracking performance. Therefore, establishing a refined sea clutter echo model is of great significance for improving the radar's target recognition ability and anti-interference performance.
[0003] In the latest radar system, due to the high temporal, spatial and frequency randomness of sea clutter, traditional sea clutter simulation methods usually adopt data-level multi-threaded parallel computing on the GPU. However, in the simulation process, due to the data dependence between tasks, it is difficult to meet the high-performance requirements of modern radar systems for computing efficiency.
[0004] The existing sea clutter parallel simulation method is to use CPU-GPU collaborative computing to make full use of the parallel processing ability of the GPU to accelerate computationally intensive tasks. The CPU is responsible for initializing data, calculating relevant parameters required for clutter simulation, initializing the GPU device and allocating video memory space, while the GPU parallelly processes the FFT of the chirp signal, the parameter calculation of scattering units (such as elevation angle, azimuth angle, scattering coefficient and antenna gain), and the accumulation of clutter amplitude. During the simulation process, the calculation of each scattering unit and the generation of clutter echoes are both executed in parallel by the GPU. Finally, the generated data is transferred to the CPU side through memory copy to complete the simulation of sea clutter. The above method realizes parallel execution at the data level through GPU multi-threaded processing, and controls multiple threads through one instruction to complete the calculation of different data. However, when there are multiple tasks without data dependence in the kernel function (such as the calculation of scattering coefficient and the calculation of antenna gain), although there is no data dependence between these two tasks, since the thread can only process one calculation task at a certain moment and the tasks have a sequential order, the calculation of multiple tasks inside the kernel function is serial, and one task needs to be processed before starting to process the next task. Therefore, for the design of the kernel function for clutter calculation, the calculation instruction of the scattering coefficient is before the calculation instruction of the antenna gain. This means that when starting the calculation of the scattering coefficient at a certain moment, it is impossible to start the calculation of the antenna gain at the same time. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a method, medium and program product for optimizing sea clutter task-level simulation. By dividing each computing task into different levels, there is no data dependence relationship between the computing tasks at the same level, and combined with the CUDA multi-stream processing technology, parallel execution at the task level is achieved, effectively improving the sea clutter simulation efficiency and providing theoretical and technical support for rapid sea clutter simulation.
[0006] In order to achieve the above technical objectives and technical effects, the present invention is realized through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for optimizing sea clutter task-level simulation, which is applied to a CPU host and includes:
[0008] Obtain a data flow graph of sea clutter computing tasks pre-constructed, where the nodes of the data flow graph represent computing tasks, and the directed edges represent the data dependence relationship between computing tasks;
[0009] Perform hierarchical processing on the data flow graph, divide each computing task into different levels, and there is no data dependence relationship between the computing tasks at the same level;
[0010] According to the hierarchical order, perform task calculations layer by layer based on the CUDA parallel computing framework.
[0011] In combination with the first aspect, optionally, when performing hierarchical processing, the topological sorting algorithm is adopted.
[0012] In combination with the first aspect, optionally, the dividing each computing task into different levels specifically includes:
[0013] Count the in-degree of each node in the data flow graph;
[0014] Add all nodes with an in-degree of 0 to the queue;
[0015] Take out nodes from the queue in turn, add them to the corresponding layer, and at the same time traverse all neighbor nodes pointed to by this node. If the in-degree of the neighbor node becomes 0, add it to the queue and add the neighbor node to the corresponding layer;
[0016] Record the hierarchical result after topological sorting, obtain the hierarchical order of each node, and complete the division of each computing task into different levels.
[0017] In combination with the first aspect, optionally, the performing task calculations layer by layer based on the CUDA parallel computing framework includes:
[0018] If the number of computing tasks in a certain layer is equal to 1, directly send it to the GPU for processing;
[0019] If the number of computing tasks in a certain layer is greater than 1, then the same number of CUDA streams are allocated according to the number of computing tasks, and the computable parallel computing tasks are allocated to different CUDA streams for execution.
[0020] Combined with the first aspect, optionally, the GPU processes the computing tasks by the following steps:
[0021] Start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel computing of the corresponding computing tasks according to the indexes;
[0022] Save the calculation results and transfer them to the CPU host through memory copy.
[0023] Combined with the first aspect, optionally, the step of allocating the same number of CUDA streams according to the number of computing tasks and allocating the computable parallel computing tasks to different CUDA streams for execution includes:
[0024] Allocate the corresponding number of CUDA streams in the CPU host according to the number of computing tasks, initialize the GPU, allocate video memory space, and set the corresponding number of thread blocks and threads;
[0025] Control the GPU to start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel computing of the corresponding computing tasks according to the indexes;
[0026] Save the calculation results and transfer them to the CPU host through memory copy.
[0027] Combined with the first aspect, optionally, the step of hierarchically processing the data flow graph and dividing each computing task into different levels includes:
[0028] Divide the data flow graph into four layers;
[0029] The first layer is responsible for calculating the slant range between the beam center and the radar;
[0030] The second layer is responsible for calculating the horizontal distance between the beam center and the radar, intermediate variables, and grazing angle;
[0031] The third layer is responsible for calculating the clutter cell block area, clutter Doppler frequency, clutter cell block backscattering coefficient, and radar antenna gain;
[0032] The fourth layer is responsible for calculating the clutter modulation coefficient matrix.
[0033] Combined with the first aspect, optionally, the computing tasks of the first layer and the fourth layer are processed in parallel by the multi-threading of the GPU; the computable parallel computing tasks in the second layer and the third layer are allocated to different CUDA streams for execution.
[0034] In a second aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the sea clutter task-level simulation optimization method described in any one of the first aspects.
[0035] In a third aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the sea clutter task-level simulation optimization method described in any one of the first aspects.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] The present invention proposes a sea clutter task-level simulation optimization method, medium and program product. By dividing each computing task into different levels, there is no data dependency between the computing tasks at the same level, and combined with the CUDA multi-stream processing technology, parallel execution at the task level is achieved, effectively improving the sea clutter simulation efficiency and providing theoretical and technical support for fast sea clutter simulation. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:
[0039] Figure 1 It is a schematic diagram of the overall process of the sea clutter task-level simulation optimization method according to an embodiment of the present invention;
[0040] Figure 2(a) is a data flow diagram according to an embodiment of the present invention;
[0041] Figure 2(b) is a schematic diagram of reasonable task division according to an embodiment of the present invention;
[0042] Figure 2(c) is a schematic diagram of unreasonable task division according to an embodiment of the present invention;
[0043] Figure 3 It is a data flow diagram of the calculation of sea clutter simulation-related tasks according to an embodiment of the present invention;
[0044] Figure 4 It is a flow chart of the topological sorting algorithm according to an embodiment of the invention;
[0045] Figure 5 It is a topological sorting hierarchical result diagram according to an embodiment of the invention;
[0046] Figure 6(a) is a schematic diagram of the comparison of the time-consuming of sea clutter simulation calculation according to an embodiment of the present invention;
[0047] Figure 6(b) is a schematic diagram of sea clutter probability density estimation according to an embodiment of the present invention;
[0048] Figure 6(c) is a schematic diagram of sea clutter power spectrum estimation according to an embodiment of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, such descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0051] Embodiment 1
[0052] An embodiment of the present invention provides a sea clutter task-level simulation optimization method, which is applied to a CPU host and specifically includes the following steps:
[0053] (1) Obtain a data flow graph of a pre-constructed sea clutter calculation task, where the nodes of the data flow graph represent calculation tasks, and the directed edges represent the data dependency relationships between the calculation tasks;
[0054] (2) Perform hierarchical processing on the data flow graph to divide each calculation task into different levels, and there is no data dependency relationship between the calculation tasks at the same level;
[0055] (3) Calculate tasks layer by layer based on the CUDA parallel computing framework in the order of levels.
[0056] In a specific implementation manner of the embodiment of the present invention, when performing hierarchical processing, the topological sorting algorithm is adopted. Specifically, the dividing of each calculation task into different levels specifically includes:
[0057] Count the in-degree of each node in the data flow graph;
[0058] Add all nodes with an in-degree of 0 to the queue;
[0059] Take out nodes from the queue one by one, add them to the corresponding layer, and at the same time traverse all neighbor nodes pointed to by this node. If the in-degree of a neighbor node becomes 0, add it to the queue and add the neighbor node to the corresponding layer;
[0060] Record the layered result after topological sorting, obtain the hierarchical order of each node, and complete the division of each computing task into different layers.
[0061] In a specific implementation manner of the embodiment of the present invention, the task calculation based on the CUDA parallel computing framework layer by layer according to the hierarchical order includes:
[0062] If the number of computing tasks in a certain layer is equal to 1, directly send it to the GPU for processing;
[0063] If the number of computing tasks in a certain layer is greater than 1, allocate the same number of CUDA streams according to the number of computing tasks, and allocate the computable parallel computing tasks to different CUDA streams for execution.
[0064] In a specific implementation manner of the embodiment of the present invention, the GPU processes the computing tasks using the following steps:
[0065] Start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel computing of the corresponding computing tasks according to the index;
[0066] Save the calculation result and transfer it to the CPU host through memory copy.
[0067] In a specific implementation manner of the embodiment of the present invention, the step of allocating the same number of CUDA streams according to the number of computing tasks and allocating the computable parallel computing tasks to different CUDA streams for execution includes:
[0068] Allocate the corresponding number of CUDA streams in the CPU host according to the number of computing tasks, initialize the GPU, allocate video memory space, and set the corresponding thread blocks and the number of threads;
[0069] Control the GPU to start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel computing of the corresponding computing tasks according to the index;
[0070] Save the calculation result and transfer it to the CPU host through memory copy.
[0071] In a specific implementation manner of the embodiment of the present invention, the step of performing layered processing on the data flow graph and dividing each computing task into different layers includes:
[0072] Divide the data flow graph into four layers;
[0073] The first layer is responsible for calculating the slant range between the beam center and the radar;
[0074] The second layer is responsible for calculating the horizontal distance between the beam center and the radar, intermediate variables, and grazing angle;
[0075] The third layer is responsible for calculating the clutter cell block area, clutter Doppler frequency, clutter cell block backscattering coefficient, and radar antenna gain;
[0076] The fourth layer is responsible for calculating the clutter modulation coefficient matrix.
[0077] In a specific implementation manner of the embodiment of the present invention, the calculation tasks of the first layer and the fourth layer are processed in parallel through multiple threads of the GPU; the calculation tasks that can be executed in parallel in the second layer and the third layer are assigned to different CUDA streams for execution.
[0078] The following details the sea clutter task-level simulation optimization method in the embodiment of the present invention in conjunction with a specific implementation manner.
[0079] As Figure 1 shown, the sea clutter task-level simulation optimization method specifically includes the following steps:
[0080] 1. Parameter initialization.
[0081] The transmit signal waveform in the embodiment of the present invention is a linear frequency modulation signal, and the antenna pattern model is a sinc function model. Specific radar parameters (transmit signal frequency, number of pulses, number of azimuth gates, etc.), sea surface environment parameters (sea surface movement speed, sea state level, etc.), and sea clutter modeling parameters (sea clutter backscattering coefficient model, amplitude distribution model, power spectrum, etc.) can be flexibly and reasonably set according to actual simulation requirements, with strong scalability.
[0082] These parameters are used for the sea clutter calculation task.
[0083] 2. Hierarchically process the sea clutter data flow graph based on the topological sorting hierarchical algorithm. By constructing a task dependency graph (DAG) and the topological sorting hierarchical algorithm, each calculation task is divided into a hierarchical structure with independent execution windows, requiring that there is no data dependency between the calculation tasks at the same level and supporting CUDA multi-stream parallel computing. As Figure 1 shown, the first layer includes tasks; the second layer includes tasks;...; the Xth layer includes tasks;
[0084] 3. Complete CUDA multi-stream parallel computing layer by layer. For each layer of calculation tasks:
[0085] (1)In the CPU host, allocate corresponding CUDA streams according to the number of tasks, initialize the GPU, allocate video memory space, and set the corresponding thread blocks and the number of threads.
[0086] (2)The GPU host completes the computing tasks, starts the corresponding kernel function, calculates the thread index and the thread block index, and completes the parallel computing of the corresponding computing tasks according to the indexes.
[0087] (3)Save the calculation results and transfer them to the CPU host through memory copy.
[0088] 4. After all computing tasks are completed, process and analyze the obtained data, compare it with the original sea clutter simulation data, and test the speedup ratio of the CUDA multi-stream parallel computing simulation method based on the topological sorting hierarchical algorithm relative to the original simulation method.
[0089] In the present invention, in the form of a data flow graph, the dependency relationship and data transmission between tasks can be intuitively displayed, which helps to understand the optimization strategy of task partitioning in the algorithm. The data flow graph can be used to effectively describe the computing process between tasks. Specifically, the data flow graph can be expressed as G=(V,E), where V is the set of nodes, representing each operation operator; is the set of directed edges, and the edge , indicates that there is a dependency relationship between node and , and is the predecessor node of. In other words, the operator must complete the computing task before the operator executes. In this way, the data flow graph can comprehensively reflect the computing relationship and dependency structure between tasks, providing a clear basis for task partitioning and optimization.
[0090] Figures 2(a) - 2(c) show the task graph and its corresponding partitioning methods, where the arrows between nodes represent the dependencies between tasks. Figure 2(b) shows a reasonable partitioning method that divides the tasks into four stages. In each stage, there are no data dependencies between tasks, so they can be executed in parallel. This partitioning method makes full use of the parallelism of tasks and helps to improve the execution efficiency of the system. In contrast, Figure 2(c) shows an unreasonable partitioning method. This partitioning violates the task dependency principle, resulting in tasks with dependencies being assigned to the same stage for execution. For example, the calculation of task V6 depends on the calculation result of task V3, but in this partitioning method, V6 and V3 are assigned to the same stage, causing the calculations of V6 and V3 to not be able to be executed in parallel at the same time. This partitioning method may cause data asynchronization problems in the actual system, thus having a significant negative impact on the correctness and performance of the calculation results.
[0091] The present invention deeply analyzes the data dependencies of the sea clutter simulation calculation tasks and obtains the data flow graph of the sea clutter simulation-related task calculations as shown, and the specific calculation tasks and parameter descriptions are also given in the figure. The focus is on studying the data dependencies of tasks such as range calculation, clutter cell area calculation, and clutter backscattering coefficient calculation. It can be observed that the calculations of tasks R_h, Temp, and Gound_angle depend on the calculation result of task R, but there are no data dependencies between tasks R_h, Temp, and Gound_angle, so these tasks can be executed in parallel. Similarly, there are no data dependencies between Area, Dp, Sea_sigma, and Antena_gain, so these tasks can be executed in parallel. Figure 3 Based on the above analysis, the topological sorting hierarchical algorithm is particularly suitable for the task partitioning of sea clutter parallel computing. Specifically, the topological sorting hierarchical algorithm constructs a hierarchical architecture design by dividing the calculation tasks into multiple levels and assigning tasks without data dependencies to the same level. This design can make full use of multiple CUDA streams to achieve parallel execution of tasks.
[0092] Figure shows the flowchart of topological sorting and layering, which is completed through the following four steps: Figure 4 Figure shows the flowchart of topological sorting and layering, which is completed through the following four steps:
[0093] (1) Calculate the in-degree: Count the in-degree of each node, that is, how many edges point to this node.
[0094] (2) Initialize the queue: Add all nodes with an in-degree of 0 to the queue, indicating that these nodes have no dependencies and can be used as starting points.
[0095] (3) Topological sorting process: Take out nodes from the queue one by one and add them to the corresponding layers. Traverse all the neighbor nodes pointed to by this node: Decrease the in-degree of the neighbor nodes. If the in-degree of a neighbor node becomes 0, add it to the queue and update its layer at the same time.
[0096] (4) Generate topological layers: Record the layered result after topological sorting, that is, the hierarchical order of each node.
[0097] By analyzing the computational data flow graph of the sea clutter simulation related tasks using the topological sorting and layering algorithm, the layered result as shown in Figure 5 is obtained. The overall computational task is divided into four layers: The first layer is responsible for calculating R, the second layer includes R_h, Temp, and Ground_angle, the third layer contains Area, Dp, Sea_sigma, and Antenna_gain, and the fourth layer processes Reflect_patch. Based on the task division result, the computationally parallelizable tasks in the second and third layers are assigned to different CUDA streams for execution. Since CUDA streams are independent execution queues, tasks in different streams can be executed in parallel, thus making full use of the computing resources of the GPU. For the computational tasks in the first and fourth layers, they are implemented through multi-threaded parallel processing of the GPU.
[0098] Figures 6(a) - 6(c) show the sea clutter simulation time-consuming using the topological sorting and layering algorithm combined with CUDA multi-stream processing and the sea clutter simulation time-consuming using only multi-threading under the same parameter settings. The comparison statistics include the time-consuming for the CPU to transfer data to the GPU, the time-consuming for calculating the clutter distribution modulation weighting, the GPU operation time-consuming, and the time-consuming for the GPU to transfer the operation result to the CPU. The results show that in the sea clutter simulation task, compared with the method of only using multi-threaded parallel processing, the method combining the topological sorting and layering algorithm with CUDA multi-stream processing shows a certain degree of optimization.
[0099] Embodiment 2
[0100] Based on the same inventive concept as in Embodiment 1, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the sea clutter task-level simulation optimization method described in any one of Embodiment 1.
[0101] Embodiment 3
[0102] Based on the same inventive concept as in Embodiment 1, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the sea clutter task-level simulation optimization method described in any one of Embodiment 1.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the present invention's claims. These all fall within the protection scope of the present invention.
[0108] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for optimizing sea clutter task-level simulation, applied to a CPU host, characterized in that Including: Obtain the data flow graph of the pre-built sea clutter calculation task, where the nodes of the data flow graph represent calculation tasks, and the directed edges represent the data dependency relationships between calculation tasks; Perform hierarchical processing on the data flow graph, divide each calculation task into different levels, and there is no data dependency relationship between calculation tasks at the same level; According to the hierarchical order, perform task calculations layer by layer based on the CUDA parallel computing framework.
2. The method for optimizing sea clutter task-level simulation according to claim 1, wherein: When performing hierarchical processing, the topological sorting algorithm is adopted.
3. A method for optimizing sea clutter task-level simulation according to claim 2, characterized in that: The specific process of dividing each calculation task into different levels includes: Count the in-degree of each node in the data flow graph; Add all nodes with an in-degree of 0 to the queue; Take out nodes from the queue in turn, add them to the corresponding layer, and at the same time traverse all neighbor nodes pointed to by the node. If the in-degree of the neighbor node becomes 0, add it to the queue and add the neighbor node to the corresponding layer; Record the hierarchical result after topological sorting, obtain the hierarchical order of each node, and complete the division of each calculation task into different levels.
4. A method for optimizing sea clutter task-level simulation according to claim 1, characterized in that: The step of performing task calculations layer by layer based on the CUDA parallel computing framework according to the hierarchical order includes: If the number of calculation tasks in a certain layer is equal to 1, directly send it to the GPU for processing; If the number of calculation tasks in a certain layer is greater than 1, allocate the same number of CUDA streams according to the number of calculation tasks, and allocate the parallelizable calculation tasks to different CUDA streams for execution.
5. A method for optimizing sea clutter task-level simulation according to claim 4, characterized in that: The GPU processes the calculation tasks using the following steps: Start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel calculation of the corresponding calculation task according to the index; Save the calculation result and transfer it to the CPU host through memory copy.
6. A method for optimizing sea clutter task-level simulation according to claim 4, characterized in that: The step of allocating the same number of CUDA streams according to the number of calculation tasks and allocating the parallelizable calculation tasks to different CUDA streams for execution includes: Allocate the corresponding number of CUDA streams in the CPU host according to the number of calculation tasks, initialize the GPU, allocate video memory space, and set the corresponding thread block and thread number; Control the GPU to start the corresponding kernel function, calculate the thread index and thread block index, and complete the parallel calculation of the corresponding calculation task according to the index; Save the calculation result and transfer it to the CPU host through memory copy.
7. A method for optimizing sea clutter task-level simulation according to claim 1, characterized in that: The step of performing hierarchical processing on the data flow graph and dividing each calculation task into different levels includes: Divide the data flow graph into four layers; The first layer is responsible for calculating the slant range between the beam center and the radar; The second layer is responsible for calculating the horizontal distance between the beam center and the radar, intermediate variables, and grazing angle; The third layer is responsible for calculating the clutter cell block area, clutter Doppler frequency, clutter cell block backscattering coefficient, and radar antenna gain; The fourth layer is responsible for calculating the clutter modulation coefficient matrix.
8. A method for optimizing sea clutter task-level simulation according to claim 7, characterized in that: The calculation tasks of the first layer and the fourth layer are processed in parallel through multi-threading of the GPU; the parallelizable calculation tasks in the second layer and the third layer are allocated to different CUDA streams for execution.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When this program is executed by a processor, it implements the sea clutter task-level simulation optimization method described in any one of claims 1 to 8.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method for optimizing sea clutter task-level simulation according to any one of claims 1 to 8 is implemented.