A storm surge ensemble numerical prediction method and device for parallel computing
By combining parallel computing and FPGA coprocessors, dynamically migrating computing nodes and quickly recovering abnormal data, the problems of low efficiency and insufficient robustness in traditional storm surge numerical forecasting methods are solved, and efficient and accurate storm surge forecasting is achieved.
Patent Information
- Application Number
- CN202511211413.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Traditional storm surge numerical forecasting methods are inefficient and inaccurate when dealing with large-scale, high-resolution computing tasks, and lack robustness, making them difficult to cope with hardware failures and data anomalies.
A parallel computing-based ensemble numerical forecasting method for storm surges is proposed. By constructing a three-dimensional hard collaboration-physical constraint-millisecond fault-tolerant architecture, a random perturbation field is generated using an FPGA coprocessor. Combined with multi-layer Monte Carlo resource dynamic configuration and dual pipeline scheduling, dynamic migration of computing nodes and rapid recovery of abnormal data blocks are achieved.
It improves the efficiency and accuracy of storm surge numerical forecasting, enhances the robustness of the computing system, enables computation to be restored in a very short time, and ensures the continuity and reliability of forecasts.
Smart Images

Figure CN120743476B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of storm forecasting, and more particularly to a parallel computing method and apparatus for ensemble numerical forecasting of storm surges. Background Technology
[0002] Storm surge is an abnormal rise and fall of sea level caused by natural factors such as strong winds, sudden changes in air pressure, or spring tides. It is often accompanied by dramatic tidal changes and poses a serious threat to coastal infrastructure, the ecological environment, and the safety of residents' lives and property. With the intensification of global climate change and the increasing frequency of extreme weather events, the intensity and frequency of storm surge disasters are showing an upward trend, placing higher demands on the accuracy of storm surge forecasting and rapid response capabilities.
[0003] Numerical storm surge forecasting, as a crucial tool for storm surge early warning, uses mathematical models to simulate the dynamic processes of storm surges, enabling early prediction of their timing, intensity, and impact range. However, traditional numerical storm surge forecasting methods face significant challenges when handling large-scale, high-resolution computational tasks. On one hand, as the precision of computational grids increases, the computational load grows exponentially, making it difficult for traditional serial computation methods to meet the demands of real-time forecasting. On the other hand, the complexity and uncertainty of storm surge dynamic processes require forecasting systems to possess high robustness and fault tolerance to cope with hardware failures and data anomalies during the computation process.
[0004] Therefore, we propose a parallel computational method and apparatus for ensemble numerical forecasting of storm surges to address the aforementioned problems. Summary of the Invention
[0005] This invention provides a parallel computing method and apparatus for ensemble numerical forecasting of storm surge, which improves the efficiency, accuracy and robustness of ensemble numerical forecasting of storm surge.
[0006] The first aspect of this invention provides a parallel computing method for ensemble numerical forecasting of storm surges. This method includes: acquiring a digital elevation model dataset; generating a terrain curvature matrix using a curvature calculation module; executing a grid partitioning algorithm to output a grid topology file, wherein the grid density for land areas is 30%-50% of that for sea areas; dividing the total duration into 2K equal-length segments based on forecast duration parameters and the grid topology file; constructing a prediction pipeline queue and a correction pipeline queue; and outputting a dual-pipeline scheduling instruction set; outputting a priority signal based on a hierarchical variance statistics table of historical storm surge ensembles; generating node migration control instructions; and migrating computing nodes from low-variance levels to high-variance levels in real time to obtain a dynamic resource allocation table; and based on real-time wind field data streams and grid data... In the FPGA coprocessor, the topology file is used to: call the KL decomposition hardware core to calculate the eigenbase of the covariance matrix; generate random coefficients through polynomial chaotic expansion of the hardware core; and perform matrix multiplication and addition operations to output a random perturbation field dataset. Based on the mesh topology file, dual-pipeline scheduling instruction set, resource dynamic configuration table, and random perturbation field dataset, the MPI process group is started according to the mesh topology file. Dual-pipeline tasks are executed alternately according to the dual-pipeline scheduling instruction set, and computing nodes are migrated in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set. Based on the original storm surge water level field time series set and the pre-stored computing node state entropy fingerprint database, abnormal data blocks are detected by the entropy value comparator, triggering the process hot migration module to resume calculation within ≤100ms to obtain the storm surge set forecast results.
[0007] Optionally, in a first implementation of the first aspect of the present invention, the grid topology file includes a grid weight vector, an adjacency table, and a boundary mapping function set; in the grid weight vector, the weight coefficient of the land region subdomain ranges from 0.3 to 0.5, and the weight coefficient of the sea region subdomain is fixed at 1.0, and computing resources are allocated when the MPI process group is started; the adjacency table includes a set of Cartesian coordinates recording the shared boundaries between non-uniform subdomains and metadata of the MPI communication interface, and the adjacency table defines the data exchange path between MPI processes; the boundary mapping function set includes a bilinear interpolation rule base at the land-sea junction and a conserved flux transfer coefficient matrix, and the boundary mapping function set ensures the continuous transfer of physical quantities at the subdomain boundaries.
[0008] Optionally, in a second implementation of the first aspect of the present invention, the dual-pipeline scheduling instruction set includes: a time slice-hardware mapping table, using a key-value pair set, where odd-order time slice IDs are bound to CPU clusters and even-order time slice IDs are bound to FPGA coprocessors; a pipeline synchronization signal sequence, using a timing instruction queue, where cross-process synchronization is triggered upon completion of each time slice, and hardware-level ready signals are inserted between odd and even queues to eliminate pipeline wait dependencies; and a fault-tolerant time window parameter set, using threshold configuration. .
[0009] Optionally, in the third implementation of the first aspect of the present invention, the resource dynamic configuration table includes: a hierarchy-node mapping matrix, which is a two-dimensional array containing a list of physical addresses of computing nodes corresponding to the coarse, medium, and fine levels, and records the node migration status in real time, including online, migrating, and faulty; a node migration control instruction queue, which is a time-sequential instruction set that generates migration instructions in response to variance priority signals and directly controls the migration of physical computing nodes; and a resource utilization monitoring vector, which is a real-time sampling array that collects the hardware resource occupancy rate of each level every second, and updates the historical variance statistics table with the data feedback to provide a benchmark for resource anomaly detection.
[0010] Optionally, in the fourth implementation of the first aspect of the present invention, the random disturbance field dataset includes: a disturbance vector bound to grid coordinates, the structure of which adopts spatially discretized data blocks, is strictly aligned with the subdomain coordinates in the grid topology file, and the physical constraints include wind speed disturbance and wind direction disturbance; a hardware-calculated checksum, the structure of which is a 32-bit cyclic redundancy checksum, which is generated in real time based on the disturbance vector data stream to provide a data integrity benchmark; and a data block partitioning index table, the structure of which is a key-value mapping, inherits the subdomain partitioning scheme of the grid topology file, defines the physical address directly accessed by the MPI process, and eliminates the data copying overhead of the MPI process.
[0011] Optionally, in the fifth implementation of the first aspect of the present invention, the obtained original storm surge water level field time series set includes: a subdomain water level field data block, which adopts a three-dimensional floating-point array, a sequence of local water level values calculated by each MPI process, with spatial resolution strictly aligned with the grid topology file; a cross-domain synchronization marker sequence, which adopts timestamp-based verification records to record subdomain boundary data exchange events, including 32-bit CRC boundary check values; and a calculation state snapshot set, which adopts a process memory image sequence, which saves the complete process state at fixed intervals, including register values and stack snapshots.
[0012] Optionally, in the sixth implementation of the first aspect of the present invention, the storm surge ensemble forecast results include: a storm surge probability distribution map, which adopts a two-dimensional floating-point matrix and includes spatially gridded water level exceedance probability values and key risk area identifiers, used to input the disaster prevention and early warning system to release inundation risk maps; an abnormal event record table, which is a spatiotemporal coordinate event set structure, records all abnormal events detected by the entropy value comparator, including hardware fault classification, used to drive fault diagnosis of the supercomputing operation and maintenance platform; and a hot migration audit log, which adopts a verification report structure, records the time consumed for each hot migration, including data consistency verification before and after migration, used to verify the real-time performance of the fault-tolerant system.
[0013] The second aspect of this invention provides a parallel computing storm surge ensemble numerical forecasting device, comprising: an acquisition module for acquiring a digital elevation model dataset, generating a terrain curvature matrix through a curvature calculation module, executing a grid partitioning algorithm to output a grid topology file, wherein the grid density for land areas is 30%-50% of that for sea areas; a processing module for dividing the total duration into 2K equal-length segments based on forecast duration parameters and the grid topology file, constructing a prediction pipeline queue and a correction pipeline queue, and outputting a dual pipeline scheduling instruction set; a migration module for outputting priority signals based on a hierarchical variance statistics table of historical storm surge ensembles, generating node migration control instructions, migrating computing nodes from low variance levels to high variance levels in real time, and obtaining a dynamic resource allocation table; and a setting module for setting based on real-time wind... The field data stream and mesh topology file are processed in the FPGA coprocessor as follows: The KL decomposition hardware core is called to calculate the eigenbase of the covariance matrix. Random coefficients are generated through polynomial chaotic expansion of the hardware core. Matrix multiplication and addition operations are performed to output a random perturbation field dataset. The sequence module, based on the mesh topology file, dual-pipeline scheduling instruction set, resource dynamic configuration table, and random perturbation field dataset, initiates an MPI process group according to the mesh topology file. Dual-pipeline tasks are executed alternately according to the dual-pipeline scheduling instruction set. Computational nodes are migrated in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set. The allocation module, based on the original storm surge water level field time series set and a pre-stored computational node state entropy fingerprint database, detects abnormal data blocks through an entropy value comparator and triggers the process hot migration module to resume computation within ≤100ms to obtain the storm surge set forecast results.
[0014] The mechanism of this invention is as follows:
[0015] A three-dimensional hard collaboration, physical constraint, and millisecond fault tolerance-integrated storm surge ensemble forecasting architecture was constructed, breaking through the bottlenecks of efficiency, accuracy, and reliability in traditional numerical forecasting.
[0016] By embedding physical constraint circuits (wind speed / direction limiting, conservation interpolation) into FPGAs, the stability problem of numerical models can be solved.
[0017] Based on the Lorenz attractor, an entropy-tolerant chain is used to achieve a reliability leap from "passive restart" to "active immunity". Beneficial effects
[0018] By using a hierarchical variance statistics table and a dynamic resource allocation table, real-time migration of computing nodes from low-variance hierarchies to high-variance hierarchies is achieved, breaking the limitations of traditional static resource allocation.
[0019] Construct a predictive pipeline queue and a corrective pipeline queue, and implement alternating execution through a dual pipeline scheduling instruction set, combined with hardware-level ready signals to eliminate pipeline waiting dependencies.
[0020] An entropy comparator is introduced to detect abnormal data blocks, and the computation is restored in a very short time through a process hot migration module.
[0021] The KL decomposition hardware core and the polynomial chaotic expansion hardware core are called in the FPGA coprocessor to accelerate the generation of random perturbation fields. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an embodiment of the parallel computational storm surge ensemble numerical forecasting method in this invention.
[0023] Figure 2 This is a schematic diagram of another embodiment of the storm surge ensemble numerical forecasting method using parallel computing in this invention.
[0024] Figure 3 This is a schematic diagram of an embodiment of the storm surge ensemble numerical forecasting device with parallel computing in this invention.
[0025] Figure 4 This is a schematic diagram of an embodiment of a storm surge ensemble numerical forecasting device that performs parallel computing in an embodiment of the present invention. Detailed Implementation
[0026] This invention provides a parallel computing method and apparatus for ensemble numerical forecasting of storm surges, which improves the efficiency, accuracy, and robustness of ensemble numerical forecasting. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the parallel computational storm surge ensemble numerical forecasting method in this invention includes:
[0028] 101. Obtain the digital elevation model dataset, generate the terrain curvature matrix through the curvature calculation module, execute the grid generation algorithm to output a non-uniform grid data structure file, wherein the grid density of the land area is 30%-50% of that of the sea area, and output a grid topology file with weight coefficients.
[0029] It is understood that the executing entity of this invention can be a parallel computing storm surge ensemble numerical forecasting device, or it can be a terminal or a server; the specific implementation is not limited here. This embodiment of the invention will be described using a server as an example.
[0030] It should be noted that the data source for DEM data acquisition and preprocessing was the Copernicus GLO-30 DEM dataset with a resolution of 30 meters, covering the coastal area of Fujian (23°-28°N, 117°-122°E). This data was acquired via radar satellite and processed to ensure hydrological consistency (flattening water bodies and correcting river flow direction).
[0031] Data preprocessing: The original DEM was resampled to 100-meter resolution using a bilinear interpolation algorithm to eliminate local noise and preserve topographic trends.
[0032] Terrain curvature matrix generation and curvature calculation module: Based on ArcGIS's raster surface analysis tools, calculates the total curvature of the terrain.
[0033] ,
[0034] in This represents the elevation value. The curvature threshold is set to 0.01: curvature > 0.01 is marked as a steep slope / complex terrain (mountainous area), and ≤ 0.01 is marked as a gentle area (coastal plain).
[0035] Output matrix: Generates a curvature grid with a resolution of 100 meters and a matrix dimension of 1200×900 (covering 120,000 square kilometers).
[0036] Non-uniform mesh generation algorithm, mesh density strategy:
[0037] Land area: Based on the curvature matrix, the grid is refined to 300 meters in steep slope areas (curvature > 0.01), and 500 meters in gentle areas, with an average density of 40% of that of the sea area (meeting the 30%-50% requirement).
[0038] Sea area: Nearshore (water depth < 50 meters) grid side length 800 meters, offshore (water depth > 50 meters) set at 1000-1500 meters.
[0039] Mesh generation tool: The SMS software is used to construct an irregular triangular mesh (TIN), which is divided according to the principle of "coarse mesh in the open sea and fine mesh near the shore". When fitting the coastline, high-resolution coastline data of GSHHS is used to ensure boundary accuracy. In the end, about 150,000 triangular cells are generated, with land area accounting for 30% and mesh transition gradient ≤0.2.
[0040] Weight allocation and topology file output, weight coefficient definition:
[0041] Terrain curvature weight Steep slope unit weight = 1.5, gentle slope area = 1.0;
[0042] Land and sea type weights Land unit = 0.7, sea area = 0.3 (reflecting the need for high resolution in land areas).
[0043] Topology file structure: Outputs a NetCDF format file, containing:
[0044] Node coordinates (longitude, latitude, elevation);
[0045] Unit connection relationship (triangle vertex index);
[0046] Weight coefficient matrix (each cell associated) × value).
[0047] Summary of key parameters:
[0048] elements parameter DEM resolution 30 meters (original) → 100 meters (resampling) Land grid density 40% of the sea area Minimum grid side length 300 meters on land, 800 meters near the sea Curvature threshold 0.01 (Distinguishing between steep slopes and gentle slopes) Total number of grid cells ≈150,000 triangles
[0049] By adaptively refining the curvature-driven mesh, a balance is struck between computational efficiency and terrain detail requirements, providing a high-precision spatial discretization foundation for subsequent storm surge simulations.
[0050] 102. Based on the forecast duration parameters and grid topology file, the total duration is divided into 2K equal-length segments (K≥12). A prediction pipeline queue (allocating odd-order segments to coarse-grained computing units) and a correction pipeline queue (allocating odd-order segments to fine-grained computing units) are constructed. A dual-pipeline scheduling instruction set with time slice identifiers is output.
[0051] It should be noted that the forecast duration parameter is set to a total duration of 72 hours (simulating the entire process of Typhoon Haikui from its formation to its landfall in Fujian).
[0052] Segment division: Taking K=12, the total duration is divided into 2K=24 segments of equal length, each segment having a duration of... =3 hours.
[0053] Mesh topology file parsing, mesh structure: based on the non-uniform mesh (approximately 150,000 triangular elements) generated in step 101 along the Fujian coast, where:
[0054] Land area: Average grid side length 300 meters (intensified areas in coastal plains such as Fuzhou and Putian);
[0055] Marine areas: Nearshore grid side length 800 meters (water depth < 50 meters), offshore grid side length 1000-1500 meters (water depth > 50 meters).
[0056] Weighting coefficients: land area unit weighting coefficient w=1.5 (complex terrain), sea area unit w=0.3 (lower computational requirements).
[0057] Dual pipeline queue construction, predicting pipeline queues (odd-order segments):
[0058] Computational units: Assigned to coarse-grained computing units (single-node CPU clusters), with each node processing ≥5,000 grids;
[0059] Task type: Perform large-scale storm surge trend prediction driven by low-resolution wind field (1000-meter resolution), ignoring local terrain details.
[0060] Correcting pipeline queues (even-order segments):
[0061] Computational units: Assigned to fine-grained computing units (GPU / FPGA accelerated nodes), with each node processing ≤500 grids;
[0062] Task type: Based on odd-numbered fragment results, inject high-resolution wind field disturbances (300-meter resolution) and couple astronomical tides and nearshore waves to correct water level extremes.
[0063] Time slice identifiers and scheduling logic; identifier design: Each time slice is bound to a triple (time sequence ID, grid partition ID, compute unit type):
[0064] Segment 1 (0-3 hours): (T1, Grid_A, CPU_Coarse);
[0065] Segment 2 (3-6 hours): (T2, Grid_B, GPU_Fine);
[0066] Scheduling rule: Alternate execution: Odd-order segments (T1, T3, ..., T23) are calculated first by the prediction queue, while even-order segments (T2, T4, ..., T24) rely on the output of the previous segment to initiate correction;
[0067] Dynamic dependency: If T1 is delayed in completion, T2 will be automatically suspended until T1 is ready to output, ensuring time continuity.
[0068] Output instruction set structure, field definitions:
[0069] field name Data types Example value TimeSlice_ID Integer 1 (odd number), 2 (even number) Grid_Subset string array ["Fujian_coast_zone1"] Compute_Unit enumerate Coarse / Fine Dependency Integer 0 (no dependency), 1 (depends on T1)
[0070] 103. Based on the hierarchical variance statistics table of the historical storm surge set, the priority signal is output through the variance comparator to generate node migration control instructions, and the computing nodes are migrated from the low variance level to the high variance level in real time to obtain the multi-level Monte Carlo (MLMC) resource dynamic configuration table.
[0071] It should be noted that the historical storm surge ensemble data was prepared from the following sources: historical storm surge data of the Fujian coast over the past 10 years (2015-2025), including 20 typhoon events (such as "Doksuri" and "Saola"), with 50 ensemble members generated for each event (a total of 1000 sets of data).
[0072] Hierarchical division: Based on grid resolution, it is divided into 5 levels (L1-L5), where L1 is coarse resolution (1km, full coverage) and L5 is ultra-high resolution (20m, focusing on easily flooded areas such as Fuzhou and Putian).
[0073] Construction of hierarchical variance statistics table, variance calculation: Calculate the variance of water level prediction for each level:
[0074] L1 (1km): Mean variance 0.02 m 2 (The overall trend is stable with minimal fluctuations);
[0075] L5 (20m): Mean variance 0.15 m 2 (Significantly affected by topography and tidal coupling, resulting in violent fluctuations).
[0076] Priority sorting: Output hierarchical priority signals through a variance comparator.
[0077] Hierarchical ID variance mean Priority signals L5 0.15 m2 Maximum (signal value = 5) L3 0.08 m2 Medium (signal value = 3) L1 0.02 m2 Minimum (signal value = 1)
[0078] Triggering condition: When the variance ratio between levels is >1.5 (L5 / L1=7.5), it is marked as "high variance requires resource allocation".
[0079] Dynamic migration control of compute nodes, initial resource allocation: a total of 200 compute nodes, allocated according to a uniform strategy (40 nodes per level).
[0080] Migration rules:
[0081] Migration Trigger: Real-time monitoring of the variance statistics table; if the variance at level L5 suddenly increases to 0.18 m... 2 (6 hours before the typhoon makes landfall), the variance comparator generates migration instructions;
[0082] Migration action: Extract 50% of the nodes (40 in total) from the low variance levels (L1, L2) and dynamically allocate them to the L5 level, increasing its number of nodes to 80;
[0083] Migration delay: Response time ≤ 10ms, ensuring resource adjustments are synchronized with typhoon evolution.
[0084] The MLMC resource dynamic configuration table output includes the following structure: fields such as hierarchy weight, node quota, and variance threshold.
[0085] Hierarchical ID Node quota Variance threshold Weighting coefficient L5 80 >0.12 m2 1.5 L3 40 0.05~0.12 m2 1.0 L1 20 <0.05 m2 0.5
[0086] Dynamic assurance: The configuration table is updated every 3 minutes, and the threshold is adjusted based on real-time wind field data (wind speed in the eyewall area of Typhoon Haikui is 45m / s).
[0087] 104. Based on the real-time wind field data stream and grid topology file, in the FPGA coprocessor: call the KL decomposition hardware core to calculate the eigenbase of the covariance matrix, generate random coefficients through the polynomial chaotic expansion (PCE) hardware core, perform matrix multiplication and addition operations to output physical constraint perturbation field data blocks, and obtain a spatially discretized random perturbation field dataset.
[0088] It should be noted that the input data preparation and real-time wind field data stream are as follows:
[0089] Source: Real-time data transmission from Fujian coastal buoy array and meteorological radar, covering the core area of the strait (24°-26.5°N, 118°-121°E), with a temporal resolution of 5 minutes and a spatial resolution of 1km×1km.
[0090] Parameters: Wind speed range 12-45 m / s (peak in the eyewall region of the typhoon), wind direction NW-SE alternation, data volume 2.4 GB / s.
[0091] Mesh topology file: Based on the Fujian non-uniform mesh (approximately 150,000 triangular cells) generated in step 101, with a weight coefficient of 1.5 for land cells (complex terrain area) and 0.3 for marine cells (smooth offshore area).
[0092] FPGA coprocessor hardware core calling, KL decomposition of hardware cores:
[0093] Input: Wind field covariance matrix (dimension 1200×1200, generated by real-time wind field data stream within a 10ms window).
[0094] Calculation process: The covariance matrix is decomposed using the Jacobi iterative algorithm (hardware optimized version), and the top 20 principal eigenvectors (cumulative contribution rate > 95%) are extracted to form the feature basis matrix. (Dimensions 1200×20).
[0095] Hardware resource usage: 16,000 DSP slices (65% of FPGA resources), latency 3.2 ms.
[0096] PCE hardware core:
[0097] Input: Feature basis matrix + Gaussian random number seed (generated by the FPGA's built-in LFSR pseudo-random generator).
[0098] Calculation process: A 3rd-order polynomial chaotic expansion (Hermite orthogonal basis) is used to generate a 20-dimensional random coefficient vector. .
[0099] The expansion operation is performed using a parallel multiply-accumulate engine (48 MAC units):
[0100] ,
[0101] Output the physical constraint perturbation field (satisfying mass conservation and curl constraints) with a delay of 1.8 ms.
[0102] Matrix multiplication and addition operations and perturbation field generation: Hardware acceleration strategies.
[0103] Data flow partitioning: The grid topology is divided into 16 spatial blocks (key areas such as Minjiang Estuary and Pingtan Island), and loaded into the FPGA cache in parallel through DMA direct memory access.
[0104] Computational optimization: Pipeline technology is used to break down multiply-accumulate operations into a three-stage pipeline (data loading → vector dot product → result accumulation), achieving a throughput of 480 GFLOP / s.
[0105] Double buffering design: When calculating the current block, the data of the next block is preloaded, hiding 60% of the data transmission latency.
[0106] Output dataset: Spatial discretized perturbation field: 150,000 elements × 20 ensemble members, data format FP32, size 120 MB.
[0107] Key indicators: Disturbance amplitude in Fuzhou Binhai New City area ±0.35 m (reflecting the asymmetric structure of the typhoon), and in Meizhou Island waters of Putian ±0.18 m (modulated by the island's topography).
[0108] Key performance indicators:
[0109] Components computation delay resource utilization Precision control KL Decomposition Core 3.2 ms 65% DSP Eigenvalue error <1% PCE random expansion kernel 1.8 ms 30% LUT Statistical moment matching >90% Matrix multiply-accumulate engine 2.5 ms 80% BRAM Floating-point rounding error ≤ 1e-6
[0110] 105. Based on the grid topology file, dual pipeline scheduling instruction set, MLMC resource dynamic configuration table, and random disturbance field dataset, start the MPI process group according to the grid topology file, execute dual pipeline tasks alternately according to the scheduling instruction set, and migrate computing nodes in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set.
[0111] It should be noted that the MPI process group initialization and mesh allocation, mesh topology resolution: load the Fujian coastal non-uniform mesh file (150,000 triangular elements, the land area mesh density is 40% of the sea area) generated in step 101.
[0112] MPI process startup: The mpirun command is called to start 256 MPI processes (corresponding to 64 compute nodes, 4 processes per node), and tasks are allocated according to grid weight: high weight land area (Fuzhou coastal area, weight=1.5): 128 processes (fine-grained cells) are allocated; low weight sea area (weight=0.3): 128 processes (coarse-grained cells) are allocated.
[0113] Dual-pipeline task alternating execution, scheduling instruction set parsing:
[0114] Predictive pipeline (odd-order segments): 12 segments, such as 0-3 hours (T1) and 6-9 hours (T3), are assigned to coarse-grained cells (5000 grids per node), and low-resolution wind field driven calculations are performed (1km grid), with a time ≤90 seconds per segment;
[0115] Correction pipeline (even-order segments): 12 segments, such as 3-6 hours (T2) and 9-12 hours (T4), are assigned to fine-grained units (500 grids per GPU). The random perturbation field generated in step 104 (perturbation amplitude ±0.35m in Fuzhou area) is injected and coupled with astronomical tide correction. The resolution is 300 meters and the time taken is ≤120 seconds per segment.
[0116] Alternating logic: T2 is triggered after T1 is completed. T2 depends on the output water level field of T1 as the initial condition and synchronizes the boundary data through MPI_Send / Recv.
[0117] Dynamic node migration and fault tolerance, dynamic resource configuration:
[0118] Initial state: The ratio of coarse to fine granular nodes is 1:1 (128 processes each);
[0119] Migration Trigger: When the eyewall of the typhoon approached the Minjiang River estuary (T8 time period), the MLMC table detected a sudden increase in variance to 0.18m in that area. 2 (threshold > 0.12m) 2 The variance comparator generates the migration instructions;
[0120] Migration action: From a low-variance sea area (variance = 0.05m) 2 ) 32 nodes were transferred to the Minjiang Estuary, increasing the number of fine-grained nodes to 160, with a migration response time of ≤10ms.
[0121] Process hot migration: If a node failure causes an abnormal entropy value (data block entropy > 5.0), the process will be migrated to a standby node within 100ms to ensure computational continuity.
[0122] Water level field time series output, data integration: each segment outputs a water level increase field matrix (150,000 elements × 20 ensemble members); key area time series recording: Minjiang Estuary outputs water level extreme values every 10 minutes (peak value of 4.2m ± 0.3m in time period T12).
[0123] Format and performance: Outputs a NetCDF format time series set containing 72 hours and 1440 time steps; total computation time is reduced from 24 hours to 6.5 hours, improving efficiency by 73%.
[0124] Technical effects:
[0125] index Before migration After migration Minjiang Estuary Calculation Speed 120 seconds / segment 85 seconds / segment Extreme error of water level ±0.45m ±0.25m resource utilization rate 78% 92%
[0126] 106. Based on the water level field time series set and the pre-stored computation node state entropy fingerprint database, abnormal data blocks are detected by the entropy value comparator, triggering the process hot migration module to resume calculation within ≤100ms, integrating effective data to generate probabilistic forecast result files, and obtaining the storm surge set forecast result after fault tolerance processing.
[0127] It should be noted that the fingerprint database for computing node state entropy is constructed by extracting state parameters every 10 seconds over 72 hours based on the historical running data (CPU load, memory usage, network latency) of the MPI process group from step 105.
[0128] Entropy calculation: Shannon entropy formula is used to quantify node stability.
[0129] ,
[0130] in This represents the probability distribution of each state parameter.
[0131] Fingerprint database parameters: The entropy values of 256 pre-stored computing nodes are within the baseline range of 0.8-3.2 (>4.0 is considered abnormal).
[0132] Anomaly detection and thermal migration, entropy value comparator: Real-time monitoring of water level field time series set (150,000 grids × 20 set members), scanning node entropy values once every 5 seconds.
[0133] Anomaly Trigger: When the entropy value of node N78 (responsible for Fuzhou Binhai area) suddenly increases to 4.8 (due to hardware failure causing data block entropy value deviation >15%), the comparator marks the abnormal data block (grid ID: FJ_Coast_Zone5).
[0134] Process hot migration: Migration target: Select standby node N203 in the same data center (entropy value 1.2, resource idle rate 85%). Migration process: Create a new process dynamically through MPI_Comm_spawn, copy the process state (memory image, computing context), and the migration time is ≤75ms (meeting the ≤100ms requirement).
[0135] Probabilistic forecast results were generated, and effective data were integrated: anomalous data (T12-T13 segments) during the N78 node failure period were removed, and the effective results of the remaining 19 set members were retained. Based on the random perturbation field dataset, the water level sequence of the Minjiang Estuary in Fuzhou (peak value 4.2m ± 0.3m) was reconstructed.
[0136] Output: Flood risk probability map: The probability of flooding in the low-lying area of the Minjiang Estuary (Mawei Port) is >85% (water depth >1.5m); Extreme water level distribution: 90% confidence interval is 3.8-4.6m, median is 4.1m; Time-sensitive warning: Red alert is triggered 3 hours before landfall (water increase rate >0.5m / h).
[0137] In this embodiment of the invention, digital elevation model (DEM) data is acquired and preprocessed to generate a non-uniform grid and topology file; a dual-pipeline scheduling instruction set is constructed based on the forecast duration and grid topology file; a multi-layer Monte Carlo (MLMC) resource dynamic configuration table is generated based on historical storm surge ensemble data; a random disturbance field dataset is generated using an FPGA coprocessor; combining the above files and datasets, the dual-pipeline tasks are executed alternately by an MPI process group and the resource dynamic configuration is responded to, resulting in the original storm surge water level field time series set; finally, a fault-tolerant probabilistic forecast result is generated through entropy comparison and process hot migration. Based on an adaptive densification grid of terrain curvature, the grid is denser on steep slopes and sparser in gentle areas of land, and finer nearshore and coarser offshore, while also considering the weighting of land and sea types. The forecast duration is divided into equal-length segments, and a dual pipeline queue for prediction and correction is constructed to execute tasks alternately, with dynamic dependency rules designed. Based on the hierarchical variance statistics of historical storm surge ensemble data, computational nodes are dynamically migrated from low-variance to high-variance levels. KL decomposition and PCE hardware cores are invoked in the FPGA coprocessor, and matrix multiplication and addition operations are combined to generate a physical constraint perturbation field. A computational node state entropy fingerprint database is constructed, and an entropy value comparator detects abnormal data blocks, triggering a process hot migration module to resume computation. During the computation process, node failures can be detected and handled promptly, ensuring computational continuity and data integrity, improving the stability and reliability of the forecasting system, and avoiding forecast interruptions or errors caused by node failures.
[0138] Please see Figure 2 Another embodiment of the parallel computational storm surge ensemble numerical forecasting method in this invention includes:
[0139] 201. Obtain the digital elevation model dataset, generate the terrain curvature matrix through the curvature calculation module, execute the grid generation algorithm to output a non-uniform grid data structure file, wherein the grid density of the land area is 30%-50% of that of the sea area, and output a grid topology file with weight coefficients.
[0140] Specifically, the output mesh topology file with weighted coefficients includes:
[0141] Grid weight vector: The weight coefficient for land subdomains ranges from 0.3 to 0.5, while the weight coefficient for sea subdomains is fixed at 1.0; it is used to allocate computing resources when starting the MPI process group, and the input source is the spatial attribute data output by the grid partitioning algorithm;
[0142] The adjacency table records the set of Cartesian coordinates of shared boundaries between non-uniform subdomains and includes metadata (process ID, memory address) of the MPI communication interface; it is used to define data exchange paths between MPI processes; the input source is the topology analysis results of the mesh generation algorithm.
[0143] Boundary mapping function set: a bilinear interpolation rule base at the land-sea boundary, containing a conserved flux transfer coefficient matrix to ensure the continuous transfer of physical quantities at the subdomain boundary. The input source is the land-sea boundary marker in the digital elevation model data.
[0144] It should be noted that the DEM data acquisition and preprocessing were based on 12.5-meter resolution DEM data (ALOS satellite data) of the coastal area of China, covering longitude 118°E-122°E and latitude 24°N-28°N.
[0145] Land-sea boundary demarcation: Land-sea boundaries are demarcated based on the Normalized Difference Ratio (NDWI). Specific process:
[0146] Calculate the reflectivity difference between the blue and green bands: A threshold > 0.2 indicates a water area.
[0147] Combined with elevation correction: water area elevation range [-5m, 0.5m] (near sea level), land area elevation > 0.5m.
[0148] Output: A binary mask (0: land, 1: ocean) as input to the set of boundary mapping functions.
[0149] Terrain curvature calculation and mesh generation, curvature matrix generation:
[0150] Using the Gaussian curvature formula:
[0151] ,
[0152] Calculate the curvature of each DEM grid point. High curvature regions ( Areas with a topographic change of >0.05 (coastal cliffs) are marked as abrupt changes in terrain.
[0153] Non-uniform mesh generation:
[0154] Sea area: Basic grid resolution 500m (weight coefficient 1.0).
[0155] Land: The grid density is 40% of that of the sea area (i.e., the resolution is 1.25km), and the weighting factor is 0.4.
[0156] Partitioning algorithm: Quadtree adaptive subdivision ensures smooth grid transition at the land-sea boundary.
[0157] Mesh topology file output, mesh weight vector:
[0158] Subdomain type Weighting coefficient Computing resource allocation Sea area 1.0 Allocate 40% of the MPI process land 0.4 Allocate 60% of the MPI process
[0159] Allocate computing resources based on weighting coefficients: ;
[0160] Adjacency table:
[0161] Subdomain ID Adjacent Subdomain ID Shared boundary coordinate set Communication metadata 101 102 {(118.5,25.1), ...} [Process ID: 3, Memory Address: 0x7FFA] 102 103 {(118.7,25.3), ...} [Process ID: 5, Memory Address: 0x8B00]
[0162] The shared boundary is the land-sea boundary line, and the communication metadata is automatically generated by MPI_Cart_create.
[0163] Boundary mapping function set: Bilinear interpolation rule: Elevation value of grid point P at the land-sea boundary;
[0164] The coefficients (a,b,c,d) are fitted by four adjacent points.
[0165] Flux transfer coefficient: The conservation matrix ensures the continuity of water flow (the flux attenuation factor from ocean to land is 0.3).
[0166] 202. Based on the forecast duration parameters and grid topology file, the total duration is divided into 2K equal-length segments (K≥12). A prediction pipeline queue (allocating odd-order segments to coarse-grained computing units) and a correction pipeline queue (allocating odd-order segments to fine-grained computing units) are constructed, and a dual-pipeline scheduling instruction set with time slice identifiers is output.
[0167] Specifically, the set of dual-pipeline scheduling instructions that output time-slice identifiers includes:
[0168] Time-Hardware Mapping Table: Structure: Key-value pair set [Time Slice ID, Computation Unit Type]; Odd-order time slice ID → Binds to CPU cluster (performs shallow water equation calculations with resolution ≤ 5km); Even-order time slice ID → Binds to FPGA coprocessor (performs turbulence model calculations with resolution ≥ 0.5km); Used to drive the "Alternating Execution of Dual Pipeline Tasks" step; Input source is the subdomain computing capability identifier in the mesh topology file;
[0169] Pipeline synchronization signal sequence: Structure: Timing instruction queue {trigger time, MPI_Barrier instruction, target process group ID}; Content: Cross-process synchronization is triggered when each time slice is completed; Hardware-level ready signals are inserted between odd and even queues; Used to eliminate pipeline wait dependencies; Input sources are time slice partitioning results and hardware delay parameters;
[0170] Fault-tolerant time window parameter set: Structure: Threshold configuration {time slice ID, maximum allowable delay} ;content: (CPU task ≤ 300s, FPGA task ≤ 50ms); Input source is hardware performance benchmark test report;
[0171] It should be noted that this is a 72-hour typhoon storm surge forecast for the East China Sea region (total duration 72 hours, K=12, 24 time slots).
[0172] Time slice partitioning and queue construction: Time slice partitioning: Divide the total duration of 72 hours into 24 segments of equal length of 3 hours (ID=1~24).
[0173] Dual pipeline distribution:
[0174] Predict the pipeline queue (odd-numbered IDs: 1, 3, ..., 23) → allocate to CPU cluster (coarse-grained);
[0175] Perform shallow water equation calculations (5km resolution) to simulate large-scale tidal movements.
[0176] Correct the pipeline queue (even-numbered IDs: 2, 4, ..., 24) → allocate to the FPGA coprocessor (fine-grained);
[0177] Perform turbulence model calculations (0.5km resolution) to capture nearshore surge details.
[0178] Time slice-hardware mapping table:
[0179] Time Slice ID Computing unit type Binding Task Instructions 1,3,...,23 CPU cluster (Xeon Gold 6348) Solving shallow water equations (time step 10s) 2,4,...,24 FPGA (Xilinx Alveo U280) Turbulence model calculation (time step 0.1s)
[0180] Input source: The marine subdomain in the mesh topology file is marked as "FPGA preferred" (because the 0.5km mesh requires hardware acceleration).
[0181] Pipeline synchronization signal sequence, synchronization mechanism:
[0182] A global MPI_Barrier is triggered at the end of each time slice (e.g., when ID=1 ends, all processes are synchronized).
[0183] A hardware ready signal is inserted when switching between odd and even queues (FPGA→CPU switching delay ≤ 5ms).
[0184] Command example:
[0185] {Trigger time: 03:00:00, Command: MPI_Barrier, Target process group: ALL};
[0186] {Trigger time: 03:00:05, Instruction: FPGA_Ready, Target process group: CPU_Group1};
[0187] Fault tolerance time window parameter set, timeout threshold (based on hardware benchmark tests):
[0188] Task type Theoretical time Maximum allowable delay (1.2 × theoretical value) CPU shallow water equation 250s ≤300s FPGA turbulence model 40ms ≤50ms
[0189] Monitoring logic: If the CPU task with ID=5 times out for 300 seconds, the hot migration module will be triggered immediately.
[0190] Key data flow and output: Inputs: mesh topology file (including subdomain hardware identifiers) and hardware benchmark report (CPU / FPGA latency parameters).
[0191] Output: Dual-pipeline scheduling instruction set (HDF5 format): including time slice mapping table, synchronization sequence, and fault tolerance threshold.
[0192] Resource utilization: CPU tasks occupy 70% of cluster resources, and FPGA tasks consume ≤200W / card.
[0193] 203. Based on the hierarchical variance statistics table of the historical storm surge set, the priority signal is output through the variance comparator to generate node migration control instructions, and the computing nodes are migrated from the low variance level to the high variance level in real time to obtain the multi-level Monte Carlo (MLMC) resource dynamic configuration table.
[0194] Specifically, the Multi-Level Monte Carlo (MLMC) resource dynamic configuration table includes:
[0195] Hierarchy-Node Mapping Matrix: Structure: Two-dimensional array [Hierarchy ID][Node Physical Address]; Content: List of physical addresses of computing nodes corresponding to the coarse, medium, and fine levels; Real-time recording of node migration status (online / migration in progress / fault); Input source is the priority signal output by the variance comparator;
[0196] Node migration control instruction queue: Structure: Time-sequential instruction set {timestamp, source level ID, target level ID, number of migrated nodes}; Content: Migration instructions are generated in response to variance priority signals; Instruction execution cycle ≤ 1μs (hardware-level response); Used to directly control the migration of physical computing nodes; Input source is the 3-bit priority code of the variance comparator;
[0197] Resource utilization monitoring vector: Structure: Real-time sampling array [layer ID][CPU%, memory%, network%]; Content: Collects hardware resource utilization rate of each layer per second; Data feedback updates the historical variance statistics table; Input source is the performance counter of the compute node;
[0198] It should be noted that for the typhoon storm surge ensemble forecast in the Bohai Sea region, a three-level MLMC architecture (coarse, medium, and fine) is adopted, with an initial allocation of 128 calculation nodes (80 nodes in the coarse level, 40 nodes in the medium level, and 8 nodes in the fine level). The historical variance statistics table shows that the variance in the Laizhou Bay area is significantly higher than that in other areas.
[0199] Hierarchical variance statistics and priority generation, historical variance statistics table (based on 10 years of storm surge data):
[0200] Hierarchical ID Area Mean variance Priority signals 0 (coarse) The entire Bohai Sea 0.12 Low (01) 1 (Chinese) Bohai Bay + Laizhou Bay 0.28 (10) 2 (details) Liaodong Bay Key Area 0.45 High (11)
[0201] Variance comparator: If the variance of the fine-level layer exceeds the threshold (>0.4), output a 3-bit priority code "110" to indicate that the migration target is the fine-level layer (ID=2).
[0202] Node migration control command generation, migration command queue (response time ≤ 1μs):
[0203] {Timestamp: 2025-06-19T08:30:00.000Z, Source level: 0, Target level: 2, Number of migrated nodes: 5};
[0204] {Timestamp: 2025-06-19T08:30:00.002Z, Source level: 1, Target level: 2, Number of migrated nodes: 3};
[0205] Migration logic: Migrate 8 nodes from the low variance level (coarse / medium) to the high variance fine level to ensure that the number of fine level nodes increases to 16, thereby improving the calculation accuracy of key areas.
[0206] Hierarchical-node mapping matrix update:
[0207] Hierarchical ID Node physical address (partial example) Migration status 0 Node001~Node075 Online 1 Node076~Node112 Online 2 Nodes 113-128 (New) Migrating in progress → Online
[0208] Real-time status: After the migration is completed, the fine-level node addresses are expanded to Node081~Node128, supporting turbulence simulation at a resolution of 1km.
[0209] Resource utilization monitoring and feedback, monitoring vector (sampled 1 second after migration):
[0210] Hierarchical ID CPU% Memory% network% 0 65% 50% 30% 1 75% 60% 40% 2 92% 85% 70%
[0211] Dynamic feedback: When the fine-level CPU utilization approaches the threshold (95%), the variance statistics table is updated—the fine-level variance drops to 0.38, validating the effectiveness of the migration strategy.
[0212] 204. Based on the real-time wind field data stream and grid topology file, in the FPGA coprocessor: call the KL decomposition hardware core to calculate the eigenbase of the covariance matrix, generate random coefficients through the polynomial chaotic expansion (PCE) hardware core, perform matrix multiplication and addition operations to output physical constraint perturbation field data blocks, and obtain a spatially discretized random perturbation field dataset.
[0213] Specifically, the spatially discretized random perturbation field dataset includes:
[0214] Perturbation vector bound to grid coordinates: Structure: Spatial discretized data block {longitude, latitude, wind speed perturbation value, wind direction perturbation value}; Content: Strictly aligned with subdomain coordinates in the grid topology file; Physical constraint range: wind speed perturbation [-15m / s, +15m / s], wind direction perturbation [-30°, +30°]; Input source is the raw data output by the matrix multiplication and addition operation unit;
[0215] Hardware-based checksum calculation: Structure: 32-bit Cyclic Redundancy Check (CRC32); Content: Real-time generation based on perturbation vector data stream; Check polynomial: 0x04C11DB7; Input source: Generated by FPGA check circuit.
[0216] Data block partitioning index table: Structure: key-value mapping {subdomain ID, memory start address, data length}; Content: inherits the subdomain partitioning scheme from the mesh topology file; defines the physical address directly accessed by the MPI process; input source is the mesh topology file parsing result;
[0217] It should be noted that the real-time forecast of Typhoon Muifa in the Yangtze River Estuary region (121°-123°E, 30°-32°N) is based on a random perturbation field generated from a 0.5km resolution grid topology file.
[0218] Input data and hardware configuration; input source: real-time wind field data stream (10-minute interval): from coastal buoy stations and meteorological satellites, including wind speed / direction observations.
[0219] Mesh topology file: Mesh density at the land-sea boundary is 1.25km (land) vs. 0.5km (sea), and the subdomain ID and memory address mapping relationship is predefined.
[0220] Hardware platform: Xilinx Alveo U280 FPGA, integrating KL decomposition core (peak computing power 52 TOPS) and PCE random core (supports 100,000 chaotic expansions per second).
[0221] FPGA core computing process, KL feature basis generation: Calculate the wind field covariance matrix (dimension 1024×1024), and decompose it into feature vectors (first 50 principal components, cumulative contribution rate ≥95%) through KL hardware kernel.
[0222] Output characteristic basis matrix (single-precision floating point, occupying 18MB of FPGA on-chip memory).
[0223] PCE random coefficient generation: Based on Hermite orthogonal polynomial expansion, a random coefficient vector (dimension 50×1) following a Gaussian distribution is generated. Physical constraints:
[0224] Wind speed disturbance range: ±15m / s (12.3m / s disturbance value);
[0225] Wind direction disturbance range: ±30° (-22.5° disturbance value);
[0226] Matrix multiplication and addition: Performs multiplication and addition operations on the characteristic basis matrix and the random coefficient vector (clock period 5ns), and outputs the original disturbance vector (grid coordinate binding, longitude 121.5° and latitude 31.2° correspond to wind speed disturbance +9.8m / s).
[0227] Output dataset generation, perturbation vector bound to grid coordinates:
[0228] Structural example:
[0229] {Longitude: 121.75, Latitude: 31.35, Wind speed disturbance: -7.2m / s, Wind direction disturbance: +15.3°};
[0230] Align the grid topology: 0.5km grid point of the sea area subdomain ID=203 (memory address 0x8A00).
[0231] Hardware CRC32 checksum: Check polynomial: 0x04C11DB7;
[0232] Real-time generation: Each data block (1024 grid points) outputs one CRC32 code (0x3DA7F1C2), which is transmitted to the entropy value comparator along with the data stream.
[0233] Data block partitioning index table:
[0234] Subdomain ID Memory start address Data length 203 0x8A00 64KB 204 0x9C00 48KB
[0235] 205. Based on the grid topology file, dual pipeline scheduling instruction set, MLMC resource dynamic configuration table, and random disturbance field dataset, start the MPI process group according to the grid topology file, execute dual pipeline tasks alternately according to the scheduling instruction set, and migrate computing nodes in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set.
[0236] Specifically, the original storm surge water level field time series set includes:
[0237] Subdomain water level field data block: Structure: 3D floating-point array [longitude index, latitude index, time step]; Content: Local water level value sequence calculated by each MPI process; Spatial resolution strictly aligned with the grid topology file; Input source is the calculation results of the MPI process group executing the dual pipeline task;
[0238] Cross-domain synchronization marker sequence: Structure: Timestamped checksum record {synchronization time, boundary checksum, source / target process ID}; Content: Records subdomain boundary data exchange events; Includes 32-bit CRC boundary checksum value; Input source is hardware log of MPI_Send / Recv communication interface;
[0239] Calculate the state snapshot set: Structure: Process memory image sequence {time point, memory address range, checksum}; Content: Complete process state saved at fixed intervals (≥5 minutes); Includes register values and stack snapshots; Input source is process state capture triggered by system timer;
[0240] It should be noted that the 72-hour storm surge forecast for the Zhoushan Islands region in the East China Sea (the grid topology file contains 128 subdomains, the dual pipeline instruction set is divided into 24 time slices, and the MLMC configuration table dynamically migrates nodes to the high variance level of the Zhoushan Islands).
[0241] MPI process group startup and task execution, process allocation: 128 MPI processes are started based on the grid topology weight vector (52 processes are allocated to the sea area with a weight of 1.0; 76 processes are allocated to the land area with a weight of 0.4).
[0242] In response to the MLMC configuration table: migrate 8 nodes to the fine-grained level (original fine-grained level 16 nodes → 24 nodes), focusing on turbulence simulation of the Zhoushan Islands 0.5km grid.
[0243] Dual pipeline scheduling:
[0244] Time slice ID=1 (CPU task): Execute shallow water equation (resolution 5km) to calculate the initial water level field in Zhoushan sea area (time step 10s, output water level value range: -0.3m to +1.2m).
[0245] Time slice ID=2 (FPGA task): Execute turbulence model (resolution 0.5km) and correct nearshore surge (peak water level at grid point of Jintang Island, Zhoushan +2.5m, error <0.05m).
[0246] Water level field data and synchronization control, subdomain water level field data block: structure: three-dimensional array [longitude × 1200][latitude × 800][time step × 864] (72 hours, step size 5 minutes).
[0247] Example value:
[0248] Subdomain ID=57 (Zhoushan Island): [121.2°E, 30.0°N, t=12h] Water level +1.8m;
[0249] Subdomain ID=62 (Daishan Island): [122.1°E, 30.3°N, t=24h] Water level +2.3m;
[0250] Cross-domain synchronization marker sequence: Synchronization event: MPI_Barrier is triggered at the end of each time slice (ID=1 ends at 03:00:00), and data is exchanged at the land and sea subdomain boundaries.
[0251] Example of verification: {Synchronization time: 2025-06-19T03:00:00.005Z, Boundary checksum: 0x8C3F, Source process: 57, Target process: 62};
[0252] CRC check: 32-bit check value verifies the consistency of boundary data (error tolerance ±0.001m).
[0253] Fault tolerance and state snapshot mechanism, calculating the state snapshot set:
[0254] Capture rules: Save the full process memory image every 5 minutes (single snapshot size ≥ 2GB).
[0255] Hot migration support: When the entropy comparator detects a node failure (CRC anomaly of the process with subdomain ID=57), it resumes calculation within 82ms based on the most recent snapshot (t=12h), with water level data deviation <0.01m.
[0256] Resource monitoring feedback: The CPU utilization of fine-level nodes is 98% (close to the upper limit), triggering MLMC dynamic adjustment: 3 new nodes are migrated to the Zhoushan Islands level.
[0257] 206. Based on the water level field time series set and the pre-stored computation node state entropy fingerprint database, abnormal data blocks are detected by the entropy value comparator, triggering the process hot migration module to resume computation within ≤100ms, integrating effective data to generate probabilistic forecast result files, and obtaining the storm surge set forecast result after fault tolerance processing.
[0258] Specifically, the storm surge ensemble forecast results obtained after fault tolerance processing include:
[0259] Storm surge probability distribution map: Structure: Two-dimensional floating-point matrix [longitude][latitude]; Content: Spatially gridded water level exceedance probability values (0.0~1.0); Key risk area identifiers calculated based on the 95th quantile; Used to input flood risk maps into disaster prevention and early warning systems; Downstream application: Coastal emergency evacuation decision-making system; Input source is the quantile calculation results of integrated effective data blocks;
[0260] Abnormal Event Log Table: Structure: Spatiotemporal coordinate event set {timestamp, longitude, latitude, entropy difference, fault type code}; Content: Records all abnormal events detected by the entropy value comparator; Includes hardware fault classification (CPU / memory / network); Used to drive fault diagnosis of the supercomputing operation and maintenance platform; Downstream application: Computing node hardware health management system; Input source is the differential output result of the entropy value comparator;
[0261] Hot migration audit log: Structure: Verification report {migration start time, end time, source node ID, target node ID, data checksum}; Content: Records the time taken for each hot migration (≤100ms); Includes data consistency verification before and after migration; Used to verify the real-time performance of the fault-tolerant system; Downstream application: Parallel computing system optimization module; Input source is the execution record of the process hot migration module;
[0262] It should be noted that the 72-hour ensemble forecast (ensemble members 50 groups) for Typhoon Haikui in the Yangtze River Estuary region is based on the water level field time series set (resolution 0.5km, time step 5 minutes) output in step 205, and the reliability of the forecast is ensured through a fault-tolerant mechanism.
[0263] Entropy anomaly detection and thermal migration recovery, entropy fingerprint database comparison: pre-stored normal state entropy value range of computing nodes (CPU temperature entropy value 0.2~0.5, memory usage entropy value 0.1~0.3).
[0264] The entropy difference of 0.32 (normal threshold <0.25) was detected in subdomain ID=57 (Waigaoqiao area, Pudong), which is determined to be a memory fault (fault code MEM_ERR).
[0265] Hot migration trigger: The process hot migration module migrates the task from the faulty node (Node057) to the standby node (Node128) within 82ms.
[0266] Data consistency verification: The boundary CRC32 checksums are consistent before and after migration (0x8C3F→0x8C3F), and the water level data deviation is <0.01m.
[0267] Probabilistic forecast results are generated, and a storm surge probability distribution map is created. Data integration: After removing outlier data blocks, 50 sets of effective water level fields are integrated, and the exceedance probability of each grid point is calculated.
[0268] Risk zone marker: Wusongkou (longitude 121.5°E, latitude 31.3°N): 95th percentile water level 3.2m (1.0m above warning level), marked as a red risk zone.
[0269] Chongming Island East Beach (longitude 121.9°E, latitude 31.6°N): Overrun probability 0.15 (low risk area).
[0270] Output format: HDF5 format two-dimensional floating-point matrix (longitude × 1200, latitude × 800), directly input into Shanghai Disaster Prevention and Early Warning System.
[0271] Fault-tolerant system auditing and optimization, abnormal event log table: {timestamp: 2025-06-19T14:30:05Z, longitude: 121.5, latitude: 31.3, entropy difference: 0.32, fault type: MEM_ERR};
[0272] Downstream application: The operation and maintenance platform located the fault as the aging of the Node057 memory module, triggering a hardware replacement work order.
[0273] Hot migration audit log: {Start time: 14:30:05.003, End time: 14:30:05.085, Source node: 057, Target node: 128, Checksum: 0x5F2A}.
[0274] In this embodiment of the invention, the grid density in the land area is 30%-50% of that in the sea area, and a grid topology file with weighted coefficients is output, including a grid weight vector, an adjacency table, and a set of boundary mapping functions. This non-uniform grid partitioning and weight allocation method fully considers the different computational needs of land and sea areas, enabling more rational allocation of computational resources. The total duration is divided into multiple equal-length segments, constructing a prediction pipeline queue and a correction pipeline queue, and outputting a dual pipeline scheduling instruction set with time-slice identifiers, including a time-slice-hardware mapping table, a pipeline synchronization signal sequence, and a fault-tolerant time window parameter set. This dual-pipeline scheduling mechanism enables the alternating execution of coarse-grained and fine-grained computational tasks. Based on the hierarchical variance statistics table of historical storm surge sets, a priority signal is output through a variance comparator to generate node migration control instructions, which migrate computing nodes from low-variance levels to high-variance levels in real time, resulting in a multi-layer Monte Carlo resource dynamic configuration table. In the FPGA coprocessor, the KL decomposition hardware core and the polynomial chaotic expansion hardware core are called to perform matrix multiplication and addition operations to output physical constraint perturbation field data blocks, resulting in a spatially discretized random perturbation field dataset. Based on the water level field time series set and the pre-stored computing node state entropy fingerprint database, an entropy value comparator detects abnormal data blocks, triggering the process hot migration module to resume computation within ≤100ms, integrating valid data to generate a probabilistic forecast result file, and generating a storm surge set forecast result after fault tolerance processing, such as a storm surge probability distribution map, an abnormal event record table, and a hot migration audit log.
[0275] The above describes the parallel computation method for storm surge ensemble numerical forecasting in embodiments of the present invention. The following describes the parallel computation device for storm surge ensemble numerical forecasting in embodiments of the present invention. Please refer to [link to relevant documentation]. Figure 3An embodiment of the parallel computing storm surge ensemble numerical forecasting device of the present invention includes: an acquisition module 301, used to acquire a digital elevation model dataset, generate a terrain curvature matrix through a curvature calculation module, and execute a grid partitioning algorithm to output a grid topology file, wherein the grid density of the land area is 30%-50% of that of the sea area; a processing module 302, used to divide the total duration into 2K equal-length segments according to the forecast duration parameters and the grid topology file, construct a prediction pipeline queue and a correction pipeline queue, and output a dual pipeline scheduling instruction set; a migration module 303, used to output a priority signal based on the hierarchical variance statistics table of historical storm surge ensembles, generate node migration control instructions, migrate computing nodes from low variance levels to high variance levels in real time, and obtain a dynamic resource allocation table; and a setting module 304, used to set the real-time wind field data stream and the grid topology file. In the FPGA coprocessor, the topology file is used to calculate the eigenbase of the covariance matrix by calling the KL decomposition hardware core, generate random coefficients through polynomial chaotic expansion, and perform matrix multiplication and addition operations to output a random perturbation field dataset. The sequence module 305 is used to start the MPI process group according to the grid topology file, the dual pipeline scheduling instruction set, the resource dynamic configuration table, and the random perturbation field dataset. It alternately executes dual pipeline tasks according to the dual pipeline scheduling instruction set, migrates computing nodes in response to the resource dynamic configuration table, and obtains the original storm surge water level field time series set. The allocation module 306 is used to detect abnormal data blocks through an entropy value comparator based on the original storm surge water level field time series set and the pre-stored computing node state entropy fingerprint library, triggering the process hot migration module to resume calculation within ≤100ms, and obtaining the storm surge set forecast result.
[0276] In this embodiment of the invention, the grid density in the land area is 30%-50% of that in the sea area. This differentiated grid partitioning effectively reduces unnecessary computation in the land area while ensuring computational accuracy, improving overall computational efficiency and reducing computational resource consumption. A prediction pipeline queue and a correction pipeline queue are constructed, and tasks are executed alternately through a dual-pipeline scheduling instruction set, fully utilizing computational resources and achieving parallelization of forecasting and correction. This significantly shortens forecast time and improves forecast timeliness. Based on the hierarchical variance statistics table of historical storm surge sets, computational nodes are migrated from low-variance levels to high-variance levels in real time, achieving dynamic resource allocation and enabling computational resources to be allocated according to local conditions. Flexible allocation based on actual computational needs improves resource utilization and further enhances computational efficiency. The use of KL decomposition and polynomial chaotic expansion hardware cores within the FPGA coprocessor fully leverages hardware acceleration, speeding up the calculation of covariance matrix eigenbases and the generation of random coefficients. This accelerates the generation of random perturbation field datasets and improves the overall computational performance of the forecasting process. Abnormal data blocks are detected by an entropy comparator, triggering a process hot migration module to resume computation within ≤100ms. This effectively ensures the stability and reliability of the computation process, avoiding computational interruptions or errors caused by abnormal data and improving the accuracy of forecast results.
[0277] above Figure 3 The parallel computing storm surge ensemble numerical forecasting device in this embodiment of the invention is described in detail from the perspective of modular functional entities. The following describes the parallel computing storm surge ensemble numerical forecasting device in this embodiment of the invention in detail from the perspective of hardware processing.
[0278] Figure 4 This is a schematic diagram of a parallel computing storm surge ensemble numerical forecasting device 400 provided in an embodiment of the present invention. The parallel computing storm surge ensemble numerical forecasting device 400 can vary significantly due to different configurations or performance characteristics. It may include one or more processors (e.g., one or more processors) and a memory 420, and one or more storage media 430 (e.g., one or more mass storage devices) for storing application programs 433 or data 432. The memory 420 and storage media 430 can be temporary or persistent storage. The program stored in the storage media 430 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the parallel computing storm surge ensemble numerical forecasting device 400. Furthermore, the processor 410 may be configured to communicate with the storage media 430 and execute the series of instruction operations in the storage media 430 on the parallel computing storm surge ensemble numerical forecasting device 400.
[0279] The parallel computing storm surge ensemble numerical forecasting device 400 may also include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4 The illustrated structure of a parallel computing storm surge ensemble numerical forecasting device does not constitute a limitation on parallel computing storm surge ensemble numerical forecasting devices, which may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0280] The present invention also provides a parallel computing storm surge ensemble numerical forecasting device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor performs the steps of the parallel computing storm surge ensemble numerical forecasting method described in the above embodiments.
[0281] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the parallel computing storm surge ensemble numerical forecasting method.
[0282] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0283] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0284] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A parallel computing method for ensemble numerical forecasting of storm surges, characterized in that, The parallel computing-based numerical forecasting method for storm surge ensembles includes: Obtain the digital elevation model dataset, generate a terrain curvature matrix through the curvature calculation module, execute a grid generation algorithm to output a grid topology file, where the grid density of the land area is 30%-50% of that of the sea area; Based on the forecast duration parameters and grid topology file, the total duration is divided into 2K equal-length segments, and a prediction pipeline queue and a correction pipeline queue are constructed to output a dual pipeline scheduling instruction set; Based on the hierarchical variance statistics table of historical storm surge sets, a priority signal is output, node migration control instructions are generated, and computing nodes are migrated from low variance levels to high variance levels in real time to obtain a dynamic resource allocation table. Based on the real-time wind field data stream and grid topology file, in the FPGA coprocessor: the KL decomposition hardware core is called to calculate the eigenbase of the covariance matrix, random coefficients are generated through the polynomial chaotic expansion hardware core, and matrix multiplication and addition operations are performed to output a random perturbation field dataset. Based on the grid topology file, dual pipeline scheduling instruction set, resource dynamic configuration table, and random disturbance field dataset, the MPI process group is started according to the grid topology file, and dual pipeline tasks are executed alternately according to the dual pipeline scheduling instruction set. The computing nodes are migrated in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set. Based on the original storm surge water level field time series set and the pre-stored computation node state entropy fingerprint database, abnormal data blocks are detected by the entropy value comparator, triggering the process hot migration module to resume calculation within ≤100ms, and obtaining the storm surge set forecast results.
2. The parallel computing-based numerical forecasting method for storm surge ensembles according to claim 1, characterized in that, The mesh topology file includes a mesh weight vector, an adjacency table, and a set of boundary mapping functions; In the grid weight vector, the weight coefficient of the land region subdomain ranges from 0.3 to 0.5, while the weight coefficient of the sea region subdomain is fixed at 1.
0. Computational resources are allocated when the MPI process group is started. The adjacency table includes a set of Cartesian coordinates recording the shared boundaries between non-uniform subdomains and metadata of the MPI communication interface. The adjacency table defines the data exchange path between MPI processes. The boundary mapping function set includes a bilinear interpolation rule base at the land-sea junction and a conserved flux transfer coefficient matrix. The boundary mapping function set ensures the continuous transfer of physical quantities at the subdomain boundary.
3. The parallel computing-based numerical forecasting method for storm surge ensembles according to claim 1, characterized in that, The dual-pipeline scheduling instruction set includes: The time slice-hardware mapping table uses a set of key-value pairs: odd-numbered time slice IDs are bound to CPU clusters, and even-numbered time slice IDs are bound to FPGA coprocessors. The pipeline synchronization signal sequence adopts a time-sequenced instruction queue. Cross-process synchronization is triggered when each time slice is completed. Hardware-level ready signals are inserted between the odd and even queues to eliminate pipeline wait dependencies. The fault tolerance time window parameter set is configured using a threshold. .
4. The parallel computing-based numerical forecasting method for storm surge ensembles according to claim 1, characterized in that, The dynamic resource configuration table includes: The hierarchy-node mapping matrix is a two-dimensional array with a list of physical addresses of computing nodes corresponding to the three levels: coarse, medium, and fine. It records the node migration status in real time, including online, migrating, and faulty. The node migration control instruction queue is structured as a time-sequential instruction set. It generates migration instructions in response to variance priority signals, directly controlling the migration of physical computing nodes. The resource utilization monitoring vector is structured as a real-time sampling array. It collects the hardware resource utilization rate of each level every second, and updates the historical variance statistics table with the data feedback to provide a benchmark for resource anomaly detection.
5. The parallel computing method for ensemble numerical forecasting of storm surges according to claim 1, characterized in that, The random perturbation field dataset includes: The perturbation vector bound to the grid coordinates is structured using spatially discretized data blocks, strictly aligned with the subdomain coordinates in the grid topology file. Physical constraints include wind speed perturbation and wind direction perturbation. The hardware-calculated checksum is a 32-bit cyclic redundancy checksum, generated in real time based on the perturbation vector data stream, providing a data integrity benchmark. The data block partitioning index table has a key-value mapping structure, inherits the subdomain partitioning scheme of the grid topology file, defines the physical address that the MPI process can directly access, and eliminates the data copying overhead of the MPI process.
6. The parallel computing-based numerical forecasting method for storm surge ensembles according to claim 1, characterized in that, The obtained original storm surge water level field time series set includes: The subdomain water level field data block uses a three-dimensional floating-point array, and the local water level value sequence calculated by each MPI process has a spatial resolution that is strictly aligned with the grid topology file. Cross-domain synchronization marker sequence, using timestamp-based verification records, records subdomain boundary data exchange events, including a 32-bit CRC boundary check value; The computation state snapshot set uses a process memory image sequence to save the complete process state at fixed intervals, including register values and stack snapshots.
7. The parallel computing-based numerical forecasting method for storm surge ensembles according to claim 1, characterized in that, The storm surge ensemble forecast results include: The storm surge probability distribution map uses a two-dimensional floating-point matrix, including spatially gridded water level exceedance probability values and key risk area markers, and is used to input the disaster prevention and early warning system to release inundation risk maps; The abnormal event log table is structured as a spatiotemporal coordinate event set, recording all abnormal events detected by the entropy value comparator, including hardware fault classification, and is used to drive fault diagnosis of the supercomputing operation and maintenance platform. The live migration audit log, using a verification report structure, records the time taken for each live migration and includes data consistency verification before and after the migration, used to verify the real-time performance of the fault-tolerant system.
8. A parallel computing-based storm surge ensemble numerical forecasting device, characterized in that, The parallel computing storm surge ensemble numerical forecasting device includes: The acquisition module is used to acquire the digital elevation model dataset, the curvature calculation module generates the terrain curvature matrix, and the grid generation algorithm is executed to output the grid topology file, where the grid density of the land area is 30%-50% of that of the sea area; The processing module is used to divide the total duration into 2K equal-length segments based on the forecast duration parameters and grid topology file, construct the prediction pipeline queue and the correction pipeline queue, and output a dual pipeline scheduling instruction set; The migration module is used to output priority signals and generate node migration control instructions based on the hierarchical variance statistics table of historical storm surge sets, and migrate computing nodes from low variance levels to high variance levels in real time to obtain a dynamic resource configuration table. The configuration module is used to calculate the eigenbase of the covariance matrix by calling the KL decomposition hardware core in the FPGA coprocessor based on the real-time wind field data stream and grid topology file, generate random coefficients by using the polynomial chaotic expansion hardware core, and perform matrix multiplication and addition operations to output a random perturbation field dataset. The sequence module is used to start the MPI process group according to the grid topology file, the dual pipeline scheduling instruction set, the resource dynamic configuration table, and the random disturbance field dataset. It executes the dual pipeline tasks alternately according to the dual pipeline scheduling instruction set, and migrates the computing nodes in response to the resource dynamic configuration table to obtain the original storm surge water level field time series set. The allocation module is used to detect abnormal data blocks by using an entropy value comparator based on the original storm surge water level field time series set and the pre-stored computation node state entropy fingerprint database, and trigger the process hot migration module to resume calculation within ≤100ms to obtain the storm surge set forecast results.
Citation Information
Patent Citations
Storm surge set data forecasting method and device based on GPU parallel computing
CN112147719A
Tidal river reach storm surge rapid forecasting method based on deep learning and AI large model
CN119200041A