Earthquake vibration model construction and evaluation method based on big data
By intelligently acquiring and deeply managing multi-source heterogeneous seismic data, a three-layer architecture for an earthquake big data platform is constructed. Combining finite difference and finite element methods, and adopting a hybrid framework of seismology and deep learning, the platform solves the problems of easy interruption of seismic data acquisition links, inefficient resource scheduling, and low model coupling. This enables efficient management and high-precision risk assessment, providing reliable data support for emergency decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI TITANIUM ROBOT CO LTD
- Filing Date
- 2025-12-04
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies suffer from issues such as easy interruption of earthquake data acquisition links, inefficient resource scheduling, difficulty in quantifying data reliability, inability of platform architecture to adapt to heterogeneous data storage and multi-role requirements, low coupling of earthquake models, single evaluation dimensions, weak prediction of extreme conditions, and lagging platform and model updates.
By intelligently acquiring and deeply managing multi-source heterogeneous seismic data, a dynamically updated management rule base is constructed. The Lasso algorithm is used to schedule hardware resources, and the Raft consensus algorithm is used to ensure data consistency. A three-layer architecture for the seismic big data platform is built. Multi-scale site models are constructed by combining finite difference and finite element methods. The model is trained using a hybrid framework of seismological theory and deep learning. Multi-dimensional model evaluation and dynamic risk assessment are carried out, and the platform and model are iteratively optimized through feedback.
It has achieved a continuous supply of high-quality data sources, improved data management efficiency and model simulation accuracy, provided multi-scenario services and high-precision risk assessment, provided reliable data support for emergency decision-making, and ensured that the technical solution is adapted to changes in demand and geological conditions in the long term.
Smart Images

Figure CN121996404A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of earthquake engineering technology, specifically to a method for constructing and evaluating earthquake motion models based on big data. Background Technology
[0002] In the field of earthquake monitoring and disaster prevention and mitigation, with the development of sensor technology, the Internet of Things and deep learning, earthquake-related data has formed a complex system of multiple sources and heterogeneity, covering real-time waveform data of earthquake monitoring networks, stratigraphic and fault parameters from geological exploration, BIM models of engineering structures and historical earthquake damage records, etc.
[0003] However, existing technologies have significant bottlenecks in practical applications: Firstly, the data acquisition process relies on a single transmission link, which is prone to data interruption due to equipment failure or environmental interference. Furthermore, it lacks an intelligent resource scheduling mechanism, resulting in hardware resource utilization of less than 50%. Additionally, the handling of conflicts between multiple data sources largely depends on manual verification, which is inefficient and cannot quantify data credibility, making it difficult to form a high-quality data source. Secondly, existing platform architectures mostly adopt a single storage mode, which cannot adapt to the differentiated storage needs of structured seismic parameters and unstructured waveform files. Furthermore, data processing and application services are disconnected, making it difficult to respond quickly to the needs of researchers, emergency departments, and other stakeholders. Third, earthquake motion models are mostly constructed separately as site or structural models, without deep coupling. The simulation accuracy can only meet the needs of macro-regional analysis and cannot support the detailed assessment of key local areas. Furthermore, model evaluation only focuses on simulation errors and ignores the matching degree between structural damage probability and risk level. At the same time, there is a lack of a mechanism to supplement extreme earthquake data, resulting in weak predictive ability for rare earthquake events. Fourth, the platform's functions and model parameter updates rely on manual triggering, making it unable to dynamically adapt to changes in geological conditions and new data. After long-term use, the accuracy deteriorates significantly, making it difficult to provide reliable support for earthquake research and emergency decision-making. Therefore, there is an urgent need for an integrated construction method that covers the entire data lifecycle to solve the above problems.
[0004] To address the aforementioned problems, this invention provides a method for constructing and evaluating earthquake motion models based on big data. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for constructing and evaluating earthquake motion models based on big data. This method solves the problems of single and easily interrupted multi-source earthquake data acquisition links, inefficient resource scheduling, difficulty in quantifying data reliability, platform architecture that cannot adapt to heterogeneous data storage and multi-role requirements, low coupling of earthquake models, single evaluation dimensions, weak prediction of extreme conditions, and lagging platform and model updates.
[0006] To achieve the above objectives, one aspect of the present invention is to provide a method for constructing and evaluating earthquake motion models based on big data, including: Intelligent acquisition and in-depth governance of multi-source heterogeneous seismic data: Seismic monitoring data, geological and source data, structural and disaster data are collected and transmitted to the data center through wired and wireless dual links. Hardware resources are scheduled based on the Lasso algorithm, and the Raft consensus algorithm is used to ensure data consistency. The collected data is cleaned, conflict is handled, and a unified format and coordinate system are implemented. Credibility weight labels are added, and a governance rule base is built and dynamically updated. A three-layer architecture for the earthquake big data platform is constructed. The storage layer uses relational and non-relational databases to classify and store structured and unstructured data, integrates a data warehouse and links it to a metadata center. The processing layer deploys association rule algorithms, clustering algorithms and an LSTM / GRU-based earthquake hazard prediction model, and allocates processing tasks in conjunction with a resource scheduling model. The application layer provides differentiated functions for researchers, emergency rescue departments and the public. A multi-source fusion seismic motion model was constructed, and a multi-scale site model was built by combining the finite difference and finite element methods. A structure-site coupled sub-model was constructed for building complexes and power transmission towers. The model was trained using a hybrid framework of seismological theory and deep learning, and the model parameters were corrected in stages. Multi-dimensional model evaluation and dynamic risk assessment: dividing the training set and test set, expanding the sample using Latin hypercube sampling, evaluating the model through indicators such as simulation accuracy and response spectrum consistency, calculating the structural failure probability based on Poisson binomial distribution, simulating multi-intensity earthquake scenarios and optimizing substandard models; The platform and model are optimized through feedback iteration, with regular updates to data and metadata. Platform functions are adjusted based on user feedback, and the model is retrained with full data every two years. Extreme working condition data is supplemented based on generative adversarial networks, and the seismic velocity model is dynamically updated.
[0007] The second invention provides a system for constructing an earthquake big data platform based on data collection and governance, used to realize an earthquake motion model construction and evaluation system based on big data, including the following modules: The module includes data acquisition and governance, platform architecture, model building, evaluation and assessment, and iterative optimization. Data acquisition and governance module: Collects earthquake monitoring, geological and seismic source, structural and disaster data, transmits them to the data center through wired and wireless dual links, schedules hardware resources based on Lasso algorithm, uses Raft consensus algorithm to ensure data consistency, cleans data, handles conflicts, unifies format and coordinate system, adds credibility weight labels and dynamically updates governance rule base; Platform architecture modules: A three-tier architecture is constructed. The storage layer uses relational and non-relational databases to classify and store structured and unstructured data, and integrates a data warehouse with a metadata center. The processing layer deploys association rules, clustering algorithms, and LSTM / GRU seismic hazard prediction models, and allocates tasks in conjunction with a resource scheduling model. The application layer provides differentiated functions for researchers, emergency rescue departments, and the public. Model building module: Combines finite difference and finite element methods to build multi-scale site models, constructs structure-site coupled sub-models for building complexes and power transmission towers, and trains the model using a hybrid framework of seismic theory and deep learning and corrects parameters in stages; Evaluation and assessment module: Divide the training set and test set, expand the sample by Latin hypercube sampling, evaluate the model by simulation accuracy and response spectrum consistency index, calculate the structural failure probability based on Poisson binomial distribution, simulate multi-intensity earthquake scenarios and optimize substandard models; Iterative optimization module: Regularly updates data and metadata, adjusts platform functions based on user feedback, retrains the model with full data every 2 years, supplements extreme working condition data based on generative adversarial networks, and dynamically updates the earthquake velocity model.
[0008] The third aspect of the present invention is to provide a computer device, the computer device comprising: a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by the processor to realize a method for constructing and evaluating earthquake vibration models based on big data.
[0009] The fourth aspect of this invention is to provide a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a method for constructing and evaluating earthquake vibration models based on big data.
[0010] Compared with existing technologies, this invention provides a method for constructing and evaluating earthquake motion models based on big data, which has the following beneficial effects: This solution utilizes intelligent acquisition and in-depth governance of multi-source heterogeneous seismic data. It ensures data transmission continuity through a dual-link system of wired and wireless connections, optimizes hardware resource scheduling using the Lasso algorithm, and ensures data consistency using the Raft consensus algorithm. Simultaneously, it improves data quality through hierarchical conflict handling, format standardization, and credibility weight labeling. It also constructs a dynamically updated governance rule base to address the issues of diverse and poor-quality data sources, providing high-quality data sources for subsequent stages. This solution constructs a three-layer architecture for an earthquake big data platform. The storage layer categorizes and stores structured and unstructured data and associates it with a metadata center. The processing layer deploys algorithms and prediction models and allocates tasks in conjunction with resource scheduling. The application layer provides differentiated functions, which solves the problems of poor platform storage adaptability and service disconnect, and achieves efficient data management and multi-scenario services. This scheme is based on building a multi-source fusion seismic motion model. It improves the model coupling by using a multi-scale site model and a structure-site coupled sub-model. It optimizes the model accuracy with a hybrid framework of "seismic theory + deep learning" and phased parameter correction, thus solving the problem of rough model simulation and providing a high-precision foundation for risk assessment. This solution employs multi-dimensional model evaluation and dynamic risk assessment. It ensures comprehensive evaluation through sample expansion, verifies model reliability using multi-dimensional indicators, and combines probability calculations and multi-intensity simulations to output risk results. This addresses the issues of singular model evaluation and inaccurate risk assessment, providing data support for emergency decision-making. Furthermore, the solution utilizes feedback-based iterative optimization of the platform and model, regularly updating data and functions, retraining the model, and generating adversarial networks to supplement extreme condition data. This resolves the issue of lagging platform and model updates, ensuring the technical solution remains adaptable to long-term changes in needs and adjustments in geological conditions. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a diagram of the overall system module architecture of the present invention; Figure 2 This is an internal flowchart of the data acquisition and management module of the present invention; Figure 3 This is a flowchart of the model construction and evaluation method of the present invention. Figure 4 This is the internal logic diagram of the iterative optimization module of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0014] To address issues such as the single and easily interrupted acquisition link of multi-source seismic data, inefficient resource scheduling, difficulty in quantifying data reliability, platform architecture inability to adapt to heterogeneous data storage and multi-role requirements, low coupling of seismic models, single evaluation dimensions, weak prediction of extreme conditions, and lagging platform and model updates.
[0015] This proposal suggests a method for constructing and evaluating earthquake motion models based on big data.
[0016] By intelligently acquiring and deeply processing multi-source heterogeneous seismic data, a dynamically updated governance rule base is constructed to solve the problems of mixed data sources and poor data quality, providing a high-quality data source for subsequent stages. By constructing a three-layer architecture for the earthquake big data platform, the problems of poor platform storage adaptability and service disconnect are solved, enabling efficient data management and multi-scenario services. By establishing a multi-source fusion seismic motion model, the problem of rough model simulation is solved, providing a high-precision basis for risk assessment; By using multi-dimensional model evaluation and dynamic risk assessment, the problems of single model evaluation and inaccurate risk assessment are solved, providing data support for emergency decision-making. By using feedback-based iterative optimization of the platform and model, the problem of lagging platform and model updates is solved, ensuring that the technical solution can adapt to changes in demand and geological conditions in the long term.
[0017] Example 1, as Figures 1-4 The following is an example illustrating a method for constructing and evaluating earthquake motion models based on big data, provided in an embodiment of this application.
[0018] S100: Collects three types of data: earthquake monitoring, geology and seismic source data, and structural and disaster data, and transmits them to the data center via a dual-link (wired + wireless) connection. It utilizes the Lasso algorithm to schedule hardware resources and the Raft consensus algorithm to ensure data consistency. The system cleans and deduplicates data, handles conflicting data in a tiered manner, standardizes the format and coordinates, adds credibility weight labels, and builds and dynamically updates a governance rule base. S200: Constructs a three-tier architecture. The storage layer uses relational and non-relational databases to store heterogeneous data, integrating a data warehouse with a metadata center. The processing layer deploys association rules, clustering algorithms, and LSTM / GRU prediction models, combined with resource scheduling to allocate tasks. The application layer provides differentiated functions for scientific research, emergency response, and the public. S300: A multi-scale site model is established using the finite difference and finite element methods to construct a structure-site coupled sub-model for building complexes and power transmission towers. The model is trained using a framework of "seismic theory + deep learning," and parameters are corrected in stages using small earthquake data in the elastic stage and large earthquake data in the elastoplastic stage. The training and test sets were divided into a 7:3 ratio using S400, and the sample size was expanded using Latin hypercube sampling. The model was evaluated using metrics such as RMSE and R, and the structural failure probability was calculated using the Poisson binomial distribution. Multi-intensity earthquake scenarios were simulated, and substandard models were optimized at the data, feature, and model levels. S500: Updates data and metadata on a regular basis and adjusts platform functions based on user feedback. The model is retrained with full data every two years, supplemented with extreme condition data using GAN, and the earthquake velocity model is updated quarterly and within 72 hours after an earthquake of magnitude Mw≥5.
[0019] Step S100 is used to transform fragmented and low-reliability data into standardized and high-value data sources by constructing a "full-dimensional collection-intelligent scheduling-hierarchical governance" system, laying the foundation for subsequent platform construction and model training, including: S110: S1101: Fixed station data includes three-component seismic waveforms (velocity, acceleration), station coordinates (latitude, longitude, altitude), and sampling rate (100Hz-1000Hz), used to obtain seismic wave propagation characteristics; Mobile station data: High-frequency waveform data (sampling rate ≥2000Hz) collected by temporarily deployed portable sensors to supplement blind spots of fixed station networks; Teleseismic cross-convolution waveform data: generated by convolving teleseismic signals with reference station signals to improve the signal-to-noise ratio of weak seismic signals; High-frequency Rayleigh wave phase velocity data: extracted from surface wave signals using the Eikonal imaging method, reflecting shallow geological structures; S1102: Geological structure data: stratigraphic distribution (sand layer, clay layer, rock layer thickness), fault parameters (strike, dip angle, burial depth), obtained through geological exploration reports and drilling data; Soil physical parameters: shear wave velocity (Vs30, average shear wave velocity from the surface to a depth of 30m), density (2.0-2.8g / cm³), used for site classification; Source parameters: magnitude (moment magnitude Mw, surface wave magnitude Ms), focal depth, fault slip model (slippage amount, rupture velocity), derived from earthquake rapid reporting and post-processing reports; S1103: Engineering structural data: Building complex BIM model (including component dimensions and material strength), power transmission tower parameters (height, member cross-section, connection method), and structural natural vibration period (obtained through pulsation testing); Historical earthquake damage data: earthquake disaster records since 1970 (building collapse rate, number of casualties), incremental dynamic analysis (IDA) curve library (relationship between structural apex displacement and strength under different seismic wave inputs).
[0020] By classifying and sorting the data, we can clarify the collection frequency (e.g., real-time collection of data from fixed stations, and geological data updated every 5 years) and storage format (SEG-Y format for waveform data, IFC format for BIM models), providing a basis for the design of subsequent data collection links.
[0021] S120: S1201: The backbone network uses fiber optic transmission with a bandwidth of ≥100Mbps and a transmission delay of ≤50ms, covering fixed stations and regional data centers. The branch line uses industrial Ethernet to connect temporary deployment points of mobile stations with nearby fixed stations, supports PoE power supply (simultaneously transmitting data and power), and is suitable for outdoor environments without mains power. S1202: The main wireless link adopts a 5G private network with a transmission rate of ≥50Mbps, which is used for real-time data backhaul between mobile stations and data centers and supports continuous transmission in mobile scenarios. The backup wireless link uses BeiDou short message service and is activated when the 5G signal is interrupted to transmit critical data (such as magnitude and epicenter location). The length of a single message is ≤1000 bytes and the latency is ≤60 seconds. S1203: Deploy a link monitoring module to collect transmission rate, packet loss rate, and latency parameters in real time, and set thresholds (such as triggering switching when packet loss rate > 5% or latency > 100ms). The system employs the OSPF (Optical Routing Policy) protocol, which automatically switches to a backup link within 100ms when the primary link fails, ensuring uninterrupted data transmission.
[0022] S130: S1301: Extract feature parameters from historical acquisition tasks: task type (real-time acquisition / batch acquisition), data volume (MB / hour), duration (minutes), sampling rate (Hz); Hardware resource parameters: CPU frequency (GHz), memory capacity (GB), storage IOPS (times / second), network bandwidth (Mbps); S1302: The regression model is constructed using the Lasso algorithm, and the formula is as follows: ; Where y represents the predicted resource demand (such as CPU utilization). These are task feature parameters. λ is the characteristic coefficient, and λ is the regularization parameter (to control overfitting, with a value of 0.01-0.1). The model training dataset consists of 1000+ data collection task records from the past 3 years, with a cross-validation accuracy of ≥85%. S1303: Perform initial screening of available hardware nodes based on remaining resources (CPU remaining ≥30%, memory remaining ≥40%); ; Select nodes with a matching degree ≥ 0.8 to assign tasks; Dynamic adjustment: When the node resource utilization rate exceeds 80%, some subtasks will be automatically migrated to nodes with lower load.
[0023] S140: S1401: All acquisition nodes are equipped with BeiDou timing modules, with a synchronization accuracy of ≤1ms, to ensure consistent timestamps for earthquake events; Time calibration is performed every hour, and automatic correction is performed when the local time of the node deviates from the standard time by more than 5ms. S1402: Select 3-5 core stations as consensus nodes, with 1 as the master node and the rest as slave nodes; The master node receives and verifies the data, generates log entries (including data content, timestamp, and checksum), and synchronizes them to the slave nodes. After a slave node confirms receipt and returns an ACK, the master node marks the data as "committed" once it receives a majority (≥50%+1) of ACKs, ensuring data consistency across multiple nodes. S1403: Monitoring metrics: CPU utilization (threshold ≤ 80%), memory utilization (threshold ≤ 70%), data transmission latency (threshold ≤ 500ms), data integrity (checksum and matching); If a node fails to meet the standard three times in a row, it will be marked as "abnormal" and automatically switched to a backup node (the number of backup nodes is 30% of the core nodes).
[0024] S150: S1501: Identify duplicate data based on the triple key value of "event ID + station ID + timestamp", such as the same earthquake event being recorded repeatedly within 30 seconds at the same station; Keep the latest record, mark the rest as "duplicate" and archive them (do not delete, for traceability purposes); S1502: For short-term missing data (<5 seconds): use linear interpolation; The formula is ,in Here are the missing values, and t is the time of the missing value. For long-term missing data (≥5 seconds): fill in the missing data using data from adjacent stations; Calculate distance weights ( (Station spacing). Missing values This can only be used when there are ≥3 valid adjacent stations; S1503: Date format: uniformly convert to "YYYY-MM-DDHH:MM:SS.XXX" (accurate to milliseconds), such as correcting "2023 / 10 / 01 15:30" to "2023-10-01 15:30:00.000"; Coordinate format: Convert the Beijing 54 and Xi'an 80 coordinate systems to the WGS84 coordinate system. The conversion formula adopts the seven-parameter method (based on regional conversion parameters). Units are standardized: magnitude is rounded to one decimal place (e.g., Mw6.2), and acceleration is standardized to gal (1 gal = 1 cm / s²).
[0025] S160: S1601: The parameter differences of the same seismic event (time difference < 30 seconds, epicentral distance < 10 km) exceed the threshold: magnitude difference > 0.5, focal depth difference > 5 km, maximum acceleration difference > 20%; Geological data conflict: Shear wave velocity difference > 10% and stratum thickness difference > 20m at the same location; S1602: Conflict Intensity Calculation: Conflict intensity ; in, Let be the parameter value of the i-th station. The mean of the parameters, The station credibility weights are set based on historical data accuracy, ranging from 0.5 to 1.0, where n is the number of stations participating in the comparison. S1603: Low-intensity conflict (conflict intensity < 0.3): Automatically uses weighted average correction, the formula is as follows: ; Medium-intensity conflict (0.3 ≤ conflict intensity < 0.5): Combined with geological background verification, such as whether the fault distribution supports the difference in focal depth, the corrected version is marked "Geologically verified"; High-intensity conflict (conflict intensity ≥ 0.5): Triggers a manual review process, where at least 3 earthquake experts consult to determine the final value, and the result is marked "manual confirmation".
[0026] S170: S1701: Earthquake Event Table: Fields include event ID, time of occurrence (WGS84), epicenter coordinates (WGS84), magnitude (Mw), focal depth (km), and number of reference stations; Waveform data: The sampling rate is uniformly set to 100Hz (low sampling) or 1000Hz (high sampling). The waveform is extracted from 10 seconds before the event to 300 seconds after the event, and the instrument response is removed. Structural data: The BIM model is simplified into a structured table of "component ID + material strength + dimensions + location", which is linked to geographic coordinates; S1702: Reliability Weight: Calculated by weighting based on station accuracy (0.8-1.0), data integrity (0.6-1.0), and conflict resolution results (0.7-1.0), using the following formula: ; Quality grades: A (w≥0.8), B (0.6≤w<0.8), C (w<0.6), where grade C data is for reference only and is not used for model training; Mark the location: Add the "quality_weight" and "quality_level" fields to the header of the data file or in the database table, and let them flow with the data; S180: S1801: Cleaning rules: threshold for duplicate data identification, methods and applicable scenarios for missing value imputation, and format conversion formulas; Conflict handling rules: conflict detection threshold (e.g., magnitude difference of 0.5), conflict handling procedures at each level, and expert review trigger conditions; Standardization rules: field definitions for each data type, unit conversion factors, and methods for calculating quality marks; S1802: Collect 1000+ cases of manual governance from the past 5 years and extract general rules; Decision tree algorithms (such as ID3) are used to transform cases into rule entries, such as "if the conflict intensity is <0.3 and the number of stations is ≥5, then automatic weighted correction is performed"; Rules are stored in JSON format and contain fields such as "rule ID, trigger condition, execution action, and priority". S1803: Monthly statistics on rule execution performance (such as conflict handling accuracy). When the accuracy of a rule is less than 80%, it is marked as "to be optimized". When a new data source (such as a new type of sensor) is added, a matching rule suggestion is generated by matching rules similarity (calculating the cosine similarity between the new data features and existing rules), and then included in the library after being confirmed by experts. The rule base is optimized annually, removing redundant rules (execution frequency < 1 time / month) and merging duplicate rules.
[0027] High-quality data processed by S100 requires an efficient platform to achieve integrated management of storage, processing, and services. Traditional platforms, due to their rigid architecture, struggle to adapt to heterogeneous data and the needs of multiple roles. This step utilizes a three-tier architecture design of "storage-processing-application" to provide a full lifecycle carrier for data, while simultaneously providing computing resources and data interface support for the S300 model construction. Specifically: S210: S2101: Structured data: Utilizes a MySQL cluster (master-slave architecture) to store earthquake event parameters, station information, structural attributes, etc., supporting transaction processing and index queries, with a single table data volume of ≤10 million rows; Unstructured data: MongoDB is used to store waveform files (SEG-Y), geological images (TIFF), and BIM models (IFC), supporting the association of binary data and metadata, with a single file size ≤10GB; Time-series data: InfluxDB is used to store real-time monitoring data (such as 100 seismic waveforms per second), supporting high write throughput (≥100,000 records / second) and time range queries; Cold data archiving: Use tape libraries to store historical data (such as waveform files from 10 years ago), with an access frequency of ≤1 time / month, reducing storage costs; S2102: Thematic Classification: Includes earthquake event themes, geological structure themes, engineering structure themes, and disaster assessment themes; Star schema: Each topic is centered around a fact table (e.g., an earthquake event fact table) and associated with dimension tables (e.g., time dimension, spatial dimension, station dimension). Partitioning strategies: Partition by time (e.g., by year) or space (e.g., by latitude and longitude grid) to improve the efficiency of large table queries (query time reduced to within 1 second); S2103: Metadata type: Data source (station ID, file path), processing history (cleaning time, conflict resolution record), quality information (reliability weight, level), correlation (e.g., a waveform data corresponds to a certain earthquake event); Search function: Supports queries by combination of "data type + time range + quality level", returns data storage location and access permissions, and has a response time of ≤2 seconds; Version management: Records the version number, update time, and updater for each data update, and supports version rollback (e.g., restoring to the state 3 days ago).
[0028] S220: S2201: Association Rule Algorithm (Apriori): Mines the correlation between earthquake parameters and structural damage, such as "when the magnitude is greater than 7 and Vs30 is less than 180 m / s, the building collapse rate is greater than 50%", with the minimum support set at 5% and the minimum confidence set at 70%; K-means clustering algorithm: Clustering historical earthquake events, classifying them according to magnitude (5-6, 6-7, and above 7) and focal depth (shallow depth <70km, intermediate depth 70-300km), with the number of clusters k=5-8, used for sample stratification; Feature selection algorithm (elastic net): Selects key features from seismic wave features (peak value, frequency, duration), with regularization parameter α=0.5 (balancing L1 and L2 regularization), retaining the top 20% of features that have a significant impact on structural damage; S2202: Input features: historical earthquake frequency, fault activity rate, geological structural complexity (level 1-5), crustal stress value (MPa). Model architecture: LSTM neural network (128-dimensional input layer, 3 hidden layers × 64 neurons, 1-dimensional output layer), activation function is ReLU, optimizer is Adam (learning rate 0.001). Training process: A sliding window method (window size of 5 years) was used for training with data from 2000 to 2020 and validation with data from 2021 to 2023. The probability of earthquakes occurring in the target area in the next 5 years was predicted with an accuracy of ≥80%. S2203: Uses Spark Streaming to process real-time data streams, with a batch processing interval of 10 seconds and parallelism = number of CPU cores × 2; Deploy data quality verification operators: calculate data credibility weights in real time and filter data with w < 0.6; Result caching: Cache the results of frequently accessed processing (such as earthquake events in the last 24 hours) in Redis, and set the cache expiration time to 1 hour to reduce redundant calculations.
[0029] S230: S2301: Task Priority Assignment P0 level (emergency): Real-time earthquake rapid reporting and processing (response time < 1 minute), emergency rescue data processing; P1 Level (High): Model training tasks, daily data summary; Level P2 (Intermediate): Historical data reprocessing and report generation; P3 Level (Low): Log analysis, system maintenance; S2302: Call the resource demand prediction model of S103 to calculate the CPU and memory resources required for each task; By employing a priority queue scheduling algorithm, P0-level tasks can preempt resources from P1-P3-level tasks, ensuring that urgent tasks are executed first. Load balancing: When the CPU utilization of a node is greater than 85%, some P2 / P3 level tasks will be migrated to nodes with lower load (utilization < 60%). S2303: Automatic verification: When the deviation between the calculated result and historical data of the same type (such as the difference between the predicted magnitude and the actual magnitude) is greater than the threshold (such as 0.3), reprocessing is triggered. Cross-validation: The model output results are repeatedly calculated using different algorithms (such as LSTM and random forest), and the consistency rate of the results is ≥90% to be considered valid; Manual sampling: 5% of the processing results are randomly selected every week (with a focus on P0 / P1 level tasks) and reviewed by analysts. If problems are found, the processing process is traced back.
[0030] S240: S2401: Data Interface: Provides a RESTful API that supports downloading data based on conditions (such as "earthquake waveforms with Mw > 6 from 2010 to 2020"), with a concurrent interface capacity of ≥100 times / second; Model Experimentation Platform: Supports uploading custom models (such as improved LSTM), calling platform data for training and validation, and provides GPU resources (single card computing power ≥10TFLOPS). Feature analysis tools: Visualize seismic wave spectrum (Fourier transform) and geological structure profiles, supporting custom parameters (such as filter frequency range). S2402: Rapid Disaster Assessment: Input earthquake parameters (magnitude, epicenter), and output the affected area (based on population density and building vulnerability) and estimated number of casualties (error ≤ 20%) within 10 minutes. Hotspot area identification: The probability of building damage is displayed through a heat map (areas with a congestion index > 80% are considered high-risk areas), overlaid with real-time traffic conditions (roads with a congestion index > 0.7 are marked in red); Rescue route planning: Based on Dijkstra's algorithm, it avoids high-risk areas and congested sections, generating 3 optimal routes (shortest time, shortest distance, and highest safety factor), with a planning accuracy of ≥90%. S2403: Earthquake Science Popularization: Explaining the causes of earthquakes and disaster avoidance knowledge with a combination of text and images, and embedding short videos (≤5 minutes). Historical earthquake visualization: 3D display of the earthquake propagation process of Wenchuan earthquake (2008) and Yushu earthquake (2010), supporting zooming and view switching; Personal Inquiry: Enter your address to inquire about the earthquake risk level (1-5) and seismic resistance recommendations for your area (such as "strengthen beams and columns").
[0031] Leveraging the high-quality data and computing resources provided by the S200 platform, traditional earthquake models, suffering from site-structure separation and low simulation accuracy, urgently need to be addressed. This step improves the accuracy of earthquake response simulation by constructing a "multi-scale site-structure coupled model" and a "seismology + deep learning" hybrid training framework, providing core analytical tools for model evaluation and risk assessment in the S400 platform. The S300 platform includes: S310: S3101: Covers the target area and its surrounding 50km, with a horizontal range of ≥100km×100km and a vertical depth of ≥30km; The finite difference method is used to divide the grid, with horizontal grid size of 500m-1km (increasing with depth) and vertical grid size of 200m-500m, for a total of ≥1 million grids; S3102: Input the geological structure data of S101 and divide the model into layers according to strata (such as surface layer, sedimentary layer, bedrock layer). Assign physical parameters to each layer: longitudinal wave velocity (Vp), transverse wave velocity (Vs), density (ρ), and damping ratio (ξ), such as bedrock layer Vp=5000m / s, Vs=3000m / s, ρ=2.6g / cm³, ξ=0.02; S3103: Use a point source model or fault model and input the source parameters (magnitude, depth, rupture direction). Seismic wave incidence: The radiation patterns of P-waves and S-waves are set according to the focal mechanism, with a time step ≤ 0.01 seconds and a simulation duration ≥ 300 seconds; S3104: The side surfaces employ absorbing boundary conditions (such as Clayton-Engquist boundary) to reduce wave reflection; The bottom is a fixed boundary, and the top is a free surface (considering the surface amplification effect).
[0032] S320: S3201: For urban core areas or key project areas (such as 10km×10km), the horizontal grid size is 50m-100m, the vertical depth is ≥5km, and the total number of grids is ≥500,000; Grid-reinforced zones: In densely built-up areas or fault zones, the grid size is reduced to 20m to capture local vibration differences; S3202: Input the high-frequency Rayleigh wave phase velocity data from S110, and use the surface wave dispersion inversion method to obtain the Vs profile from the surface to a depth of 500m. Calculate the correction coefficient by comparing the initial Vs value with the macroscopic model. Correcting shallow mesh parameters ; S3203: Input digital elevation model (DEM) data, convert the surface undulations into the top boundary of the model, and refine the grid in areas with a slope > 15°; An equivalent linearization method is adopted to consider the wave scattering and amplification effects caused by topography (e.g., the amplification factor of 1.5-2.0 at the top of the mountain).
[0033] S330: S3301: Low-rise buildings (≤3 stories): Single-degree-of-freedom (SDOF) model is used, lumped mass m = total building mass, stiffness ; Multi-story buildings (4-10 stories): A multi-degree-of-freedom (MDOF) shear model is adopted, with each story as a single mass point, the mass concentrated in the floor, and the stiffness determined by lateral force resisting members (walls, columns); High-rise buildings (>10 stories): The MDOF bending-shear coupling model is adopted, considering translational and rotational coupling, and the stiffness matrix includes bending and shear contributions. S3302: Elastic stage: A linear elastic model is adopted, with the stiffness being the initial stiffness k0; Elastic-plastic stage: A bilinear model is adopted, with yield strength Fy=0.8Fu (Fu is the ultimate strength) and stiffness after yielding k1=0.1k0; Damage accumulation: Introducing damage variables When D≥0.8, it is determined to be a structural failure; S3303: The dynamic time history analysis method is adopted, and the ground acceleration time history output by the S320 fine model is used as input; Considering the interaction between the foundation and the subgrade (SSI effect), a spring-damped model is used for simulation, with spring stiffness ks = Gs / r0 (Gs is the shear modulus of the subgrade, and r0 is the equivalent radius of the foundation). S340: S3401: Using a spatial truss model, the iron tower is decomposed into nodes (foundation, tower body nodes, crossarm nodes) and members (main members, diagonal members, crossarms). The nodes are mass-concentrated, and the members are elastic rod elements. The cross-sectional properties (area, moment of inertia) are assigned according to the design drawings. S3402: Traveling wave effect: Seismic waves of different phases are input along the line direction, with wave velocity of 300-500m / s and phase difference Δϕ=2πd / λ (d is the tower spacing and λ is the wavelength). Coherence effect: The correlation of ground motion at different tower locations is corrected by using the coherence function γ(d,f)=exp(−0.1df) (d is the distance and f is the frequency); S3403: The conductor uses catenary elements, taking into account gravitational stiffness and geometric nonlinearity; The insulator string uses rigid rod units with hinged ends, which transmit force but not bending moment; Dynamic equation: ; Where [M] is the mass matrix, [C] is the damping matrix, [K] is the stiffness matrix, and [F(t)] is the seismic load.
[0034] S350: S351: Large-scale model → Small-scale model: Output the displacement time history of the boundary of the small-scale model as the input of the small-scale model; Small-scale model → structural model: Output the acceleration time history at the foundation of the structure as the input of the structural model; Data format: uniformly ASCII format, including time (s) and physical quantity (m / s² or m), sampling rate 100Hz; S352: Time synchronization: All models use the same time step (0.01 seconds) and start time alignment; Spatial mapping: Through coordinate transformation (WGS84 → local coordinate system), the grid points of the site model are accurately matched with the foundation positions of the structural model (error ≤ 1m). S353: A parallel computing platform built on Python+MPI, supporting batch input of model parameters and automatic concatenation of results; Deploy a visualization module to display real-time animations of seismic wave propagation (10 frames per second) and structural deformation cloud maps.
[0035] S360: S3601: Seismology Module: Input geological and source parameters, and output initial ground motion and structural response through the S301-305 model; Deep learning module: The residual between the initial output and the actual monitoring data is used as input, and a GRU neural network (64-dimensional input layer, 2 hidden layers × 32 neurons, 1-dimensional output layer) is used for correction. Fusion output: y final = y seismology + Δy deep learning, where Δy is the residual correction amount; S3602: Input features: magnitude, focal depth, Vs30, building height, structure type (10 categories); Output labels: Actual monitored peak ground acceleration (PGA), maximum inter-story drift angle (MIDR); Data volume: Contains 5000+ earthquake events and 10000+ building monitoring data from 1990 to 2023, divided into training and validation sets in a 7:3 ratio. The data comes from a high-quality dataset in the S201 storage layer. S3603: Normalize labels such as PGA and MIDR (range 0-1), the formula is y′=(y−ymin) / (ymax−ymin); The sample size was doubled by using methods such as random noise injection (signal-to-noise ratio ≥20dB) and parameter scaling (±10%).
[0036] S370: S3701: Target parameters: initial structural stiffness k0, site damping ratio ξ; Corrected data: Structural natural vibration period T (error ≤ 5%) and velocity response spectrum monitored during minor earthquakes (Mw < 5); Optimization algorithm: Particle Swarm Optimization (PSO); Objective function; With 50 particles, 100 iterations, and a convergence accuracy of 1e-4; S3702: Target parameters: yield strength Fy, post-yield stiffness ratio k1 / k0; Correction data: structural damage data of major earthquakes (Mw≥6) (e.g., MIDR=0.02-0.1), IDA curve library of S101; Matching method: Ensure the root mean square error between the simulated IDA curve and the measured curve is ≤10%, the formula is:
[0037] S3703: Validate the model using uncorrected seismic data (such as 100+ events from 2020-2023) to ensure that the PGA simulation error of the corrected model is ≤15% and the MIDR simulation error is ≤20%.
[0038] S380: S3801: Number of hidden layers: 2-4 layers; Number of neurons: 32-128; Learning rate: 0.0001-0.01; Batch size: 16-128; S3802: Use grid search to traverse all parameter combinations (total number of combinations ≤ 100). Using the root mean square error (RMSE) of the validation set as the evaluation metric, the parameter combination with the smallest RMSE is selected. S3803: Hidden layers: 3; Neuron: 64-32-16; Learning rate: 0.001; Batch size: 32; Number of iterations: 500 (early stopping mechanism: if the RMSE of the validation set does not decrease for 10 consecutive rounds, the iteration will stop).
[0039] The multi-source fusion model constructed by S300 needs to undergo scientific evaluation to verify its reliability, and the simulation results need to be transformed into quantitative risk indicators. This step identifies model defects through a multi-dimensional evaluation system and generates dynamic risk assessment results by combining probabilistic analysis. This provides direction for model optimization and specific improvement targets for the iterative upgrade of S500, which includes: S410: S4101: Time-stratified sampling: Data from 1990 to 2023 is divided into strata of 5 years each, and the training set and test set are extracted from each stratum at a ratio of 7:3 to avoid time distribution bias; Category balance: Ensure that the test set includes samples of different magnitudes (5-6, 6-7, and above 7) and different site types (hard, medium-hard, and soft), with each category accounting for ≥20% of the samples; S4102: Earthquake parameters: magnitude range 4.0-8.5 (average 6.2), focal depth 5-30km (average 15km); Site parameters: Vs30 range 150-800m / s (average 350m / s), covering Class I-V sites; Structural parameters: Building height 5-100m (average 25m), including frame, shear wall, and hybrid structures; S4103: The difference in feature distribution between the training set and the test set is ≤10% (passes the KS test). The test set contains ≥70% high-confidence data (w≥0.8) to ensure the reliability of the evaluation results. The data comes from the quality tag data of the S201 storage layer.
[0040] S420: S4201: Determine the sampling parameters: magnitude (5-8), focal depth (5-30km), Vs30 (150-800m / s), and building height (5-100m), for a total of 4 parameters; Number of strata: Each parameter is divided into 20 strata, and the total number of samples is 20 (to ensure that there is at least 1 sample in each stratum). Sampling process: Randomly select one value from each layer of each parameter and combine them into a new sample to ensure uniform coverage of the parameter space; S4202: Calculate the feature correlation between the expanded sample and the original sample (Pearson coefficient ≥ 0.8); Select 10% of the expanded sample and calculate its "theoretical response value" through a physical model to ensure that the response pattern is consistent with that of the original sample of the same type. S4203: Establish sample labels: label "original sample" and "expanded sample", and count them separately during evaluation; Regular updates: After adding actual earthquake data each year, the sample is resampled and expanded to maintain the timeliness of the sample database.
[0041] S430: S4301: Root Mean Square Error (RMSE) ; in, For PGA (gal) or MIDR, n is the sample size. RMSE should be ≤ 10% of the industry average (e.g., PGA RMSE ≤ 5gal). S4302: Correlation coefficient (R): ; S4303: Calculate the relative error between the simulated peak value and the measured peak value. The requirement is that PE ≤ 15%.
[0042] S440: S4401: Calculate the response spectrum values for different periods (0.1-10 seconds) for simulated and measured acceleration time histories: Sd (displacement response spectrum), Sv (velocity response spectrum), Sa (acceleration response spectrum); Damping ratio is taken as 5% (a standard value for building structures); S4402: Spectral ratio: The requirement is 0.85≤R≤1.15 (period 0.1-2 seconds, corresponding to the natural vibration period of most buildings). Spectral similarity: Calculate the Euclidean distance between two response spectra. ,Require .
[0043] S4403: Within a period of 0.1-2 seconds, 90% of the samples meet the spectral ratio requirements; The average value of spectral similarity D is The reaction spectra showed good consistency.
[0044] S450: S4501: Minor damage: MIDR < 0.01; Moderate injury: 0.01 ≤ MIDR < 0.05; Severe damage: 0.05 ≤ MIDR < 0.1; Collapse: MIDR ≥ 0.1; S4502: Accuracy: Acc = Number of correctly predicted samples / Total number of samples, with Acc ≥ 80% required; Unsafe probability bias:
[0045] Confusion matrix: Statistically calculate the false positive rate for minor → moderate, severe → collapse, etc., requiring the main false positive rate to be ≤10%; S4503: Overall accuracy rate is 85%, with prediction accuracy for minor damage and collapse exceeding 90%; The average unsafety probability deviation is 3.2%, which meets the engineering requirements; The main misjudgment was moderate to severe (8%), due to the high sensitivity of parameters in the elastoplastic stage.
[0046] S460: S4601: Level 1 (Low Risk): Building failure probability <10%; Level 2 (Low to Medium Risk): Failure probability ≤ 10% < 30%; Level 3 (Medium Risk): 30% ≤ probability of failure < 50%; Level 4 (Medium-High Risk): 50% ≤ Failure Probability < 70%; Level 5 (High Risk): Failure probability ≥ 70%; S4602: Matching degree calculation: Among them, the historical assessment level is based on the local seismic zoning map or the results of a special assessment; S4603: In the validation across 5 pilot areas (total area 5000 km²), the average matching rate was 88%; The high-risk area (level 5) had the highest matching rate (92%) because the geological and structural characteristics of the high-risk area are obvious and easy to identify.
[0047] S470: S4701: If the simulation error is large in a certain area, supplement the mobile station data for that area (add 5-10 stations) and re-acquire high-frequency Rayleigh wave data; Low-quality data (w < 0.6) are manually reviewed, corrected, or removed to improve the quality of training data; S4702: If the response spectrum consistency is poor, add characteristic dimensions (such as seismic wave duration and spectral characteristics). The Recursive Feature Elimination (RFE) algorithm is used to remove redundant features (importance < 5%) and reduce noise interference. S4703: If the simulation accuracy is low, adjust the finite element mesh density (refine it to 1 / 2 of the original size). If the damage prediction is inaccurate, correct the elastoplastic constitutive parameters (e.g., adjust the post-yield stiffness ratio from 0.1 to 0.15). If the deep learning module performs poorly, increase the number of hidden layers or change the activation function (e.g., change from ReLU to Swish).
[0048] S480: S4801: Building structure: MIDR ≥ 0.1 is considered a failure; Transmission towers: Members with stress ≥ yield strength are considered to have failed; S4802: Assume that the failure probability of the structure under n potential seismic events is p1, p2, ..., pn (independent events); Total failure probability , where pi is calculated by the S300 model (inputting the i-th earthquake parameter); S4803: A building may be affected by three earthquakes within the next 50 years: Mw6.0 (p1=5%), Mw6.5 (p2=15%), and Mw7.0 (p3=30%). The overall failure probability corresponds to risk level 3 (medium risk).
[0049] S490: S4901: Seismic intensity: Set at 6, 7, 8, 9 and 10 degrees according to the Chinese Seismic Intensity Scale (GB / T17742-2020); Corresponding magnitudes: In the target area (Vs30=300m / s), 6 degrees ≈ Mw5.0, 7 degrees ≈ Mw5.5, 8 degrees ≈ Mw6.0, 9 degrees ≈ Mw6.5, 10 degrees ≈ Mw7.0; S4901: For each intensity, run the S300 model and output the failure probability of all buildings in the area; Statistics on the "cumulative number of failed buildings" and "percentage of high-risk area" under different intensities; S4901: The horizontal axis represents the earthquake intensity (6-10 degrees), and the vertical axis represents the cumulative failure probability (0-100%). Curve fitting: The Logistic function P(I)=1 / (1+exp(−a(I−b))) is used, where I is the intensity and a and b are the fitting parameters (determined by the least squares method).
[0050] The evaluation results and risk assessment requirements of S400 revealed the dynamic adaptation problem between platform functionality and model performance. This step, through mechanisms such as data updates, functional iterations, and model retraining, enables the platform and model to continuously adapt to new data, new requirements, and changes in geological conditions, forming a complete closed loop of "data-platform-model-evaluation-optimization." S500 includes: S510: S5101: Real-time data: seismic waveforms, station status (updated every second); Regular data: Earthquake early warning results (updated daily), structural monitoring data (updated monthly); Batch data: geological data, historical earthquake damage data (updated every 5 years); The amount of data updated each time should be ≤10% of the total data volume to avoid system overload; S5102: Data Acquisition Terminal: New data is processed by S105-S107 (cleaning, standardization, quality marking) to generate an update package; Transmission end: Incremental transmission (transmitting only the changed parts) is used, and the encryption algorithm is AES-256 to ensure data security; Storage: The database adopts a "backup first, update later" strategy, with backups retained for 30 days and rollback supported; S5103: Automatic synchronization: Metadata (source, quality, and relationships) of newly added data is written to the metadata center in real time; Conflict handling: If there is a metadata conflict (such as different descriptions of the same event), the latest collected high-confidence data (w≥0.8) shall prevail; Synchronous verification: Daily comparison of data with metadata (such as data volume and time range), triggering an alarm when inconsistencies occur.
[0051] S520: S5201: Built-in feedback module: Users can submit feature suggestions (text + screenshots) and bug reports (including operation logs); Regular surveys: Questionnaires (sample size ≥ 100) are distributed to researchers and emergency response departments every quarter to collect needs; Expert review: Each year, 5-8 earthquake experts will be organized to evaluate the scientific validity and practicality of the platform's functions; S5202: Functional defect (e.g., path planning error): Level P0, response within 24 hours; User experience optimization (e.g., interface lag): Level P1, response within 7 days; New feature suggestion (e.g., adding a tsunami warning): Level P2, feasibility assessed within 30 days; S5203: Defect Fix: After the technical team locates the problem, they release a patch version (minor version number + 1), such as V1.0 → V1.1; User experience optimization: Adjust the interface layout and optimize algorithm efficiency (e.g., reduce query time from 3 seconds to 1 second); Requirements assessment: For new feature suggestions, calculate the ratio of development cost to user benefit; if the ratio is ≥1.5, include it in the development plan. S530: S5301: Write a requirements specification document: clearly define the functional objectives (such as "post-earthquake recovery assessment of transmission lines"), inputs and outputs, and performance indicators (response time < 30 seconds); Architecture design: A microservice architecture is adopted, with new modules deployed independently and communicating with the existing platform via APIs to avoid affecting core functions; S5302: Technology stack: Backend uses Java / SpringBoot, frontend uses Vue.js, and the database is compatible with existing storage (MySQL / MongoDB). Test cases include functional tests (input exceptions, boundary conditions), performance tests (concurrent users ≥ 50), and compatibility tests (adaptation to Chrome / Edge browsers). S5303: Gray-scale release: First, open the new features to 10% of users and collect user feedback; Full rollout: After confirming there are no major issues, switch traffic to the new module using a load balancer; Documentation updates: User manuals and API documentation are updated in sync, and online training is organized (for research and emergency users).
[0052] S540: S5401: Cycle: Full retraining every 2 years, incremental retraining every 6 months (using only new data); Data: Full retraining uses all the data after governance up to the present (≥1 million records), and incremental retraining uses data from the past 6 months (≥50,000 records). The data comes from the storage layer updated by S501. S5402: Data preparation: Call the governance rule base of S108 to standardize the newly added data; Model training: Using the S306 hybrid framework, keeping the hyperparameters unchanged (to ensure comparability), retraining was performed; Performance evaluation: Calculate evaluation metrics (RMSE, accuracy, etc.) using the latest test set (data from the last 6 months); S5403: Baseline setting: Use the performance metrics from the first training session as the baseline (e.g., RMSE=5gal). Bias warning: When the deviation of the metrics after retraining is greater than 10% (e.g., RMSE=5.6gal), deep optimization (e.g., adjusting the model architecture) is triggered. Baseline update: If the performance improvement after retraining is ≥15% (e.g., RMSE=4.2gal), update the baseline to the new metric; S550: S5501: Mega-earthquake: Mw≥8.5 (such as an enhanced version of the Wenchuan earthquake, Mw8.0); Special earthquake sources: sudden fault slip (slippage > 5m), twin earthquakes (two earthquakes of magnitude Mw ≥ 7 with an interval of < 24 hours); Complex sites: near fault zones (distance from fault < 5km), soft soil areas (Vs30 < 150m / s); S5502: Generator: Input random noise + extreme working condition parameters (such as Mw8.5, fault slip 6m), output simulated ground motion time history; Discriminator: Distinguishes between generated data and real data (historical major earthquake data); the loss function used is cross-entropy. Training iteration: The generator and discriminator are trained adversarially for 500 rounds until the discriminator accuracy is approximately 50% (unable to distinguish between real and fake data); S5503: Mix the generated data (10% of the training set) with the real data and retrain the S300 model; The focus is on optimizing parameters for major earthquakes (such as yield strength and ultimate deformation) to ensure that the simulation error under extreme conditions is ≤20%. S560: S5601: Dense monitoring points (spacing ≤ 5km) are set up in fault zones and active fracture areas. The sensor is a three-component accelerometer (range ±2g). Data sampling rate of 100Hz, transmitted to the platform processing layer in real time; S5602: Input real-time monitoring seismic wave travel time data (P-wave and S-wave arrival times). A joint inversion algorithm (tomography + travel time inversion) is used to update the subsurface velocity structure (Vs, Vp). Update frequency: Regular quarterly updates, with emergency updates within 72 hours of an earthquake of magnitude Mw≥5; S5603: Internal validation: Calculate the travel time residuals of the updated model (≤0.1 seconds). External validation: The simulation epicenter error was ≤5km, verified by using newly occurring earthquake events (not involved in the inversion).
[0053] Through the organic integration of the above five steps, this technical solution realizes closed-loop management of the entire process of earthquake big data, from acquisition and governance to platform construction, model training, evaluation and optimization. The high-quality data of S100 lays the foundation for subsequent steps, the platform architecture of S200 provides efficient support, the model of S300 realizes the core simulation function, the evaluation and risk assessment of S400 verifies the value of the model and points out the direction for optimization, and the iterative mechanism of S500 ensures long-term adaptability, ultimately providing comprehensive technical support for earthquake scientific research, emergency rescue and disaster prevention and mitigation.
[0054] Experimental example: Experimental objective: Verify the effectiveness of the entire technical solution of "data acquisition-governance-platform construction-model training-evaluation and optimization", specifically including: the data quality compliance rate after data governance, the operational performance of the platform's three-layer architecture, the simulation accuracy of the multi-source fusion seismic motion model, the quantitative accuracy of dynamic risk assessment, and the long-term adaptability of the iterative optimization mechanism.
[0055] Experimental area: A pilot city was selected within the Longmenshan Fault Zone in Sichuan Province (geographical range: 103°-104°E, 31°-32°N, area approximately 5000 km²). This region has a history of frequent seismic activity (such as the 2008 Wenchuan earthquake and the 2013 Lushan earthquake), possesses abundant seismic monitoring stations and historical earthquake damage data, and includes densely populated areas (city center) and power transmission line corridors (suburbs), meeting the requirements for multi-scenario verification.
[0056] Experimental environment: Data acquisition terminal: Hardware configuration: 5 fixed stations (deploying PCB393B05 three-component accelerometers, sampling rate 1000Hz), 3 mobile stations (deploying Nanometrics Trillium Compact sensors, sampling rate 2000Hz), BeiDou time synchronization module (synchronization accuracy ≤1ms); Software configuration: data acquisition software (SeisComP3), link monitoring tool (Zabbix). Platform servers: storage servers (2 units, CPU Intel Xeon Gold 6348, memory 128GB, hard disk 40TB SSD), compute servers (4 units, GPU NVIDIA A100, memory 256GB); software configuration: operating system (CentOS 8.5), database (MySQL 8.0 cluster, MongoDB 5.0, InfluxDB 2.0), computing framework (Spark 3.3, TensorFlow 2.10); Application terminals: Research staff workstations (CPU i7-13700K, 64GB memory), emergency department terminals (industrial tablets, supporting 4G / 5G); Software configuration: visualization software (ParaView 5.11), GIS tools (ArcGIS 10.8), and custom platform client (developed with Vue.js).
[0057] Detailed experimental steps: Step 1: Real-time seismic waveforms (including P-wave and S-wave time-domain signals) were collected by 5 fixed stations, and high-frequency Rayleigh wave data were collected by 3 mobile stations. The data was transmitted to the data center via a dual-link system of "fiber optic + 5G" and timed by BeiDou. Historical earthquake data for the region from 2000 to 2023 (a total of 526 earthquakes, of which 28 were of magnitude ≥ 5) were obtained from the China Earthquake Networks Center.
[0058] Geological profiles of the pilot city were obtained from the Sichuan Provincial Geological Survey (including sand layer thickness of 3-8m, clay layer thickness of 5-12m, bedrock layer Vs=2800-3200m / s), and Longmenshan fault parameters (strike 310°, dip angle 55°, burial depth 8-15km); the Vs30 distribution (180-450m / s) of the target area was extracted from the "China Seismic Ground Motion Parameter Zoning Map". BIM models of 200 building complexes in the pilot city (including 150 multi-story buildings with less than 10 floors and 50 high-rise buildings with more than 10 floors) and parameters of 50 220kV transmission towers (height 35-55m, main material cross-section Q345 steel) were collected; earthquake damage records from 2008 to 2023 were obtained from the local emergency management bureau (e.g., the building collapse rate in this area was 12% during the 2008 Wenchuan earthquake). Based on the Lasso algorithm analysis of data acquisition task records from 2021 to 2023, a resource requirement model was constructed, predicting that the "real-time waveform acquisition" task requires a CPU clock speed of ≥3.0GHz and memory of ≥16GB. Four suitable computing nodes were selected, and the hardware resource utilization rate was improved from 48% in the traditional method to 82%. Twelve duplicate waveform data were removed, and eight short-term missing data (2-4 seconds in duration) were filled in using the "interpolation of the mean of adjacent stations". Six conflicting data were identified (such as Mw6.2 recorded by a fixed station and Mw6.8 recorded by a mobile station for a certain earthquake event), and the conflict intensity was calculated (0.32-0.65). Four medium- and low-intensity conflicts (<0.5) were corrected through geological fault data verification, and the final values of two high-intensity conflicts (≥0.5) were confirmed by consultation among three earthquake experts. After the data was processed, the data format was 100% uniform, the data with a credibility weight of ≥0.8 accounted for 92.3%, and the conflict resolution rate was 98.5%, meeting the input requirements of the subsequent platform.
[0059] Step 2: The MySQL cluster stores structured data (such as event parameters of 526 earthquakes, with 86,000 rows in a single table), with a query response time of ≤0.5 seconds; MongoDB stores unstructured data (such as earthquake waveform files, with a single file size of 100-500MB), supporting queries by "event ID + quality level" combination; InfluxDB stores real-time monitoring data (100 waveform data entries written per second, with a write throughput of ≥12,000 entries / second); the Apriori association rule algorithm is deployed to discover the association rule that "when Vs30 < 250m / s and Mw ≥ 6.0, the building damage rate increases by 45%" (minimum support 5%, confidence 72%); an LSTM earthquake hazard prediction model is constructed (128-dimensional input layer, 3 hidden layers × 64 neurons), trained with data from 2000-2020 and validated with data from 2021-2023, achieving an accuracy of 83% in predicting the probability of Mw ≥ 5 earthquakes in the region within the next 5 years; It provides researchers with an interface for adjusting model parameters (supporting modification of the LSTM learning rate from 0.001 to 0.01); it offers emergency departments a "disaster hotspot identification" function, allowing them to input parameters of a Mw4.2 earthquake in March 2024 and output a heat map of high-risk areas within 10 seconds (with a deviation of ≤8% from actual inspection results); and it provides the public with a 3D visualization module for the Wenchuan earthquake (supporting view zooming and vibration process playback). Functional testing: The entire data acquisition, storage, and processing chain was uninterrupted, with a conflict data resolution rate of 96.2% (meeting threshold ≥95%) and a rescue route planning accuracy rate of 92.3% (meeting threshold ≥90%).
[0060] Performance testing: 100 researchers simultaneously called the data interface with a 100% concurrent response rate and an average latency of 1.8 seconds; the batch processing interval for real-time waveform data processing was 10 seconds, consistent with the preset threshold.
[0061] Step 3: Large-scale macroscopic model: The finite difference method is used, with a grid size of 1km×1km (vertical depth of 30km). Geological structure and fault parameters are input to simulate the propagation of regional seismic waves and output the ground acceleration time history (correlation coefficient R=0.81 with the measured data of the fixed station). Small-scale refined model: For a 10km×10km area in the city center, the grid was refined to 50m×50m, and the shallow Vs was corrected by combining high-frequency Rayleigh wave phase velocity data (the deviation between the corrected Vs and the borehole measured Vs is ≤10%), and the local simulation accuracy is improved by 22% compared with the macro model. Building complex: For buildings with fewer than 10 stories, the MDOF shear model (natural period T = 0.3-1.5s) is used; for buildings with more than 10 stories, the bending-shear coupling model (T = 1.6-3.5s) is used; and for the elastoplastic stage, the bilinear constitutive model (post-yield stiffness ratio 0.1) is used. Transmission towers: Construct a spatial truss model, consider the traveling wave effect (wave speed 400m / s), extract an equivalent single tower mass point model (3 mass points), and calculate the stress error of the members to be ≤15%; Using the "seismological theory + GRU" framework, the data after treatment (526 earthquakes + parameters of 200 buildings) were input, and the parameters were corrected in stages: in the elastic stage, the initial stiffness was corrected using the data of the Mw4.2 minor earthquake in March 2024 (natural vibration period simulation error ≤5%), and in the elastoplastic stage, the skeleton curve was corrected using the damage data of the Wenchuan earthquake in 2008 (MIDR simulation error ≤18%). The results were validated using data from 10 earthquakes (Mw4.0-6.1 magnitude) that were not used in the training between 2021 and 2023, and are as follows: Step 4: The 526 earthquake data were divided into a training set (368 earthquakes) and a test set (158 earthquakes) in a 7:3 ratio. The test set covers Mw4.0-8.0 (including 2 historical earthquakes with Mw≥7), and high-confidence data (w≥0.8) accounts for 75%. Latin hypercube sampling was used to stratify and sample three parameters: magnitude (5-8), Vs30 (180-450 m / s), and building height (5-100 m), generating 100 expanded samples with a feature correlation coefficient ≥0.83 with the original samples. Simulation accuracy: RMSE=4.1gal, R=0.87 for the PGA test set, and RMSE=0.014, R=0.83 for the MIDR test set; Damage prediction: The accuracy rate of building damage level is 86%, and the deviation of unsafe probability is 3.1% (the standard threshold is ≤5%). Risk matching: The risk level output by the model matches the "Pilot City Earthquake Risk Zoning Report" with a degree of 89% (meeting the threshold of ≥85%). Based on the Poisson binomial distribution, the probability of failure of the central building complex in the pilot city over the next 50 years is calculated as follows: encountering a Mw6.0 earthquake (p1=8%), a Mw6.5 earthquake (p2=18%), and a Mw7.0 earthquake (p3=35%), the total probability of failure is P=1-(1-0.08)(1-0.18)(1-0.35)=52.7%, corresponding to a risk level of 4 (medium-high risk); The "intensity-cumulative failure probability" risk curve was generated, and the results showed that the failure probability was 2% at intensity 6, 8% at intensity 7, 25% at intensity 8, 58% at intensity 9, and 82% at intensity 10, which deviated from the historical earthquake damage statistics (such as the actual failure probability of 23% at intensity 8) by ≤8%. Step 5: New earthquake monitoring data from April to July 2024 (12 minor earthquakes) has been added; the credibility weight of the metadata center has been updated (91% of the new data has a w≥0.8); based on feedback from emergency departments that "the identification of disaster hotspots is delayed by 5 minutes," the calculation frequency has been shortened from 5 minutes to 2 minutes, improving response time by 60%. Full retraining: The model was retrained using full data from 2000 to 2024 (538 earthquakes), reducing the PGA simulation error from 12% to 10%. Extreme case supplement: Based on GAN, 20 sets of Mw8.5 magnitude super earthquake data (fault slip 6m) were generated. After supplementary training, the PGA simulation error of the model for extreme earthquakes was reduced from 25% to 18%. Velocity model update: Using real-time travel time data from a Mw4.5 earthquake in July 2024, the subsurface Vs structure was updated through joint inversion, reducing the epicenter location error of subsequent earthquakes from 8km to 4km. The overall performance of the model improved by 15%-20% after iteration, and the platform's functional coverage increased from 80% to 92%, meeting the needs of long-term earthquake monitoring and risk assessment.
[0062] Experimental conclusion: Data governance phase: Multi-link acquisition (fiber optic + 5G) enables data transmission reliability to reach 99.9%, and the Lasso resource scheduling algorithm improves hardware utilization to 82%. Hierarchical conflict handling and quality labeling ensure that the proportion of data with a credibility weight ≥0.8 after governance reaches 92.3%, providing a high-quality data source for subsequent stages. Platform architecture phase: A three-tier architecture enables efficient management of heterogeneous data: structured data query latency is ≤0.5 seconds, and unstructured data storage costs are reduced by 30%; the processing layer algorithm (such as LSTM hazard prediction) achieves an accuracy rate of 83%, and the application layer's multi-role functions meet the differentiated needs of scientific research, emergency response, and the public, with a rescue route planning accuracy rate of 92.3%. Model building phase: The multi-scale site-structure coupled model, combined with a "seismology + deep learning" hybrid framework, achieves PGA simulation error ≤12%, MIDR simulation error ≤18%, and response spectrum consistency meets the engineering requirements of 0.85-1.15, improving accuracy by 20%-25% compared to traditional single models. Evaluation and optimization phase: The multi-dimensional evaluation indicators (RMSE, R, damage accuracy) all meet the standards, and the deviation of the "intensity-failure probability" curve of dynamic risk assessment from historical data is ≤8%; the iterative optimization mechanism (data update, model retraining, extreme working condition supplementation) reduces the long-term accuracy decay rate of the model to 5% / year, and the platform's functional adaptability continues to improve. Overall conclusion: This technical solution, through a closed-loop design of the entire process, solves the pain points of traditional methods such as "low data quality, poor platform performance, insufficient model accuracy, and difficulty in risk quantification." It can effectively support earthquake scientific research and analysis, emergency rescue decision-making, and disaster prevention and mitigation work, and has engineering application value.
[0063] Example 2: A system for constructing an earthquake big data platform based on data collection and governance, used to realize an earthquake motion model construction and evaluation system based on big data, including the following modules; The module includes data acquisition and governance, platform architecture, model building, evaluation and assessment, and iterative optimization. Data acquisition and governance module: Collects earthquake monitoring, geological and seismic source, structural and disaster data, transmits them to the data center through wired and wireless dual links, schedules hardware resources based on Lasso algorithm, uses Raft consensus algorithm to ensure data consistency, cleans data, handles conflicts, unifies format and coordinate system, adds credibility weight labels and dynamically updates governance rule base; Platform architecture modules: A three-tier architecture is constructed. The storage layer uses relational and non-relational databases to classify and store structured and unstructured data, and integrates a data warehouse with a metadata center. The processing layer deploys association rules, clustering algorithms, and LSTM / GRU seismic hazard prediction models, and allocates tasks in conjunction with a resource scheduling model. The application layer provides differentiated functions for researchers, emergency rescue departments, and the public. Model building module: Combines finite difference and finite element methods to build multi-scale site models, constructs structure-site coupled sub-models for building complexes and power transmission towers, and trains the model using a hybrid framework of seismic theory and deep learning and corrects parameters in stages; Evaluation and assessment module: Divide the training set and test set, expand the sample by Latin hypercube sampling, evaluate the model by simulation accuracy and response spectrum consistency index, calculate the structural failure probability based on Poisson binomial distribution, simulate multi-intensity earthquake scenarios and optimize substandard models; Iterative optimization module: Regularly updates data and metadata, adjusts platform functions based on user feedback, retrains the model with full data every 2 years, supplements extreme working condition data based on generative adversarial networks, and dynamically updates the earthquake velocity model.
[0064] Example 3: A computer device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a method for constructing and evaluating earthquake motion models based on big data.
[0065] Example 4: A computer-readable storage medium, characterized in that the storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a method for constructing and evaluating earthquake motion models based on big data.
[0066] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for constructing and evaluating seismic motion models based on big data, characterized in that, include: S100: Intelligent acquisition and deep governance of multi-source heterogeneous seismic data. It collects seismic monitoring data, geological and source data, and structural and disaster data, and transmits them to the data center through wired and wireless dual links. It schedules hardware resources based on the Lasso algorithm and uses the Raft consensus algorithm to ensure data consistency. It cleans the collected data, handles conflicts, unifies the format and coordinate system, adds credibility weight labels, and builds and dynamically updates the governance rule base. S200: Construct a three-layer architecture for the earthquake big data platform. The storage layer uses relational and non-relational databases to classify and store structured and unstructured data, integrates a data warehouse and associates it with a metadata center. The processing layer deploys association rule algorithms, clustering algorithms and an LSTM / GRU-based earthquake hazard prediction model, and allocates processing tasks in conjunction with a resource scheduling model. The application layer provides differentiated functions for researchers, emergency rescue departments, and the public; S300: Construct a multi-source fusion seismic motion model, combine finite difference and finite element methods to construct a multi-scale site model, construct a structure-site coupled sub-model for building complexes and power transmission towers, use a hybrid framework of seismological theory and deep learning to train the model, and correct the model parameters in stages. S400: Multi-dimensional model evaluation and dynamic risk assessment, dividing the training set and test set, using Latin hypercube sampling to expand the sample, evaluating the model through indicators such as simulation accuracy and response spectrum consistency, calculating the structural failure probability based on Poisson binomial distribution, simulating multi-intensity earthquake scenarios and optimizing substandard models. S500: The platform and model are optimized through feedback iteration, data and metadata are updated regularly, platform functions are adjusted based on user feedback, the model is retrained with full data every 2 years, extreme working condition data is supplemented based on generative adversarial networks, and the seismic velocity model is dynamically updated.
2. The method for constructing and evaluating seismic motion models based on big data according to claim 1, characterized in that: In S100, the conflict intensity is calculated during data conflict handling. vᵢ represents the parameter value of the i-th data source. Here, wᵢ represents the mean of the parameters, and wᵢ represents the data source credibility weight. Automatic weighted correction is applied when the conflict intensity is <0.3; verification is performed in conjunction with geological data when the conflict intensity is 0.3 ≤ conflict intensity <0.5; and manual review is conducted when the conflict intensity is ≥0.
5.
3. The method for constructing and evaluating earthquake motion models based on big data according to claim 1, characterized in that: In S200, the relational database in the storage layer is MySQL, which stores structured data such as earthquake event parameters and station information; the non-relational database is MongoDB, which stores unstructured data such as earthquake waveform files and BIM models; the data warehouse supports multi-dimensional queries based on "credibility weight + time range".
4. The method for constructing and evaluating earthquake motion models based on big data according to claim 1, characterized in that: In S300, the phased correction of model parameters is as follows: in the elastic stage, the natural vibration period and elastic displacement of the structure monitored by minor earthquakes are used to correct the initial stiffness through particle swarm optimization algorithm; in the elastoplastic stage, the structural skeleton curve is corrected using historical damage data from major earthquakes and an incremental dynamic analysis curve library.
5. The method for constructing and evaluating seismic motion models based on big data according to claim 1, characterized in that: In S400, simulation accuracy indicators include root mean square error (RMSE) and correlation coefficient (R), requiring R ≥ 0.8 and RMSE ≤ 10% of the industry average; response spectrum consistency indicators include maximum displacement response (Sd) and maximum velocity response (Sv), requiring the deviation between simulated and actual values to be ≤ 15%.
6. The method for constructing and evaluating seismic motion models based on big data according to claim 1, characterized in that: In S500, the frequency of data updates is as follows: real-time seismic waveform data is updated every second, structural monitoring data is updated monthly, and geological and historical earthquake damage data is updated every 5 years; when dynamically updating the seismic velocity model, it is updated quarterly as usual, and an emergency update is performed within 72 hours after an earthquake of magnitude Mw≥5.
7. The method for constructing and evaluating seismic motion models based on big data according to claim 1, characterized in that: In S100, the credibility weight is calculated by weighting the station accuracy (0.8-1.0), data integrity (0.6-1.0), and conflict handling results (0.7-1.0), with the formula w=0.4w_+0.3w_integrity+0.3w_conflict. Data with a weight <0.6 is marked as low quality and re-collected.
8. A seismic motion model construction and evaluation system based on big data, used to implement the method described in any one of claims 1-7, characterized in that, Includes the following modules: The module includes data acquisition and governance, platform architecture, model building, evaluation and assessment, and iterative optimization. Data acquisition and governance module: Collects earthquake monitoring, geological and seismic source, structural and disaster data, transmits them to the data center through wired and wireless dual links, schedules hardware resources based on Lasso algorithm, uses Raft consensus algorithm to ensure data consistency, cleans data, handles conflicts, unifies format and coordinate system, adds credibility weight labels and dynamically updates governance rule base; Platform architecture modules: A three-tier architecture is constructed. The storage layer uses relational and non-relational databases to classify and store structured and unstructured data, and integrates a data warehouse with a metadata center. The processing layer deploys association rules, clustering algorithms, and LSTM / GRU seismic hazard prediction models, and allocates tasks in conjunction with a resource scheduling model. The application layer provides differentiated functions for researchers, emergency rescue departments, and the public; Model building module: Combines finite difference and finite element methods to build multi-scale site models, constructs structure-site coupled sub-models for building complexes and power transmission towers, and trains the model using a hybrid framework of seismic theory and deep learning and corrects parameters in stages; Evaluation and assessment module: Divide the training set and test set, expand the sample by Latin hypercube sampling, evaluate the model by simulation accuracy and response spectrum consistency index, calculate the structural failure probability based on Poisson binomial distribution, simulate multi-intensity earthquake scenarios and optimize substandard models; Iterative optimization module: Regularly updates data and metadata, adjusts platform functions based on user feedback, retrains the model with full data every 2 years, supplements extreme working condition data based on generative adversarial networks, and dynamically updates the earthquake velocity model.
9. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the earthquake motion model construction and evaluation method based on big data as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a method for constructing and evaluating earthquake motion models based on big data as described in any one of claims 1 to 7.