Data resource pool building system

By real-time monitoring and dynamic adjustment of the seismic data resource pool, the problems of uncontrolled data quality and low efficiency in system fault location in existing technologies have been solved. This has enabled efficient and accurate data processing and system stability, making it suitable for seismic monitoring scenarios with high-frequency observations and high-concurrency writing.

CN121995442AInactive Publication Date: 2026-05-08EARTHQUAKE ADMINISTRATION OF BEIJING MUNICIPALITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EARTHQUAKE ADMINISTRATION OF BEIJING MUNICIPALITY
Filing Date
2025-12-05
Publication Date
2026-05-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack the ability to monitor earthquake data resource pools in real time, and cannot effectively detect physical layer failures such as hardware damage, storage logic errors, or architectural imbalances, resulting in data quality fluctuations and system response delays. Furthermore, the fixed storage duration is insufficient to cope with sudden loads that could cause data accumulation or loss.

Method used

By combining a data acquisition module, a non-seismic event exclusion module, a physical performance detection module, and an anomaly cause tracing module, along with waveform structure feature comparison, spatial residual calculation, and environmental disturbance analysis, the system enables real-time monitoring and dynamic adjustment of seismic waveform data, identifies and eliminates abnormal data, and optimizes storage architecture performance.

Benefits of technology

It improves the accuracy and system stability of seismic data processing, ensures data quality and system efficiency in high-frequency observation and high-concurrency writing scenarios, and provides high adaptability and scalability support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121995442A_ABST
    Figure CN121995442A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of geographic information system and seismological observation application, in particular to a data resource pool building system which comprises a data acquisition module, a non-seismic event exclusion module, a physical performance detection module and an abnormal reason tracing module. According to the method, seismic waveform original data are periodically collected, storage duration comparison is carried out, seismic observation data categories are judged based on multi-dimensional parameters, physical performance fault categories are identified through hardware, software and architecture state parameters, and the method specifically comprises the steps of extracting and analyzing data geographic positions to judge scheduling paths, identifying and clearing repeated observation contents, and determining the physical performance fault categories. Calculating waveform stacking pressure and signal writing skew, judging frequency domain sorting delay, calculating a time adjustment factor based on frequency drift and correcting a storage time threshold value; through multi-dimensional data verification and load optimization, dynamic adjustment of a data storage time threshold is realized, and data storage efficiency and system stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of geographic information systems and earthquake monitoring application technology, and in particular to a data resource pool construction system. Background Technology

[0002] With the rapid expansion of earthquake observation networks, the massive amounts of seismic waveform data generated have exceeded the capacity of traditional storage architectures and processing models. Currently, the management of earthquake data both domestically and internationally largely relies on static data resource warehouses or offline archiving methods, lacking a comprehensive solution for real-time monitoring and dynamic adjustment of data acquisition cycles, storage durations, data quality, and system operating status. For example, existing technologies often use standard storage durations for unified control, which cannot cope with sudden peak observation loads or abnormal data accumulation introduced by interference factors, leading to data write blockages, redundant storage, or failures in misclassification and cleaning operations, thereby causing a decline in the operating efficiency of the terminal system.

[0003] Furthermore, traditional seismic data processing systems primarily focus on data archiving and format standardization, as well as lifecycle management and metadata cataloging. However, most solutions lack automatic detection and fault location mechanisms for physical and behavioral layer performance degradation, such as system hardware degradation, logical anomalies, and resource distribution imbalances. In addition, existing massive seismic data resource pools typically lack the ability to analyze and dynamically optimize real-time performance indicators such as frequency domain sorting latency and write channel load distribution.

[0004] Therefore, existing technologies often suffer from problems such as system response delays, data quality fluctuations, or uneven resource allocation in scenarios involving high-concurrency seismic observation, high-frequency storage requests, and environmental disturbances. There is an urgent need for a data resource pool construction system capable of collaboratively operating across multiple dimensions, including acquisition, interference elimination, performance testing, anomaly location, and dynamic threshold adjustment, to improve the real-time performance, accuracy, and system stability of seismic data processing.

[0005] Chinese Patent Publication No. CN106777271A discloses an automatic system construction method based on a service resource pool, including: S1: establishing an automatic system container framework; S2: constructing three major resource pools: a data resource pool, a plugin resource pool, and an interface resource pool; S3: setting the title, logo, base map, scope, scale, navigation bar, and copyright information of the WebGIS system based on the data resource pool; S4: setting various functional components of the WebGIS system based on the plugin resource pool; S5: setting the interface skin style of the WebGIS system based on the interface resource pool; S6: previewing the automatically constructed WebGIS system based on the above resource pools; and S7: saving the automatically constructed WebGIS system.

[0006] Meanwhile, existing technologies only focus on static resource assembly and lack the ability to monitor the real-time status of the data pool. They cannot detect physical layer failures such as hardware damage, storage logic errors or architectural imbalances, nor can they identify anomalies such as write blocking or duplicate accumulation at the data behavior layer. At the same time, direct storage of raw data causes environmental noise to pollute the data pool, and the fixed standard storage duration can lead to the risk of accumulation or loss when there is a sudden surge in data, resulting in low efficiency in locating operational and maintenance faults. Summary of the Invention

[0007] To address this, the present invention provides a data resource pool construction system to overcome the problems in the prior art where the lack of a real-time operation monitoring mechanism for data pool storage leads to the pollution of raw data causing quality control failure, and the inability of fixed parameters to cope with sudden loads, resulting in data accumulation or loss and low efficiency in fault location.

[0008] To achieve the above objectives, the present invention provides a data resource pool construction system, comprising: The data acquisition module is used to periodically acquire raw seismic waveform data and obtain the comparison results between the actual storage time and the standard storage time of the current acquisition cycle. The non-seismic event exclusion module is connected to the data acquisition module and is used to determine the category of the original seismic waveform data based on the interference event exclusion parameters after the first storage duration comparison result. The physical performance detection module, which is connected to the non-seismic event exclusion module, is used to perform a hierarchical detection operation based on hardware operating status parameters, software operating status parameters and architecture operating status parameters when the original seismic waveform data is the first observation data category, so as to determine whether there is a fault in the data pool, and to determine the corresponding physical performance fault category when there is a fault. Among them, physical performance failures include hardware performance failures, storage logic performance failures, or architecture performance failures. The anomaly cause tracing module, which is connected to the non-seismic event exclusion module, is used to remove interfering events in the second observation data category and perform anomaly cause judgment operation to determine whether there is a fault in the data pool, and to determine the corresponding data behavior performance fault when there is a fault. Specifically, when a performance failure result with no data behavior is obtained, the standard storage time is adjusted based on the time adjustment factor.

[0009] Furthermore, the non-seismic event exclusion module includes a parameter discrimination unit and a category determination unit; The parameter discrimination unit is used to perform multi-dimensional feature analysis on the original seismic waveform data based on interference event exclusion parameters in order to identify suspected abnormal data. The category determination unit is used to classify suspected abnormal data according to the data category to which the original seismic waveform data belongs, and output classification results, including a first classification result and a second classification result. The data category corresponding to the first classification result is the first observation data category, and the data category corresponding to the second classification result is the second observation data category.

[0010] Furthermore, the parameter discrimination unit includes a waveform feature analysis subunit, a spatial residual calculation subunit, an environmental interference discrimination subunit, and a scene enhancement recognition subunit; The waveform feature analysis subunit is used to extract several waveform structure feature parameters from the original seismic waveform data, and to determine waveform structure anomaly data based on the comparison results of each waveform structure feature parameter with the corresponding judgment threshold or interval. The spatial residual calculation subunit is used to calculate the travel time residual value of the source inversion fitting based on the seismic phase arrival time information, and to determine the waveform structure anomaly data based on the comparison result of the source positioning deviation and the preset tolerance threshold. The environmental interference discrimination subunit is used to identify abnormal waveform structure data caused by environmental factors based on the comparison results of environmental interference characteristic parameters and corresponding judgment thresholds. The environmental disturbance characteristic parameters include meteorological disturbance index, electromagnetic field strength and traffic harmonic characteristics; The scene enhancement recognition subunit is used to re-perform parameter comparison on the data processed by the waveform feature analysis subunit, the spatial residual calculation subunit, and the environmental interference discrimination subunit.

[0011] Furthermore, the physical performance detection module includes a fault detection unit and an architecture defect detection unit; The fault detection unit is used to detect the hardware and software operating status parameters in the storage system to determine whether there is a hardware performance fault or a storage logic performance fault in the data pool. The architecture defect detection unit is used to detect the architecture operating status parameters in the storage pool in order to determine whether there are any data pool architecture performance failures.

[0012] Furthermore, the fault detection unit includes a hardware fault detection subunit and a software fault detection subunit; The hardware fault detection subunit is used to detect the hardware operating status parameters in the storage system in order to determine whether there is a data pool hardware performance fault. The software fault detection subunit is used to detect the software running status parameters in the storage system in order to determine whether there is a data pool storage logic performance fault.

[0013] Furthermore, the anomaly cause tracing module includes a distribution rationality verification unit and a write load distribution identification unit; The distribution rationality verification unit is used to extract and analyze the geographical location information of the original seismic waveform data corresponding to the second observation data category, and based on the waveform content, timestamp and spatial coverage within the same acquisition cycle, identify duplicate data, extract the area of ​​the overlapping region and perform difference processing. The write load distribution identification unit is used to identify and schedule the write load status of the data pool.

[0014] Furthermore, the distribution rationality verification unit includes a path scheduling identification subunit; The path scheduling identification subunit is used to analyze the geographical location information of the observation data.

[0015] Furthermore, the distribution rationality verification unit also includes a duplicate observation content cleanup subunit; The duplicate observation content cleaning subunit is used to compare waveform data, timestamps and spatial coverage under the same acquisition period. If the data are completely consistent and the geographical coverage overlaps, the difference processing is performed to remove the redundant parts.

[0016] Furthermore, the write load distribution identification unit includes a waveform stacking pressure determination subunit; The waveform stacking pressure determination subunit is used to obtain the ratio of waveform writing amount to cleaning residue per unit time, obtain the actual waveform stacking ratio, identify whether there is writing stacking pressure based on the comparison result of the actual waveform stacking ratio and the waveform stacking ratio threshold, and call the backup pool when it is determined that there is writing stacking pressure.

[0017] Furthermore, the write load distribution identification unit includes a threshold calibration subunit; The threshold calibration subunit is used to adjust the storage time threshold to the calibration storage time based on the time adjustment factor when the first data write distribution judgment result is obtained.

[0018] Compared with existing technologies, the advantages of this invention are as follows: by continuously monitoring waveform data within a collection period, and combining non-seismic event identification, physical performance detection, and anomaly tracing mechanisms, it distinguishes performance problems at the hardware, software, architecture, and data behavior levels; by introducing waveform structure feature comparison, spatial residual calculation, and environmental disturbance analysis, it achieves accurate identification and filtering of non-seismic events, ensuring the accuracy of data classification; in terms of physical fault identification, the system effectively locates structural or component-level defects in the storage pool through a hierarchical parameter detection model; simultaneously, by combining frequency domain energy drift analysis and write distribution status judgment, it performs dynamic time threshold correction and scheduling optimization operations to achieve real-time adjustment of the data pool load status; the system has high adaptability and scalability, improves the processing efficiency and reliability of raw data, and is suitable for earthquake monitoring and earthquake emergency application scenarios with high-frequency observation and high-concurrency writing.

[0019] Furthermore, by setting multidimensional feature parameters and judgment thresholds, and combining waveform structure analysis, seismic phase travel time residual calculation and environmental interference identification, the system can automatically identify and classify waveform structure anomaly data, effectively eliminating non-seismic events caused by sensing errors, source location deviations and external disturbances, improving the accuracy of seismic observation data and providing high-quality data support for subsequent seismic event extraction and emergency response.

[0020] Furthermore, by determining the architecture performance based on the write amplification factor, the performance bottleneck of the data pool architecture can be accurately identified, effectively avoiding the computational burden caused by complex multi-parameter monitoring. By setting reasonable thresholds, performance anomalies of the storage system under high concurrency or large data volume write conditions can be detected in a timely manner, promoting the rational allocation and optimization of storage resources, improving the overall write efficiency and stability of the system, and ensuring the efficient and reliable operation of the data pool. Attached Figure Description

[0021] Figure 1 This is a connection diagram of the system built based on the data resource pool according to an embodiment of the present invention; Figure 2 This is a connection diagram of the non-earthquake event exclusion module according to an embodiment of the present invention; Figure 3 This is a connection diagram of the physical performance testing module according to an embodiment of the present invention; Figure 4 This is a schematic diagram showing the connection of the rationality verification unit in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0023] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0024] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0025] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0026] Please see Figure 1 As shown, it is a connection diagram of a data resource pool construction system according to an embodiment of the present invention. The present invention provides a data resource pool construction system, including: The data acquisition module is used to periodically acquire raw seismic waveform data and obtain the comparison results of the actual storage time and the standard storage time for the current acquisition cycle. The non-seismic event exclusion module is connected to the data acquisition module and is used to determine the category of the original seismic waveform data based on the interference event exclusion parameters after the first storage duration comparison result. The physical performance detection module, which is connected to the non-seismic event exclusion module, is used to perform a hierarchical detection operation based on hardware operating status parameters, software operating status parameters and architecture operating status parameters when the original seismic waveform data is the first observation data category, so as to determine whether there is a fault in the data pool, and to determine the corresponding physical performance fault category when there is a fault. Among them, physical performance failures include hardware performance failures, storage logic performance failures, or architecture performance failures. The anomaly cause tracing module, which is connected to the non-seismic event exclusion module, is used to remove interfering events in the second observation data category and perform anomaly cause judgment operation to determine whether there is a fault in the data pool, and to determine the corresponding data behavior performance fault when there is a fault. Specifically, when a performance failure result with no data behavior is obtained, the standard storage time is adjusted based on the time adjustment factor.

[0027] In this embodiment, multiple types of seismic sensing devices are deployed at the acquisition site, including short-period strong motion meters, broadband seismographs, and high-precision accelerometers. Based on the frequency band response characteristics of different devices, raw seismic waveform data of specific frequency bands are acquired. Furthermore, all sensing devices are equipped with GNSS synchronization devices to ensure data time consistency between devices. The standard storage duration is set to 1 second. The actual storage duration within the current period is obtained and compared with the standard storage duration. If the actual storage duration exceeds the standard storage duration, the first storage duration comparison result is obtained, and the storage time for this period exceeds the normal setting. If the actual storage duration is less than or equal to the standard storage duration, a second storage duration comparison result is obtained. Subsequent analysis steps are then performed to determine if there are any data behavior performance failures, and the standard storage duration is adjusted if such failures are found.

[0028] By continuously monitoring waveform data collected over a period of time, and combining non-seismic event identification, physical performance testing, and anomaly tracing mechanisms, the system differentiates performance issues at the hardware, software, architecture, and data behavior levels. Through waveform structure feature comparison, spatial residual calculation, and environmental disturbance analysis, it achieves accurate identification and filtering of non-seismic events, ensuring accurate data classification. In terms of physical fault identification, the system effectively locates structural or component-level defects in the storage pool using a hierarchical parameter detection model. Simultaneously, by combining frequency domain energy drift analysis and write distribution status judgment, it performs dynamic time threshold correction and scheduling optimization operations to achieve real-time adjustment of the data pool load status. The system possesses high adaptability and scalability, improving the processing efficiency and reliability of raw data, and is suitable for high-frequency observation, high-concurrency write seismic monitoring, and earthquake emergency application scenarios.

[0029] See Figure 2 As shown, it is a connection diagram of the non-earthquake event exclusion module in an embodiment of the present invention; Specifically, the non-seismic event exclusion module includes a parameter discrimination unit and a category determination unit; The parameter discrimination unit is used to perform multi-dimensional feature analysis on the original seismic waveform data based on interference event exclusion parameters in order to identify suspected abnormal data. The category determination unit is used to classify suspected abnormal data according to the data category to which the original seismic waveform data belongs, and output classification results, including a first classification result and a second classification result. The data category corresponding to the first classification result is the first observation data category, and the data category corresponding to the second classification result is the second observation data category.

[0030] In this embodiment, the interference event exclusion parameters include waveform structure characteristic parameters, seismic phase arrival time information, and environmental interference characteristic parameters. The multidimensional feature analysis process involves analyzing the interference event exclusion parameters to determine whether there are waveform structure anomalies in the original seismic waveform data, and when waveform structure anomalies are present, treating them as suspected anomalies. Specifically, the parameter discrimination unit includes a waveform feature analysis subunit, a spatial residual calculation subunit, an environmental interference discrimination subunit, and a scene enhancement recognition subunit; The waveform feature analysis subunit is used to extract several waveform structure feature parameters from the original seismic waveform data, and to determine waveform structure anomaly data based on the comparison results of each waveform structure feature parameter with the corresponding judgment threshold or interval. The spatial residual calculation subunit is used to calculate the travel time residual value of the source inversion fitting based on the seismic phase arrival time information, and to determine the waveform structure anomaly data based on the comparison result of the source positioning deviation and the preset tolerance threshold. The environmental interference discrimination subunit is used to identify abnormal waveform structure data caused by environmental factors based on the comparison results of environmental interference characteristic parameters and corresponding judgment thresholds. The environmental disturbance characteristic parameters include meteorological disturbance index, electromagnetic field strength and traffic harmonic characteristics; The scene enhancement recognition subunit is used to re-perform parameter comparison on the data processed by the waveform feature analysis subunit, the spatial residual calculation subunit, and the environmental interference discrimination subunit. The waveform structure characteristic parameters in this embodiment include the maximum value of waveform amplitude distribution, the rate of change of envelope structure, amplitude duration, and frequency content main frequency bandwidth; The maximum waveform amplitude distribution value is used to extract the amplitude corresponding to the largest amplitude value among the three channels. The standard waveform amplitude distribution maximum value is set to ±5×10. 4 When the maximum value of the waveform amplitude distribution is greater than the maximum value of the standard waveform amplitude distribution, the seismic waveform structure is determined to be abnormal. The envelope structure change rate is used to calculate the slope of the envelope line change based on the moving window method. The standard envelope structure change rate is set to 0.15 counts / ms. When the envelope structure change rate is greater than the standard envelope structure change rate, the seismic waveform structure is abnormal. Amplitude duration is used to determine the duration of continuous samples that are more than 3 times the average amplitude. The standard amplitude duration range is set as [0.5s, 15s]. When the amplitude duration is outside the standard amplitude duration range, the seismic waveform structure is abnormal. The frequency content main frequency bandwidth is used to extract the main frequency energy range based on FFT analysis. The standard frequency content main frequency bandwidth is set as [0.3Hz, 12Hz]. When the frequency content main frequency bandwidth is outside the standard frequency content main frequency bandwidth, the seismic waveform structure is abnormal. The data is compared item by item with each threshold or interval using the above four waveform structure characteristic parameters. If any characteristic parameter does not meet the judgment threshold or interval requirements, it is considered an earthquake waveform structure anomaly. The arrival time information of the seismic phases includes the first arrival time of the P-wave, the first arrival time of the S-wave, the observed travel time, the theoretical travel time, the residual values ​​of the two seismic phases, and the preset tolerance threshold. The travel time residual is the residual value between the two phases. The spatial residual calculation subunit processes the three-component raw waveform data using the STA / LTA algorithm, automatically identifying the first arrival times of the P-wave and S-wave, and denoting them as follows: and ; Based on the geographic coordinates of the data acquisition sites and the regional 1D velocity model, the theoretical travel time under the assumed source location was calculated using the least squares inversion method. and ; Calculate the residual values ​​of the two phases for the observed travel time and predicted travel time at each acquisition station: ; Set the preset tolerance threshold to 0.6 seconds; The residual values ​​of the two seismic phases are compared with the preset tolerance thresholds respectively; If the residual value of any seismic phase is greater than the preset tolerance threshold, the source location corresponding to the current seismic waveform is inaccurate, and abnormal waveform structure data is identified. If the residual values ​​of both seismic phases are less than or equal to the preset tolerance threshold, the current seismic waveform has reasonable source consistency under the 1D velocity model of the corresponding region and is not considered as waveform structure anomaly data. The environmental interference discrimination subunit is used to identify abnormal waveform structure data caused by non-earthquake sources during the process of receiving seismic waveform data by setting several environmental interference characteristic parameters and their corresponding judgment thresholds, and to remove or mark them as interference events. Environmental disturbance characteristic parameters include wind speed and thunderstorm activity indicators; The wind speed index is used to obtain real-time wind speed information based on the meteorological sensing equipment deployed at the collection site, and compare the current wind speed value with a preset wind speed judgment threshold, which is set to 10.8 m / s. When the wind speed value is greater than the preset wind speed judgment threshold, the waveform data of the current period is abnormal waveform structure data caused by environmental factors of the wind speed index. When the wind speed value is less than or equal to the preset wind speed judgment threshold, the waveform data for the current period is normal. The thunderstorm activity index is used to identify thunderstorm activity based on lightning detection data within the collection period. The thunderstorm activity index includes the number of lightning strikes and the change in lightning amplitude per unit time. The preset thunderstorm activity frequency threshold is set to 1 time / min, and the number of lightning strikes is compared with the preset thunderstorm activity frequency threshold. If the number of lightning strikes exceeds the preset thunderstorm activity frequency threshold, the current waveform data is abnormal waveform structure data caused by environmental factors of thunderstorm activity; If the number of lightning strikes is less than or equal to the preset thunderstorm activity frequency threshold, the current waveform data is normal, and normal earthquake waveform structure data is obtained. If waveform structure anomaly data is obtained, the first classification result is obtained, and the corresponding data category is the first observation data category; If normal seismic waveform structure data is obtained, the second classification result is obtained, and the corresponding data category is the second observation data category; Among them, the abnormality of the waveform structure data corresponding to the first classification result is caused by environmental factors; The abnormality of the waveform structure data corresponding to the second classification result is not caused by environmental factors.

[0031] By setting multidimensional feature parameters and judgment thresholds, and combining waveform structure analysis, seismic phase travel time residual calculation and environmental interference identification, the system can automatically identify and classify waveform structure anomaly data, effectively eliminate non-seismic events caused by sensing errors, source location deviations and external disturbances, improve the accuracy of seismic observation data, and provide high-quality data support for subsequent seismic event extraction and emergency response.

[0032] See Figure 3 The diagram shown is a connection schematic of the physical performance testing module in an embodiment of the present invention. Specifically, the physical performance testing module includes a fault detection unit and an architecture defect detection unit; The fault detection unit is used to detect the hardware and software operating status parameters in the storage system to determine whether there is a hardware performance fault or a storage logic performance fault in the data pool. The architecture defect detection unit is used to detect the architecture operating status parameters in the storage pool in order to determine whether there are any data pool architecture performance failures.

[0033] In this embodiment, the architecture defect detection unit uses the write amplification factor as the criterion for judgment. The write amplification factor is the ratio of the actual amount of data written to the storage system during this acquisition period to the amount of data submitted for writing by the business. Set the standard write amplification factor to 1.8; Compare the write amplification factor with the standard write amplification factor; If the write amplification factor is greater than the standard write amplification factor, the current data pool has an architectural performance failure, indicating that the storage system has unnecessary overhead such as data copying, rewriting or log replay when processing write tasks. If the write amplification factor is less than or equal to the standard write amplification factor, the current data pool architecture is operating normally.

[0034] By using write amplification factor to determine the architecture performance, the system can accurately identify performance bottlenecks in the data pool architecture, effectively avoiding the computational burden caused by complex multi-parameter monitoring. By setting reasonable thresholds, it can promptly detect performance anomalies in the storage system under high concurrency or large data volume write conditions, promote the rational allocation and optimization of storage resources, improve the overall write efficiency and stability of the system, and ensure the efficient and reliable operation of the data pool.

[0035] Specifically, the fault detection unit includes a hardware fault detection subunit and a software fault detection subunit; The hardware fault detection subunit is used to detect the hardware operating status parameters in the storage system in order to determine whether there is a data pool hardware performance fault. The software fault detection subunit is used to detect the software running status parameters in the storage system in order to determine whether there is a data pool storage logic performance fault.

[0036] In this embodiment, the hardware fault detection subunit uses SMART technology to monitor the temperature and bad sector quantity parameters of the storage device to determine whether there is a hardware performance fault. The software fault detection subunit collects storage system operation logs based on log anomaly analysis technology, and uses rule engines or anomaly detection algorithms to identify storage logic anomalies, such as data consistency errors, file system crashes or deadlocks, thereby determining whether there are storage logic performance failures.

[0037] By employing SMART technology to monitor hardware operating status in real time, timely warnings and accurate identification of storage device hardware failures are achieved, improving the stability and reliability of the data pool. Furthermore, based on log anomaly analysis technology, logical anomalies in the storage system are automatically identified, effectively preventing data corruption or service interruption caused by storage logic failures, and enhancing the overall security of the data pool operation.

[0038] Specifically, the anomaly cause tracing module includes a distribution rationality verification unit and a write load distribution identification unit; The distribution rationality verification unit is used to extract and analyze the geographical location information of the original seismic waveform data corresponding to the second observation data category, and based on the waveform content, timestamp and spatial coverage within the same acquisition cycle, identify duplicate data, extract the area of ​​the overlapping region and perform difference processing. The write load distribution identification unit is used to identify and schedule the write load status of the data pool.

[0039] See Figure 4 As shown, it is a connection diagram of the distribution rationality verification unit in an embodiment of the present invention; Specifically, the distribution rationality verification unit includes a path scheduling identification subunit; The path scheduling identification subunit is used to analyze the geographical location information of the observation data.

[0040] Specifically, the distribution rationality verification unit also includes a duplicate observation content cleanup subunit; The duplicate observation content cleaning subunit is used to compare waveform data, timestamps and spatial coverage under the same acquisition period. If the data are completely consistent and the geographical coverage overlaps, the difference processing is performed to remove the redundant parts.

[0041] Specifically, the duplicate observation content cleaning subunit includes: Obtain the current data acquisition cycle information; Raw seismic waveform data within the same natural day and historical acquisition period were selected as comparison objects. Compare the waveform content and timestamp of the current earthquake observation data with the comparison object to determine whether they are completely consistent; Analyze the geographical coverage of the comparison objects to determine whether there is geographical coverage overlap, and extract the overlapping area data when the waveforms are consistent, the timestamps are the same and the geographical coverage overlaps. Perform interpolation calculations on the overlapping data to remove redundant parts.

[0042] In this embodiment, the acquisition cycle information includes the corresponding natural day label and timestamp. The corresponding natural day label is extracted and used as the retrieval basis. The original seismic waveform data received in the historical acquisition cycle within the natural day is selected from the seismic data cache module as the comparison object set. The current seismic observation data is the data corresponding to the second observation data category acquired in the current monitoring cycle. The raw seismic waveform data of the current acquisition period is used to generate waveform summary information using the standard hash fingerprint generation algorithm. Based on the timestamp order, the raw seismic waveform data of the historical acquisition period are selected sequentially from the most recent to the oldest for comparison. The hash fingerprint is compared to determine whether the waveform content is completely consistent. Simultaneously, the sampling start time and sampling frequency are compared to determine whether the timestamps are consistent; If there is a historical acquisition period that is identical to the waveform content and timestamp of the current acquisition period, then the original seismic waveform data in that historical acquisition period is considered to be repeated in terms of timestamp and waveform. Based on the above repetition, the geographic coordinate information corresponding to the comparison object and its preset service influence radius are extracted to construct its geographic coverage. Extract the coordinates and corresponding data acquisition radius of the original seismic waveform data within the current acquisition period to establish the geographical coverage of the current data; Based on the GIS spatial overlay analysis method, the overlap of the two geographical coverage areas is calculated to determine whether there is any overlap in geographical coverage areas. The spatial overlap threshold is set at 5 km². If the calculated area of ​​the overlapping region is greater than the spatial overlap threshold, it is determined that there is geographical overlap, and the current observation data is considered to have redundancy in the spatial domain, requiring a cleanup operation. The calculation of the overlapping part can be based on existing vector buffer overlay analysis techniques, such as spatial intersection operation based on GeoJSON format; When the three conditions of waveform content consistency, timestamp consistency and spatial coverage overlap are met simultaneously, the system determines that the data constitutes duplicate observation content, extracts the overlapping area data, and the overlapping area data is the original seismic waveform data corresponding to the geographical coverage overlapping area. For repeated observation data that meet the conditions of consistent waveform content, consistent timestamps, and overlapping geographical coverage, the original seismic waveform data corresponding to the overlapping geographical coverage area is extracted as the target for difference processing. Prioritize retaining data records with earlier access order in historical collection cycles, perform sample point difference calculation based on a unified time window, and directly set overlapping samples in the current collection cycle to zero.

[0043] By jointly determining waveform content, timestamps, and spatial coverage, duplicate observation data can be effectively identified and removed, avoiding interference from redundant data in source location and event identification. The introduction of hash fingerprint algorithm and GIS spatial overlay analysis technology improves the accuracy and automation of duplicate data identification.

[0044] Specifically, the write load distribution identification unit includes a waveform stacking pressure determination subunit; The waveform stacking pressure determination subunit is used to obtain the ratio of waveform writing amount to cleaning residue per unit time, obtain the actual waveform stacking ratio, identify whether there is writing stacking pressure based on the comparison result of the actual waveform stacking ratio and the waveform stacking ratio threshold, and call the backup pool when it is determined that there is writing stacking pressure.

[0045] In this embodiment, the waveform stacking pressure determination subunit is used to evaluate the data writing status within the current acquisition cycle; Obtain the waveform write volume W per unit time within the current acquisition period, and extract the total amount of residual data R that has not been cleaned within the current period. Combined with the system's preset data cleaning capacity C per unit time, calculate the cleaning residue ratio. The calculation formula is as follows: ; The ratio of waveform writing amount W per unit time to cleaning residue Multiplying them together gives the waveform stacking ratio P; The preset standard waveform stacking ratio is 200; Compare the waveform stacking ratio with the standard waveform stacking ratio; If the waveform stacking ratio is greater than the standard waveform stacking ratio, there is write stacking pressure in the current acquisition cycle, triggering the backup pool to be called. If the waveform stacking ratio is less than or equal to the standard waveform stacking ratio, the current write status is normal and the backup pool switch will not be triggered.

[0046] Based on the product of the waveform writing volume and the cleaning residue ratio per unit time, the waveform accumulation ratio is dynamically calculated and compared with the preset standard waveform accumulation ratio in real time. This accurately identifies whether there is writing accumulation pressure in the data pool and automatically triggers the backup pool when accumulation pressure is detected, thereby switching the writing path and avoiding data writing blockage, delay or loss caused by excessive load in the main data pool.

[0047] Specifically, the write load distribution identification unit includes a threshold calibration subunit; The threshold calibration subunit is used to adjust the storage time threshold to the calibration storage time based on the time adjustment factor when the first data write distribution judgment result is obtained.

[0048] Specifically, the process by which the threshold calibration subunit adjusts the storage time threshold to calibrate the storage time includes: Extract the center frequency energy peak within the current acquisition period, and obtain the left and right frequency energy peaks of the center frequency energy peak at fixed intervals, and calculate the left drift difference and right drift difference respectively; Choose the larger of the left drift difference and the right drift difference as the time delay difference; Calculate the time adjustment factor; The calibrated storage time is obtained by multiplying the time adjustment factor by the standard storage time. The time adjustment factor is the absolute value of the ratio of the difference between the time delay difference and the standard storage time to the standard storage time.

[0049] In this embodiment, the threshold calibration subunit is used to dynamically calibrate the original standard storage time based on the frequency energy distribution characteristics within the current acquisition cycle, provided that no data behavior performance fault is detected by the system. Extract the center frequency energy peak within the current acquisition period, and obtain the left and right frequency energy peaks of the center frequency energy peak at fixed intervals to calculate the left and right drift differences respectively; The center frequency energy peak refers to the frequency component corresponding to the maximum energy in the spectrum of the seismic waveform signal; the system uses Fast Fourier Transform (FFT) to perform frequency domain analysis on the waveform data acquired in the current period, and identifies the main peak point in the frequency energy distribution curve as the center frequency energy peak. The left frequency energy peak and the right frequency energy peak refer to the energy peak values ​​of adjacent frequency points sampled at a set frequency interval on both sides of the center frequency, respectively representing the energy shift of the center frequency towards lower and higher frequencies. The frequency spacing is set to ±0.5Hz; The difference between the peak energy of the left frequency and the peak energy of the center frequency is calculated as the left drift difference, and the difference between the peak energy of the right frequency and the peak energy of the center frequency is calculated as the right drift difference. The larger of the left and right drift differences is used as the time delay difference of the current acquisition cycle to evaluate the asymmetry or offset intensity of the waveform signal in the frequency domain. The difference between the time delay difference and the standard storage time, divided by the standard storage time and taken as the absolute value, constitutes the time adjustment factor for this period, i.e.: ; in, For time adjustment factors; This represents the delay difference, which is the larger of the left drift difference and the right drift difference; Standard storage time; Obtain the post-calibration standard storage time, which is the original standard storage time multiplied by the time adjustment factor, and use it as the new reference value for time control of data writing and cleaning in the current cycle.

[0050] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0051] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A data resource pool construction system, characterized in that, include: The data acquisition module is used to periodically acquire raw seismic waveform data and obtain the comparison results between the actual storage time and the standard storage time of the raw seismic waveform data in the current acquisition period. The non-seismic event exclusion module is connected to the data acquisition module and is used to determine the category of the original seismic waveform data based on the interference event exclusion parameters after the first storage duration comparison result. The physical performance detection module, which is connected to the non-seismic event exclusion module, is used to perform a hierarchical detection operation based on hardware operating status parameters, software operating status parameters and architecture operating status parameters when the original seismic waveform data is the first observation data category, so as to determine whether there is a fault in the data pool, and to determine the corresponding physical performance fault category when there is a fault. The abnormal cause tracing module is connected to the non-seismic event exclusion module. It is used to remove interference events when the original seismic waveform data is the second observation data category, and to perform an abnormal cause judgment operation to determine whether there is a fault in the data pool, and to trigger data behavior performance fault detection when there is a fault, and to determine whether to adjust the standard storage time according to the time adjustment factor based on the detection result. The physical performance failure categories include hardware performance failures, storage logic performance failures, and architecture performance failures.

2. The data resource pool construction system according to claim 1, characterized in that, The non-seismic event exclusion module includes a parameter discrimination unit and a category determination unit; The parameter discrimination unit is used to perform multi-dimensional feature analysis on the original seismic waveform data based on interference event exclusion parameters in order to identify suspected abnormal data. The category determination unit is used to classify suspected abnormal data according to the data category to which the original seismic waveform data belongs, and output classification results, including a first classification result and a second classification result. Wherein, the data category corresponding to the first classification result is the first observation data category, and the data category corresponding to the second classification result is the second observation data category.

3. The data resource pool construction system according to claim 2, characterized in that, The parameter discrimination unit includes a waveform feature analysis subunit, a spatial residual calculation subunit, an environmental interference discrimination subunit, and a scene enhancement recognition subunit. The waveform feature analysis subunit is used to extract several waveform structure feature parameters from the original seismic waveform data, and to determine waveform structure anomaly data based on the comparison results of each waveform structure feature parameter with the corresponding judgment threshold or interval. The spatial residual calculation subunit is used to calculate the travel time residual value of the source inversion fitting based on the seismic phase arrival time information, and to determine the waveform structure anomaly data based on the comparison result of the source positioning deviation and the preset tolerance threshold. The environmental interference discrimination subunit is used to identify abnormal waveform structure data caused by environmental factors based on the comparison results of environmental interference characteristic parameters and corresponding judgment thresholds. The environmental disturbance characteristic parameters include meteorological disturbance index, electromagnetic field strength and traffic harmonic characteristics; The scene enhancement recognition subunit is used to re-perform parameter comparison on the data processed by the waveform feature analysis subunit, the spatial residual calculation subunit, and the environmental interference discrimination subunit.

4. The data resource pool construction system according to claim 1, characterized in that, The physical performance testing module includes a fault detection unit and an architecture defect detection unit. The fault detection unit is used to detect the hardware and software operating status parameters in the storage system to determine whether there is a hardware performance fault or a storage logic performance fault in the data pool. The architecture defect detection unit is used to detect the architecture operating status parameters in the storage pool in order to determine whether there are any data pool architecture performance failures.

5. The data resource pool construction system according to claim 4, characterized in that, The fault detection unit includes a hardware fault detection subunit and a software fault detection subunit; The hardware fault detection subunit is used to detect the hardware operating status parameters in the storage system in order to determine whether there is a data pool hardware performance fault. The software fault detection subunit is used to detect the software running status parameters in the storage system in order to determine whether there is a data pool storage logic performance fault.

6. The data resource pool construction system according to claim 1, characterized in that, The abnormal cause tracing module includes a distribution rationality verification unit and a write load distribution identification unit; The distribution rationality verification unit is used to extract and analyze the geographical location information of the original seismic waveform data corresponding to the second observation data category, and based on the waveform content, timestamp and spatial coverage within the same acquisition cycle, identify duplicate data, extract the area of ​​the overlapping region and perform difference processing. The write load distribution identification unit is used to identify and schedule the write load status of the data pool.

7. The data resource pool construction system according to claim 6, characterized in that, The distribution rationality verification unit includes a path scheduling identification subunit; The path scheduling identification subunit is used to analyze the geographical location information of the observation data.

8. The data resource pool construction system according to claim 6, characterized in that, The distribution rationality verification unit also includes a duplicate observation content cleanup subunit; The duplicate observation content cleaning subunit is used to compare waveform data, timestamps and spatial coverage under the same acquisition period. If the data are completely consistent and the geographical coverage overlaps, the difference processing is performed to remove the redundant parts.

9. The data resource pool construction system according to claim 6, characterized in that, The write load distribution identification unit includes a waveform stacking pressure determination subunit; The waveform stacking pressure determination subunit is used to obtain the ratio of waveform writing amount to cleaning residue per unit time, obtain the actual waveform stacking ratio, identify whether there is writing stacking pressure based on the comparison result of the actual waveform stacking ratio and the waveform stacking ratio threshold, and call the backup pool when it is determined that there is writing stacking pressure.

10. The data resource pool construction system according to claim 6, characterized in that, The write load distribution identification unit includes a threshold calibration subunit; The threshold calibration subunit is used to adjust the storage time threshold to the calibration storage time based on the time adjustment factor when the first data write distribution judgment result is obtained.

Citation Information

Patent Citations

  • Automatic-building system establishment method based on service resource pools

    CN106777271A