Big data governance analysis method applied to water conservancy data center

By constructing spike pin belts, spark fragment pools, and uncompressible corridors in the water conservancy data center, and combining them with trigger trajectory strings and tidal threshold pendulums, the problem of loss of key event features caused by the non-stationary nature of hydrological data was solved. This enabled high-fidelity retention and dynamic adjustment of extreme signals, improving the accuracy of flood control scheduling and water resource management.

CN121598014APending Publication Date: 2026-03-03SHANDONG QIANYUAN ENG GRP CO LTD KENLI BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511780083.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

When faced with the non-stationary nature of hydrological monitoring data, the existing water conservancy data center suffers from a weakening or disappearance of key event characteristics due to the long-term dynamic compression and archiving mechanism. This leads to delays or inaccuracies in flood control scheduling and risk response, affecting the reliability and security of the information system.

Method used

By constructing spike pin bands, spark fragment pools, fragment passports, and uncompressible corridors, abrupt signals are identified and preserved. Furthermore, by dynamically adjusting the compression threshold through trigger trajectory strings and tidal threshold balance wheels, high-fidelity retention of extreme signals is achieved.

Benefits of technology

It significantly enhances the water conservancy data center's memory and early warning sensitivity to extreme events, providing more complete and timely high-value data support to ensure the accuracy of flood control scheduling and water resource optimization allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598014A_ABST
    Figure CN121598014A_ABST
Patent Text Reader

Abstract

The invention discloses a big data governance analysis method applied to a water conservancy data center, and the method comprises the following steps: carrying out the peak detection and marking of a sudden swelling and sudden falling signal along a timeline, and connecting the marking results according to a time adjacent relation to form a peak pin band; and on the basis of the spike pin strips, intercepting the original flow data within the preset duration before and after each spike pin strip to obtain extremely short segments, and splicing the intercepted extremely short segments according to a time sequence to form a spark segment pool. According to the method, the abrupt change signal is identified by constructing the pin belt, the extremely short segment is extracted, the passport is guided to enter the uncompressed corridor, and the compression threshold is dynamically adjusted by combining the trigger track string and the tidal threshold balance wheel, so that the high-fidelity retention of abnormal data is realized, and the extreme event identification and early warning capability is improved; and the data support efficiency of the water conservancy data center is enhanced without significantly increasing the storage pressure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a big data governance and analysis method applied to a water conservancy data center. Background Technology

[0002] Big data governance and analysis applied to water conservancy data centers refers to the construction of a comprehensive technical approach integrating data governance, quality control, and intelligent analysis for the vast, complex, and heterogeneous data system in the water conservancy industry. This method uses the water conservancy data center as its core operating platform, and establishes a data governance system covering the entire business chain by uniformly collecting, standardizing formats, semantically associating, and quality-verifying multi-dimensional data from hydrology, water resources, water ecology, water environment, flood control scheduling, and water project operation. Based on this, distributed computing, big data mining, and artificial intelligence analysis technologies are used to achieve in-depth analysis and dynamic visualization of key elements such as spatiotemporal characteristics, operational patterns, and risk trends, thereby supporting the intelligent and scientific management of water conservancy departments in flood warning, water resource optimization scheduling, water ecological protection, and decision-making.

[0003] The existing technology has the following shortcomings: In existing technologies, water conservancy data centers generally employ long-term dynamic compression and archiving mechanisms to reduce storage costs and computational load. However, this strategy has significant limitations when facing the non-stationary characteristics of hydrological monitoring data. When compression algorithms sample based on time density or statistical thresholds, low-frequency but high-value data fragments (such as short-duration flood peak signals, sudden rise and fall water level waveforms, and transient responses of gate micro-opening) are often treated as noise or anomalies and excessively thinned, resulting in the weakening or even disappearance of key event features in historical data. As information entropy continuously decays during compression, the overall data diversity and mutation sensitivity of the system are significantly reduced, making it impossible for analytical models trained or run on this data to identify the triggering sources and early evolutionary characteristics of extreme events. When sudden floods or non-periodic hydrological disturbances occur, the system lacks high-resolution historical anomaly samples, creating a decision-making gap. This leads to delays or inaccuracies in flood control scheduling, flow prediction, and risk response, seriously affecting the reliability and security of the water conservancy information system. Summary of the Invention

[0004] The purpose of this invention is to provide a big data governance and analysis method for water conservancy data centers to solve the problems mentioned in the background.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a big data governance and analysis method applied to a water conservancy data center, comprising the following steps: Peak detection and marking are performed on signals with sudden increases and decreases along the timeline. The marking results are connected according to the time adjacency to form a spike pin band. Based on the spike pin band, the original flow data is extracted within a predetermined time period before and after each spike pin band to obtain a very short segment, and the extracted very short segments are spliced ​​together in chronological order to form a spark segment pool. Generate unique identification information for each fragment in the spark fragment pool and form a fragment passport. Map the fragment passport back to the original flow data to identify the identity of the corresponding fluctuation segment. Uncompressed corridors are opened in the archiving path for data with fragment passports, and the data with fragment passports are stored in the uncompressed corridors according to the hierarchical archiving strategy to maintain the integrity of mutation details; By linking the corresponding segments of the passport in chronological order without compressing the corridor, a continuous trigger trajectory is formed as a precursory guidance chain for extreme events. The tidal threshold balance wheel is activated based on the trigger trajectory string to dynamically adjust the data compression threshold: the compression threshold is relaxed when the trigger trajectory string is in the rising phase and tightened when the trigger trajectory string is in the falling phase, thereby guiding the archiving path to complete adaptive control and achieve long-term fidelity retention of extreme signals.

[0006] Preferably, the peak pin band formation step is as follows: The original water flow data is read point by point in chronological order, the water level difference between consecutive time points is calculated, and sudden rise or fall signal points are marked according to the condition that the water level rises or falls continuously within a preset time interval and exceeds a predetermined range. Using each mutation signal point as the center, water level data is extracted forward and backward as the confirmation interval to determine whether there is a continuous mutation trend, and mutation peak points with complete trends are selected. The effective mutation peaks are arranged in chronological order, and adjacent peaks with time intervals not exceeding a preset range are grouped together to form a continuous spike pin band. For each spike pin band, calculate the maximum mutation amplitude, average mutation rate, total duration, and minimum peak interval, and record its start and end times, time interval, and internal peak point number.

[0007] Preferably, the steps for constructing the spark fragment pool are as follows: The data is extended forward by a predetermined time based on the start time of each spike pin band, and backward by a predetermined time based on the end time, forming a continuous data interval that includes the pre-mutation segment, the middle segment, and the post-mutation segment. Each data interval is divided into multiple extremely short time segments with fixed lengths, continuous start and end times, and no overlap between segments. Complete original pipeline data is retained within each extremely short time segment. All extremely short time segments are pieced together in chronological order to form a continuous data set covering an extended time interval, thus creating a spark fragment pool; Feature statistics are performed on all extremely short time segments in the spark fragment pool to generate structured descriptive information such as maximum water level fluctuation, minimum water level fluctuation, average rate of change, fluctuation density, and trend distribution.

[0008] Preferably, the steps for forming a fragment passport are as follows: Traverse each extremely short fragment in the spark fragment pool and extract the time identification field and hydrological statistical feature field to construct an identification field set; Based on the set of identification fields, a structured fragment passport is generated for each very short fragment, containing information such as number, start and end time, fragment pool number, mutation magnitude, change trend, fragment position and peak value. Map each segment passport back to the original stream data, and append the segment number and the position marker field within the segment to the original data fields; A unified passport index table is established for all passport fragments, which are then categorized, sorted, and have their core identification information stored according to the spark fragment pool number.

[0009] The preferred method for opening a corridor without compression is as follows: Extract the extremely short data segments with attached passport identifiers as the target dataset for protection and then merge and strip out the regular data stream; In the data archiving structure, uncompressible corridors are defined for the protected target data set, and dedicated storage paths are set according to time, event level, and source pool number; Implement a tiered archiving strategy based on the intensity of mutation characteristics, data scarcity, and correspondence with historical events, and configure storage protection measures corresponding to the archiving level. Establish a data mapping table to record the storage path, number of backups, retention period, and subsequent application status of each data segment, and realize controllable retrieval of data.

[0010] Preferably, in the hierarchical archiving strategy, data segments with high mutation intensity, high data scarcity and correlation with historical emergencies are classified as core mutation segments. They are stored in different physical storage paths using a dual-copy redundant storage method, and integrity verification operations are performed periodically.

[0011] Preferably, the continuous trigger trajectory string formation process is as follows: Extract data segments containing fragmented passports from the uncompressed corridor, construct a passport time list sorted in ascending order of start time, and build an initial index structure to trigger trajectory construction; Based on the passport time list, determine whether adjacent segments meet the concatenation condition and filter out segments that can be continuously constructed into trajectories; Connect the segments that meet the concatenation conditions in sequence to construct the basic trigger trajectory segment and record the time range and trend characteristics of the trajectory segment; The basic trigger trajectory segments are judged for time continuity and trend compatibility and aggregated into trigger trajectory strings; Generate a precursor guide chain identifier for the trigger trajectory string and establish its mapping relationship in the original data.

[0012] Preferably, the data compression threshold is dynamically adjusted based on the trigger trajectory string to start the tidal threshold balance wheel: the compression threshold is relaxed when the trajectory string is in the rising phase and tightened when the trajectory string is in the falling phase, so as to drive the archiving path to perform adaptive control and complete the long-term fidelity storage of extreme signals. The steps are as follows: Based on the trigger trajectory string, all segment passports are extracted and the trend phase attributes of the trajectory string, namely the rising phase, falling phase, and oscillation phase, are identified in chronological order. The compression threshold adjustment strategy is set according to the trend phase of the trajectory string. The compression threshold is relaxed in the rising phase, tightened in the falling phase, and a neutral compression threshold is maintained in the oscillation phase. Map the compression threshold adjustment strategy to the data archiving channel and write the compression strategy corresponding to the execution trend stage into the data of each segment; The compression control results of the tidal threshold balance wheel mechanism are periodically evaluated, and the compression threshold adjustment parameters are fine-tuned based on the evaluation results. Create a compression strategy execution file for each trigger trajectory string and record the history of compression threshold changes, trend stage switching time points, compression effect indicators, and archive path numbers.

[0013] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention constructs a spike-pin band to locate sudden fluctuations in signals, extracts key, extremely short segments, and assigns them unique identification passports. During data archiving, these segments are guided into an uncompressible corridor, enabling independent identification and high-fidelity preservation of abrupt fluctuations. Simultaneously, by constructing a trigger trajectory string and activating a tidal threshold oscillator, the data compression threshold is dynamically adjusted according to trends, ensuring a flexible balance between data compression and anomaly retention. The overall solution significantly improves the water conservancy data center's memory and early warning sensitivity to extreme events without significantly increasing overall storage pressure, providing more complete and timely high-value data support for subsequent flood control scheduling, water resource optimization and allocation, and intelligent model training. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the big data governance and analysis method of the present invention applied to a water conservancy data center. Detailed Implementation

[0015] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0016] This invention provides, for example Figure 1 The big data governance and analysis method applied to a water conservancy data center, as shown, includes the following steps: Peak detection and marking are performed on signals with sudden increases and decreases along the timeline. The marking results are connected according to the time adjacency to form a spike pin band. This step provides a method for identifying abrupt changes in water level in a water conservancy data center. The method aims to identify sudden rises and falls in water level within continuous raw flow data and mark these changes as a set of anomalous signals with defined time intervals, forming a time baseline for subsequent key data extraction. This method fully leverages the non-stationary nature of hydrological data, capturing and structurally identifying important data that might be overlooked by traditional compression mechanisms through well-defined steps. The specific steps are as follows: Based on the raw water flow data collected by the water conservancy data center, each data point is read sequentially over time. A continuous water level data sequence is extracted at a minute-by-minute time granularity, for example, 60 water level points from 10:00 to 11:00. For each time point, the water level difference between it and the previous time point is calculated, and the positive / negative direction and magnitude of the difference are recorded. A lower limit for determining a surge is set as a continuous rise in water level exceeding 1.0 meter within five minutes, and a sudden drop is defined as a drop exceeding 1.0 meter within five minutes. A sliding window is used to continuously scan the entire data sequence based on these criteria. Upon detecting a data segment that meets the conditions, the midpoint of that segment is recorded as an initial abrupt change signal point. For example, if the water level rises from 1.3 meters to 2.6 meters between 10:12 and 10:16, this data segment is recorded as a surge signal, marked at 10:14.

[0017] All identified sudden rises or falls in water level are trend-confirmed to eliminate false positives and select truly representative outliers. Specifically, for each sudden rise signal point, five minutes of water level data are taken both forward and backward as confirmation intervals. The trend structure of the signal point and the preceding and following data is compared to determine if a continuous sudden trend exists. For example, if the water level is stable before the signal point, and then continues to rise or fall for a period of time after the signal point with an amplitude of not less than 0.4 meters, then the signal point is identified as a peak point of a complete sudden trend. This confirmation process effectively eliminates occasional numerical deviations caused by measurement errors or minor fluctuations, ensuring that only high-value fluctuations with stable evolutionary characteristics are retained. For example, if a sudden rise signal point occurs at 10:28, supported by the preceding data from 10:23 to 10:27 as a stable level, and the subsequent data from 10:29 to 10:33 shows a continued rise of 0.5 meters, then 10:28 can be identified as a valid sudden rise point.

[0018] All confirmed spike or drop peaks are arranged chronologically, and the time difference between adjacent signal points is compared point by point. If the interval between two adjacent spike peaks does not exceed thirty minutes, they are considered as components of the same continuous spike process and grouped into the same group, forming a pin band. The start time of each pin band is the earliest spike peak, and the end time is the latest spike peak. All points within the band are temporally continuous or approximately continuous. For example, if three valid spike points appear at 10:05, 10:18, and 10:31, and the time difference between them is within thirty minutes, these three points are merged into one group, forming a pin band from 10:05 to 10:31. All peak points within each pin band are also numbered chronologically for easy subsequent location and reference.

[0019] Statistical attributes and structural descriptions are performed on each stitch band to provide a basis for decision-making in subsequent data extraction. Statistical attributes include: the maximum abrupt change amplitude in the stitch band, i.e., the absolute value of the maximum water level change at each abrupt change point; the average abrupt change rate, i.e., the total change in all peak values ​​within the band divided by the total duration; the total duration, i.e., the time length from the initial peak point to the final peak point; and the minimum peak interval, i.e., the time difference between the two closest abrupt change points within the band. For example, a stitch band may contain five abrupt change points, with a maximum water level change of 1.8 meters, an average water level change rate of 0.15 meters per minute, a total duration of 24 minutes, and a minimum peak interval of 3 minutes. These statistical attributes will serve as the basis for subsequently determining the key segment extraction window and will be stored together with the time interval of the stitch band to form a complete and traceable set of abrupt change features.

[0020] Based on the spike pin band, the original flow data is extracted within a predetermined time period before and after each spike pin band to obtain a very short segment, and the extracted very short segments are spliced ​​together in chronological order to form a spark segment pool. This step proposes a method for extracting neighboring data of abrupt change signals based on spike pin bands, in order to construct a high-resolution dataset covering the entire process before, during, and after the abrupt change. This method, through fixed-time-range truncation, continuous extremely short segment division, time-series splicing, and structured statistics, ultimately forms a dataset with identifiable value, providing fundamental support for subsequent hydrological event reconstruction and data compression and protection. Specifically, it includes the following steps: The complete time range of each spike pin band is determined, and the data extraction interval before and after is expanded accordingly. After detecting and marking abrupt fluctuations in anomalies, multiple spike pin bands have been obtained, each containing a start time and an end time. Based on the start time of each pin band, a data interval of seven to ten minutes is constructed as the extraction window by extending forward three minutes and backward three minutes from the end time. For example, if the start time of a pin band is 10:12 and the end time is 10:34, the expanded time interval is from 10:09 to 10:37. This expanded time interval fully covers the prelude, main mutation, and post-mutation stabilization stages of the mutation event, helping to capture the entire process of the generation, expansion, and decay of anomalous features. This method uses pin bands as anchor points to expand the extraction range, avoiding reliance on local windows of single data points for judgment, thereby enhancing the continuity and analyzability of the data.

[0021] The raw flow data within each extended time interval is divided into several extremely short time segments. Each time segment is no longer than one minute; in this embodiment, it is set to thirty seconds. Segments do not overlap, and their start and end times are strictly continuous. For example, from 10:09 to 10:37, a total of 28 minutes, this is divided into 56 thirty-second segments. The data content corresponding to each segment includes all raw sampled data within that thirty seconds, such as water level, flow velocity, and gate opening. Assuming a data sampling frequency of once per second, each segment will contain 30 data points. All segments record their start and end times, segment number, and raw data content to ensure clear sequence, complete content, and no loss of change information at any given moment. Establishing a set of extremely short segments using a fixed-length division method effectively captures the rapid evolution of abrupt changes, significantly improving the temporal sensitivity and local identification capabilities of the data.

[0022] Each pin band's corresponding extremely short segments are spliced ​​together chronologically to construct a complete, continuous time-segment data set, named the Spark Fragment Pool. The Spark Fragment Pool contains all extremely short segments from the start to the end of the expansion, its structure being an ordered sequence of segments, each containing multiple data points. Segments are automatically numbered according to their chronological order; for example, segment 1 is from 10:09 to 10:09:30, segment 2 is from 10:09:30 to 10:10, and so on, until the last segment. At the outermost layer of the Spark Fragment Pool, additional information is recorded, including the pin band number from which the pool originates, the original time interval, the number of segments, the start and end times, and a description of the data source, ensuring data traceability and manageable structure. By aggregating multiple extremely short segments into a pool, the Spark Fragment Pool provides continuous recording capabilities for mutation events, enabling subsequent processing steps to conduct correlation analysis, evolutionary identification, and compression control with higher timeliness and accuracy.

[0023] After the spark fragment pool is constructed, preliminary feature statistics and descriptive analysis are performed on its internal fragment content to form a data structure index and feature labels. Statistical indicators include the maximum and minimum water level fluctuations within the pool, the average rate of change, the maximum rate of change for a single fragment, the number of sudden rises and falls in fragments, and the fluctuation density within fragments. For example, in a spark fragment pool, the maximum water level rise is 2.1 meters, the minimum intra-fragment change is 0.02 meters, the maximum rate of change is 0.18 meters per 10 seconds, and the average fragment fluctuation density is 0.04 meters per second. Furthermore, the number of upward and downward trend fragments in the spark fragment pool can be counted, and their positions within the pool can be marked. All statistical results do not modify the original fragment data but are added as metadata to the fragment pool labels for use by subsequent compression control strategies, signal recognition strategies, and event prediction mechanisms.

[0024] Generate unique identification information for each fragment in the spark fragment pool and form a fragment passport. Map the fragment passport back to the original flow data to identify the identity of the corresponding fluctuation segment. This step proposes a method to construct an independent identifier for each extremely short fragment in the spark fragment pool and embed it into the original pipeline data. The aim is to achieve precise location, continuous tracking, and multi-dimensional referencing of abrupt change signals in the time series, providing a data foundation for subsequent archiving, classification, compression protection, and event modeling. This method is implemented sequentially through information extraction, structure generation, data mapping, and index archiving steps, detailed below: For the constructed spark fragment pool, each extremely short fragment is sequentially traversed, and its basic identification information is extracted and organized into a structured field set. Each extremely short fragment has a unique start and end time, forming the basic fields for its time identification component. Based on this, hydrological data points within the fragment are read, and key statistical features within the fragment are calculated, including maximum water level, minimum water level, average water level, total water level fluctuation, direction of water level change, rate of water level change, and duration of water level change. Taking an extremely short fragment from 10:12:30 to 10:13 as an example, this fragment contains 30 data points (one per second). The calculation shows that the water level rose from 2.10 meters to 2.38 meters, with a total fluctuation of 0.28 meters, an average rise of 0.0093 meters per second, showing a continuous upward trend, and reaching its maximum rate of change at the 22nd second. These values ​​are then aggregated to form a unique set of identification fields, serving as the basis for constructing subsequent identity information. Unlike existing technologies that use fuzzy indexing or sequential numbering for segment recognition, this method constructs feature attributes based on the data content itself, ensuring the irreplaceability of each segment in the global data.

[0025] Based on the set of identification fields, a complete and well-structured identity record, named a segment passport, is generated for each extremely short segment. This segment passport consists of multiple fields, each corresponding to a key feature of the segment, including but not limited to: segment number (e.g., F0072), start time, end time, spark segment pool number (e.g., SP003), corresponding spike pin number (e.g., B005), maximum water level, minimum water level, total amplitude, average rate, trend direction (rising, falling, or stable), segment position (segment number), whether it contains abrupt change signal points, peak position, and water level fluctuation variance. For example, the passport for fragment F0072 can be described as follows: start time 10:12:30, end time 10:13, originating from spark fragment pool SP003, belonging to pin zone B005, water level rose from 2.10 meters to 2.38 meters, with a total fluctuation of 0.28 meters, the maximum rate of change occurring at the 22nd second, the trend being a steady rise, located in the 9th fragment, containing one surge signal point, and the water level fluctuation variance is 0.00034. All field values ​​are stored in a standardized format and uniformly numbered, forming a complete and clearly structured set of fragment passports. Compared to traditional techniques where fragment identification relies solely on timestamp numbers, this method significantly improves the ability to characterize fluctuation features, changing trends, and abrupt changes, enhancing the accuracy and relevance of data citation.

[0026] Each generated segment passport is mapped back to its corresponding original flow data, and field-level marking is performed in the original data structure to achieve bidirectional binding between extremely short segments and the original data. In each piece of original flow data, each data point contains original attribute fields such as sampling time, water level value, and measurement point number. This step adds two marker fields: the segment number and the segment's relative position. Continuing with the example of segment F0072, its corresponding 30 data points, from 10:12:30 to 10:13, will have their segment number F0072 added to their attribute fields, and the current position recorded as the 1st to 30th second within the segment. Through this marking method, the segment passport to which a data point belongs can be quickly located at any point in time, and the original data content can be directly traced through the segment passport, achieving accurate data tracing and segment management. This marking strategy solves the technical problems of existing technologies where data compression results in the loss of contextual information and the inability to track the evolution of local data, making it particularly suitable for data scenarios requiring long-term storage and multi-dimensional analysis.

[0027] After generating and labeling all fragment passports, a unified passport index table covering all spark fragment pools is created, and all fragments are categorized, sorted, and structurally summarized. This passport index table uses the spark fragment pool number as the primary index field. Each record contains core information such as fragment number, start and end time, corresponding pin zone, direction of change, maximum fluctuation amplitude, rate of change, fragment location, and number of mutation points. For example, a record in the index table might be: Fragment number F0072, time interval 10:12:30 to 10:13:00, source SP003-B005, upward trend, maximum water level 2.38 meters, minimum water level 2.10 meters, number of mutation points 1, fluctuation rate 0.28 meters per 30 seconds. This table allows for filtering by time interval, fluctuation type, trend direction, and change intensity, facilitating subsequent data extraction, trend modeling, compression decision-making, and event reconstruction workflows. While maintaining the integrity of the original data, this index table constructs a high-resolution structural labeling system, serving as a bridge between mutation feature identification and data protection and regulation, significantly improving the accuracy and efficiency of data governance.

[0028] Uncompressed corridors are opened in the archiving path for data with fragment passports, and the data with fragment passports are stored in the uncompressed corridors according to the hierarchical archiving strategy to maintain the integrity of mutation details; This step proposes a high-fidelity data processing method for the archiving process of a water conservancy data center. Based on a pre-constructed fragment passport, raw flow data carrying mutation characteristics is allocated to an independent uncompressed storage path and stored differentially through a hierarchical archiving method to maintain the integrity, continuity, and retrieval of mutation data in the long term. The specific implementation steps of this method are as follows: Based on the completed raw data labeling results, all extremely short data segments with segment passport identifiers are extracted as the protection target data set in this step. Each data segment with a passport has been bound to the raw data through information such as passport number, start and end time, water level change characteristics, trend category, and abrupt change signal intensity level. Therefore, by traversing the newly added segment identifier field in the entire flow data, all data areas that need to enter the uncompressed protection path can be accurately located. For example, in the flow data segment from 10:12 to 10:38, five segments with abrupt increases in trend are identified by the passport number, located at 10:15 to 10:15:30, 10:17 to 10:17:30, 10:18 to 10:18:30, 10:20 to 10:20:30, and 10:21 to 10:21:30. These data segments will be separated from the regular data stream and processed separately during subsequent archiving. This step differs from traditional logging data archiving methods, which often compress the entire data segment as a whole, making it difficult to extract structured data to address abnormal fluctuations. This method completes the screening of key data segments before the archiving process, ensuring the foundation for subsequent fidelity protection measures.

[0029] A dedicated uncompressed path, called the uncompressed corridor, is designated within the data archiving structure to store passport fragment data. This corridor ensures the original accuracy and detail of high-value data. It exists as a physical directory partition, for example, storing uncompressed data uniformly in a path named "Corridor" on the D drive partition of the archiving device, further subdivided by time, event level, and source pool number. For instance, the data fragment with passport number F0063, with a time range of 10:18 to 10:18:30, sourced from Spark Fragment Pool SPK005, and an event level of Level 1, will be saved to the path D: / Corridor / Level1 / SPK005 / 20251101-F0063.dat. This file retains the second-by-second sampling data from the original water level record of this fragment, including timestamps, water level values, flow rates, and mutation identification tags, without any format conversion, precision compression, or structural reduction operations. By assigning a dedicated storage path and fine-grained naming rules to each segment, the corresponding data segment can be quickly located in various usage scenarios such as data query, comparison, verification, and model invocation, eliminating the risk of content loss, format corruption, and retrieval failure caused by compression.

[0030] For all data fragments entering the uncompressed corridor, a hierarchical archiving strategy is implemented based on dimensions such as the intensity of their mutation characteristics, data scarcity, and correspondence with historical events. Data fragments of different levels are stored in different priority archiving levels, and corresponding storage protection measures are configured. This embodiment sets three archiving levels: Level 1 is core mutation fragments, which must be retained for more than three years and use a dual-redundant replica storage mechanism; Level 2 is secondary abnormal fragments, retained for one to three years, stored with only a single replica, but compression is prohibited; Level 3 is fragments with small fluctuations but located in the edge area of ​​abnormal events, retained for no less than six months, and whether to enter the cold data pool can be determined based on the analysis and evaluation results after the retention period expires. For example, passport number F0063, due to its sudden rise rate exceeding 0.45 meters per minute and its historical event being a sudden flood discharge process during a typhoon, is classified as a Level 1 fragment. Synchronous replicas will be generated on two storage devices, stored respectively at D: / Corridor1 / Level1 / and E: / Corridor2 / Level1 / , and integrity verification tasks will be performed periodically. This tiered strategy achieves a match between storage resources and data importance, solving the problem in existing technologies where high-value data and low-frequency data are treated equally due to a single archiving strategy, resulting in resource waste or overwriting of critical information.

[0031] After storing all passport fragment data in the uncompressed corridor, a data mapping table linked to the passport fragments is constructed to record the location, storage method, and protection status of each fragment in the archiving path, enabling precise retrieval and efficient searching. This mapping table is sorted by fragment number, and each record includes the following fields: fragment number, archiving start time, archiving end time, data event level, storage path, number of backups, retention period, whether it participates in model training, and whether it participates in visualization analysis. For example, F0063's mapping record is: F0063, 10:18 to 10:18:30, Level 1, path D: / Corridor / Level1 / SPK005 / 20251101-F0063.dat, dual copies, retention for three years, training use is enabled, visualization analysis is enabled. This mapping table serves as the foundation for the data retrieval process and can be embedded into upper-layer application interfaces to achieve various application scenarios such as event tracing, model supplementary training, and high-density visualization, fully connecting the logical links between data identification, storage, and retrieval. Unlike existing technologies that generate random archive directories automatically through compression strategies, this method uses passport control as its core to construct a controllable, searchable, and interpretable archive network.

[0032] By linking the corresponding segments of the passport in chronological order without compressing the corridor, a continuous trigger trajectory is formed as a precursory guidance chain for extreme events. This step proposes a method for constructing continuous trajectories of abrupt change signals within an uncompressed corridor. Based on chronological order, multiple data segments with fragment passports are sequentially concatenated to generate a trigger trajectory string with a complete evolutionary relationship. This trajectory string is further aggregated into a precursor guidance chain reflecting the evolution of extreme hydrological events, thereby improving the accuracy and timeliness of sudden event identification. The method includes the following steps: Based on all archived passport fragments in the uncompressed corridor, the time range, trend direction, mutation characteristic level, and corresponding spark fragment pool number of each passport fragment are extracted to construct a passport time list sorted in ascending order of start time. Each record includes fragment number, start time, end time, trend attribute, water level change magnitude, and affiliated fragment pool. For example, fragment number F0091 has a start time of 10:15:00 and an end time of 10:15:30, with a trend attribute of continuous rise, and belongs to spark fragment pool number SPK005. Through this time-driven sorting mechanism, fragments from different sources are established with a unified time axis basis, forming the initial index structure for trigger trajectory construction. This time list not only identifies the time structure of fragment distribution but also provides a unified reference system for subsequent time continuity judgment, which is significantly different from existing methods that isolate data fragments and lack evolutionary clues.

[0033] Based on the information of adjacent segments in the time list, it is determined whether the concatenation condition is met, and segment pairs that can continuously construct trajectories are selected. The judgment criteria include three specific standards: First, the time interval between consecutive segments must not exceed the set maximum tolerance, for example, the time interval must be less than 60 seconds; second, the trend direction of the two segments must be consistent, both being an upward, downward, or oscillating trend; third, there must be a time overlap or continuity in the segment pool sequence between the two segments. If the time interval between segments F0091 and F0092 is 45 seconds, both have an upward trend, and they belong to SPK005 and SPK006 respectively, and the time intervals of these two spark segment pools are 10:14:30 to 10:16:00 and 10:15:45 to 10:17:15 respectively, they clearly have time continuity. Therefore, F0091 and F0092 meet the concatenation condition and are included in the initial trigger pair set. Segments that do not meet all three conditions will be excluded to ensure that the concatenated trajectory has a clear time progression logic and trend continuity.

[0034] After screening the fragment pairs, the connectable fragments are sequentially linked according to the time order of the trigger pairs to construct basic trigger trajectory segments. Each trajectory segment consists of two or more consecutive adjacent fragments and includes the segment's start time, end time, fragment number sequence, total fluctuation amplitude, average rate of change, trend direction, and trajectory segment number. Continuing with the example above, if fragments F0091, F0092, and F0093 all meet the linking condition, they constitute trajectory segment T001, with a start time of 10:15:00 and an end time of 10:16:30, a total fluctuation amplitude of 0.94 meters, an average water level change rate of 0.31 meters every 30 seconds, and a continuous upward trend. Each fragment in trajectory segment T001 is recorded as a logical substructure in the trajectory segment index, and its status label in the original data is updated to indicate that it has been included in a specific trajectory. This trajectory segment construction method ensures logical connectivity between fragments, avoiding the context loss problem caused by the single-point recording method for abrupt events in traditional technologies.

[0035] After constructing multiple trajectory segments, their temporal continuity and trend compatibility are further evaluated. Adjacent trajectory segments are then aggregated to form trajectory strings with stronger continuity and higher event levels. Aggregation conditions include a time interval of no more than 120 seconds between the end of a trajectory segment and the beginning of the next segment, no significant reversal in the trajectory trend, and continuous and unbroken time intervals between the spark fragment pools to which the trajectory segments belong. For example, the time intervals of trajectory segments T001, T002, and T003 are 10:15 to 10:16:30, 10:16:45 to 10:17:30, and 10:17:35 to 10:18:15, respectively. All three have a sudden increase in trend direction, and their time range in the spark fragment pool is 10:14:30 to 10:18:30, which meets the aggregation requirements. Therefore, they are merged into trajectory string TS001. The trajectory string TS001 contains 9 passport segments with a total fluctuation amplitude of 2.86 meters and an overall upward trend. This trajectory string will serve as an important candidate for a potential extreme event trigger chain and will proceed to the next stage of the event attribution and risk assessment process. This aggregation method strengthens the evolutionary logic between data and overcomes the limitations of traditional single-segment analysis in dynamic modeling.

[0036] The constructed trajectory strings are labeled as precursor guidance chains, and their mapping relationship in the original data sequence is recorded. Each precursor guidance chain has an independent identifier in the data structure, including chain number, start time, end time, number of component segments, cumulative fluctuation amplitude, trend type, associated event category, and reference status. For example, the precursor chain number is TS001, the start time is 10:15:00, the end time is 10:18:15, the total water level fluctuation is 2.86 meters, the event type is rapid rise, it belongs to the seasonal rainstorm warning scenario, and the current status is activated and used for model analysis. In the original flow data, all data sampling points covered by this precursor chain will have a new precursor chain number field added, indicating which trigger trajectory the data point belongs to. This information is also synchronously written to the tracking index file in the corresponding directory of the uncompressed corridor for various purposes such as model reading, risk warning, and visualization.

[0037] The tidal threshold balance wheel is activated based on the trigger trajectory string to dynamically adjust the data compression threshold: the compression threshold is relaxed when the trigger trajectory string is in the rising phase and tightened when the trigger trajectory string is in the falling phase, thereby guiding the archiving path to complete adaptive control and achieve long-term fidelity retention of extreme signals. This step proposes a method based on a dynamic compression strategy for trajectory trends. It utilizes the trend change phases of the triggered trajectory sequence as the control basis and adjusts the data compression threshold in real time through a tidal threshold balance wheel mechanism. This guides the archiving path to achieve adaptive adjustment of the compression scale, thereby enabling long-term faithful retention of signals with extreme abrupt changes. The method includes the following steps: Based on each constructed trigger trajectory string, all its segment passports are extracted, and the segments are traversed and analyzed in chronological order to identify the trend phase attributes of the trajectory string. The trend phase division is based on the direction, amplitude, and rate of water level change of continuous segments. When the water level of three or more adjacent segments shows a continuous unidirectional rise, and the average water level amplitude of each segment exceeds 0.10 meters and the average rise rate is not less than 0.002 meters per second, the trajectory string is determined to be in the rising phase. When the segment water level continuously falls and the drop is significant, and the trend stability is higher than a set threshold, it is marked as the falling phase. When the segment water level fluctuates frequently but the amplitude is weak and no significant rising or falling trend is formed, it is marked as the oscillation phase. For example, trajectory string TS001 contains segments F0111 to F0117, where the segment water level continuously rises from 2.35 meters to 3.12 meters, with an average rise rate of 0.0023 meters per second, which meets the conditions for the rising phase. This trend phase serves as the input for the next compression strategy adjustment.

[0038] Based on the trend phase of the trajectory string, a corresponding compression threshold adjustment strategy is set. When the trajectory string is in the rising phase, to ensure that high-frequency abrupt changes are not lost during compression, the compression threshold is widened, i.e., the granularity of archived data is increased, the time span of the compression sliding window is extended, the sampling frequency is reduced, and data integrity is improved. For example, under normal archiving conditions, data points are sampled every 10 minutes. When the trajectory string is identified as being in the rising phase, the sampling frequency is changed to once every 30 minutes. At the same time, noise reduction and redundant point merging operations are disabled during this time period, and all sampled points are recorded at the original granularity. When the trajectory string is in the falling phase, to reduce the storage overhead of redundant stable data, the system tightens the compression threshold, adjusts the sliding window to 5 minutes, prioritizes the retention of key turning points, water level extreme points, and points of accelerated change, and compresses or merges other duplicate data. During the oscillation phase, a neutral compression strategy is maintained, the default sampling frequency is kept, and all extreme points within the fluctuation range are retained to ensure that the potential for abrupt changes is preserved without significantly increasing the data volume.

[0039] The compression threshold adjustment strategy is mapped to the write process of the data archiving channel in real time. During the archiving process, the data write of each segment will refer to the trend stage of its corresponding trajectory string and the compression control status. For example, segment F0113 is located in the rising phase of trajectory string TS001, and its compression threshold is set to the most lenient state. The corresponding archiving strategy is to retain all samples taken per second and store them in the designated path D: / Corridor / Level1 / TS001 / F0113.dat in the uncompressed corridor. Segment F0124, located in the falling phase of trajectory string TS002, has its compression strategy set to a high compression rate state, retaining only one extreme point every five minutes, and is stored in the path D: / Archive / Level3 / TS002 / F0124.dat after compression. This kind of compression adjustment takes effect dynamically in real time according to the trend of the trajectory string, forming a compression decision basis at the archiving point, thereby ensuring that data compression and mutation evolution process are closely linked, improving the accuracy of data protection and the adaptability of compression strategies.

[0040] Periodic evaluations of the compression control effect based on the tidal threshold balance mechanism analyze compression efficiency, data fidelity, and abrupt fragment reconstruction capabilities under various trend phases. The next round of compression threshold adjustment parameters is then fine-tuned based on the evaluation results. Evaluation content includes: the percentage change in total data volume before and after compression, the incompleteness rate of abrupt fragments in the data, and the model's accuracy in identifying the compressed data. For example, in a heavy rainfall event, trajectory string TS004 used a 30-minute sampling strategy during the rising phase. After compression, the data volume decreased by only 14%, but the abrupt fragment detail integrity rate reached 99.2%, indicating that relaxing the compression threshold effectively preserved key information. In contrast, trajectory string TS005 achieved a compression rate of 68% during the falling phase, but the model reconstruction rate decreased significantly, suggesting that the compression threshold should be slightly increased during this phase to avoid excessive loss of potential information. All evaluation results are fed back to the compression strategy control mechanism to continuously optimize the dynamic balance between compression and fidelity.

[0041] For each triggered trajectory string, a compression strategy execution archive is created, and its compression threshold change history, trend phase switching time points, compression effect indicators, and archiving path numbers are uniformly summarized to form a compression behavior tracking index table. This index table is used to track the compression strategy changes of the trajectory string during the archiving process over a long period of time, ensuring a complete and verifiable data path for future event backtracking, data recovery, model calibration, or strategy reconstruction. For example, the compression behavior index table records: trajectory string number TS001, start time November 2, 2025, 10:15 AM, end time 10:23:45 AM, containing 7 segments. The first 4 segments are in the rising phase, using a relaxed compression strategy with a compression rate of less than 10%, and the last 3 segments are in the falling phase with a compression rate of 65%. The archiving paths are D: / Corridor / Level1 / TS001 and D: / Archive / Level2 / TS001, respectively. This record file not only has auditing functions but also provides important support for subsequent model interpretation and data evaluation.

[0042] The following section will use a water conservancy scenario as an example to explain the technical solution of this invention, compare the effects of existing technologies with specific data, and finally summarize the key indicators in a table.

[0043] I. Typical Scenario Background: A Heavy Rainfall Event at a Mountain Reservoir Suppose a mountain reservoir has a normal pre-flood water level of approximately 220.00 meters, a drainage area of ​​1500 square kilometers, and 20 rain gauge stations, 8 water level stations, and 4 flow stations deployed upstream, with a monitoring frequency of once per second. Traditional water conservancy data centers, in order to save space, typically follow this approach: Real-time data is temporarily stored for 7 days, maintaining a 1-second resolution. Seven days later, the data was averaged and extreme values ​​were compressed in 10-minute intervals, retaining only the maximum, minimum, and average values ​​for each 10-minute interval; After 30 days, a second compression is performed, lasting 1 hour, retaining only the average hourly water level and flow rate. This approach works well under stable operating conditions, but it is very unfavorable for sudden flood peaks.

[0044] On the night of July 15th, from 20:00 to 23:00, a short-duration heavy rainfall occurred in the mountainous area. The actual measured data is as follows (simplified to a segment from an upstream water level station): 20:10: Water level 220.35 meters; 20:20: Water level 221.10 meters; 20:25: A short-term rainstorm passed through, and the water level rose from 221.10 meters to 222.80 meters within 5 minutes; 20:30: Water level 222.80 meters; 20:40: The water level continues to rise to 223.80 meters; 21:10: The water level reached a local flood peak of 224.60 meters; It gradually declined after 21:30; There are two key characteristics: The sharp rise in the 5 minutes around 20:25 (from 221.10 to 222.80 meters) is a typical short-term surge. The continuous rise from 20:40 to 21:10 constitutes the dominant process in the formation of the flood peak; Under traditional compression strategies: After 7 days, the data from this hour will be compressed into several 10-minute intervals, for example: The average water level between 20:20 and 20:30 may be reduced to 222.10 meters; The average water level between 20:30 and 20:40 may be reduced to 223.30 meters; In another month, this data will only be the hourly average from 8 PM to 9 PM, for example, 222.90 meters; The result is that the crucial process of a sharp rise of 1.7 meters within 5 minutes was almost completely erased in the historical data, and the model can only see "a slight increase on average over this hour", completely losing the high-resolution samples that are sensitive to extreme fluctuations.

[0045] II. The complete processing procedure of the present invention in this scenario: 1. Pin identification: Identify the small section that suddenly bulges. In the present invention, the water conservancy data center first traverses the entire 1-second water level sequence from 20:00 to 23:00, analyzes the difference between each second and the previous second, and accumulates the change within a 5-minute sliding time window.

[0046] When the system detects that the water level rises from 221.10 meters to 222.80 meters within 5 minutes around 20:25, with a total fluctuation of 1.70 meters and a rise of more than 0.30 meters per minute, these moments are marked as a series of sharp peak points, and the short period from 20:23 to 20:28 is marked as a "sudden rise point band".

[0047] In the same round of testing, the system also identified a long rise in water level from 20:30 to 21:10, from 222.80 meters to 224.60 meters, as another "slow and continuous rising stitch band".

[0048] The result of this step is that the 5-minute spike that would have been treated as ordinary data in the past has become an "abnormal band with a clear start and end time".

[0049] 2. Spark Fragment Extraction: High-resolution "sparks" are extracted before and after the pin band. With the stitch tape, this invention extends the time by 3 minutes before and after each stitch tape, forming a key extraction window: The sudden increase in stitch length from 20:23 to 20:28 is expanded to 20:20 to 20:31. The sustained upward band from 20:30 to 21:10 is expanded to 20:27 to 21:13; Within each extraction window, the method divides the raw 1-second-level data into extremely short segments of 30 seconds each, for example: Segment P1: 20:20:00–20:20:30, average water level 220.40 meters; Segment P2: 20:20:30–20:21:00, average water level 220.45 meters; ... Excerpt P13: 20:23:00–20:23:30, average water level 221.00 meters; Excerpt P17: 20:25:00–20:25:30, average water level 222.10 meters, maximum water level 222.35 meters; Excerpt P18: 20:25:30–20:26:00, average water level 222.70 meters, maximum water level 222.80 meters; All these 30-second clips were pieced together in chronological order into a "spark clip pool," in which each clip retains the original 1-second level of detail.

[0050] 3. Fragment Passport and Uncompressed Corridor: Registering and protecting key fragments separately. The invention then generates a segment passport for each 30-second segment. For example, the following information is generated for segment P17: Segment ID: FP_20250715_0017; Start time: 20:25:00; End time: 20:25:30; Maximum water level: 222.35 meters; Minimum water level: 221.95 meters; Total amplitude within this segment: 0.40 meters; Trend type: Clearly upward; Is it located inside the pin tape: Yes; Spark Fragment Pool: Pool ID SPK_20250715_A; Then, this passport is "pasted back" to the original data. That is, every second-level data point in the original data from 20:25:00 to 20:25:30 has the field "belonging segment number FP_20250715_0017".

[0051] During the archiving stage, segments with passports do not follow the normal compression path; instead, they are sent to a special uncompressed corridor reserved for high-value segments, for example: Storage path: D: / Corridor / 2025-07-15 / SPK_20250715_A / FP_20250715_0017.dat Internally, each record is fully saved based on the original 1-second sampling, totaling 30 records; At the same time, the fragment passport records its hierarchical archiving level. For example, if a fragment is in an extreme surge at the beginning of a segment, it will be marked as a Level 1 protected fragment, retained for 5 years, and kept in its original precision without any averaging or resampling.

[0052] 4. Triggering trajectory strings and precursor chains: stringing together scattered fragments into a "flood arrival route". Next, this invention, without compressing the corridor, sequentially strings together all passport-carrying segments within the hour-plus period of 20:20–21:13. For example: A total of 10 sudden spikes were identified between 20:23 and 20:28. 20:28–20:40 15 “continuously rising” segments were identified; 20:40–21:10 30 “slowly rising” segments were identified; These segments, arranged chronologically, form a trigger trajectory string TS_20250715_A, which gradually evolves from a "slight rise" to a "sudden surge" and then transitions to a "slow rise at a high level".

[0053] The system calculates the overall metrics for this trajectory sequence: Coverage time: 20:20–21:10, 50 minutes in total; The total water level rose by 4.25 meters during this period (from 220.35 meters to 224.60 meters). Short-term maximum rate of ascent: 0.80 meters / 5 minutes; Number of burst-level segments: approximately 18; This trajectory string is marked as a flood precursor guidance chain, and subsequent compression control and scheduling models will be adjusted based on this guidance chain.

[0054] 5. Tidal threshold balance wheel driven compression strategy: "relaxed" during rising phase, "tightened" during falling phase. Once the current indicator chain TS_20250715_A is identified as a "flood arrival chain," this invention activates the tidal threshold balance wheel: During the rising phase from 20:20 to 21:10, the compression threshold is significantly relaxed; Normally, only one point is retained every 10 minutes in the regular area; Within the phase increase region, change the setting to "save all segments on the critical link without compression" and retain all 1-second data in the associated spark segment pool for a long time. After 21:30, the water level dropped significantly, and the system determined that it had entered the falling phase. For the slowly declining data segment after 21:30, the compression threshold is gradually tightened; You can retain only the extreme values ​​and average values ​​every 5 minutes, and use this part as the low-risk slow-release phase. The result is that the data from the truly critical "surge phase" and "concentrated rise phase before the flood peak" are almost unaffected by compression, while the subsequent long and slow decline phase does not occupy a large amount of storage space.

[0055] III. Comparison of Quantitative Effects with Existing Technologies To demonstrate the practical effect of this invention, it is assumed that a reservoir experienced six similar short-duration heavy rainfall events within a complete flood season (approximately 3 months). The results of the flood peak prediction models trained using two different historical data processing methods are compared below.

[0056] 1. Model performance under existing compression schemes Because rapid spikes within 5 minutes are eliminated by 10-minute averaging and smoothing, high-resolution extreme samples are severely lacking. The trained model performed as follows when reproducing six historical flood peaks: The average peak flow prediction error is approximately 22%. The average peak occurrence time prediction lags by approximately 15 minutes. In two of the six incidents, the flood peak was underestimated by more than 30%, resulting in the dispatch strategy showing "untimely discharge" in the simulation evaluation.

[0057] 2. Model performance after applying the technical solution of this invention In the present invention: During each emergency, the critical 30-second segment was captured at a 1-second sampling interval. For these six events, a total of six flood precursor guidance chains were formed, and a total of about 580 high-resolution passport segments were preserved, each segment being 30 seconds long, for a total of more than 17,000 second-level samples. After retraining the model based on these high-resolution samples: The average peak flow prediction error has been reduced to approximately 5%. The average lead time for peak occurrence has been improved from a 15-minute lag to an average of 20 minutes. In all six incidents, there were no further instances of severe underestimation exceeding 15%. From the perspective of flood control and dispatch: In the traditional approach, there were three instances during the flood season drills where the gate opening command was delayed, with the delay time ranging from 10 to 20 minutes. In the new scheme, the proportion of early scheduling suggestions issued has been significantly increased by utilizing early warning chains and more sensitive model predictions, and there is no longer a significant lag in gate opening commands during simulation exercises.

[0058] 3. Overall balance between storage overhead and revenue Within the same three-month flood season, taking the entire water conservancy data center as a unit: In the traditional approach, compressed historical data would occupy approximately 10TB of storage space. Under the proposed solution, normal stable operating conditions are still handled using the original compression strategy, with only the critical guide chain regions from the six heavy rainfall events remaining uncompressed. The overall storage space is approximately 12TB, an increase of about 20%. but: The number of high-resolution extreme event samples has increased by more than 8 times; The ability to identify and warn of extreme events has been greatly improved; The model's reliability in extreme scenarios is significantly improved, effectively avoiding the problem of not being able to see the most dangerous scenarios due to a lack of abnormal samples.

[0059] IV. Key Indicator Comparison Table

[0060] As can be seen from the table above, the present invention increases the number of high-resolution samples related to extreme events by about 8 times while only increasing the total storage overhead by about 20%, and significantly reduces the flood peak prediction error and scheduling lag risk. This fully demonstrates the substantial beneficial effects of the present invention's technical solution compared to the prior art in "seeing the critical moment clearly, preserving the key details, and improving the reliability of decision-making".

[0061] This invention constructs a spike-pin band to locate sudden fluctuations in signals, extracts key, extremely short segments, and assigns them unique identification passports. During data archiving, these segments are guided into an uncompressible corridor, enabling independent identification and high-fidelity preservation of abrupt fluctuations. Simultaneously, by constructing a trigger trajectory string and activating a tidal threshold oscillator, the data compression threshold is dynamically adjusted according to trends, ensuring a flexible balance between data compression and anomaly retention. The overall solution significantly improves the water conservancy data center's memory and early warning sensitivity to extreme events without significantly increasing overall storage pressure, providing more complete and timely high-value data support for subsequent flood control scheduling, water resource optimization and allocation, and intelligent model training.

[0062] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A big data governance and analysis method applied to a water conservancy data center, characterized in that, Includes the following steps: Peak detection and marking are performed on signals with sudden increases and decreases along the timeline. The marking results are connected according to the time adjacency to form a spike pin band. Based on the spike pin band, the original flow data is extracted within a predetermined time period before and after each spike pin band to obtain a very short segment, and the extracted very short segments are spliced ​​together in chronological order to form a spark segment pool. Generate unique identification information for each fragment in the spark fragment pool and form a fragment passport. Map the fragment passport back to the original flow data to identify the identity of the corresponding fluctuation segment. Uncompressed corridors are opened in the archiving path for data with fragment passports, and the data with fragment passports are stored in the uncompressed corridors according to the hierarchical archiving strategy to maintain the integrity of mutation details; By linking the corresponding segments of the passport in chronological order without compressing the corridor, a continuous trigger trajectory is formed as a precursory guidance chain for extreme events. The tidal threshold balance wheel is activated based on the trigger trajectory string to dynamically adjust the data compression threshold: the compression threshold is relaxed when the trigger trajectory string is in the rising phase and tightened when the trigger trajectory string is in the falling phase, thereby guiding the archiving path to complete adaptive control and achieve long-term fidelity retention of extreme signals.

2. The big data governance and analysis method applied to a water conservancy data center according to claim 1, characterized in that, The steps for forming the spiked pin band are as follows: The original water flow data is read point by point in chronological order, the water level difference between consecutive time points is calculated, and sudden rise or fall signal points are marked according to the condition that the water level rises or falls continuously within a preset time interval and exceeds a predetermined range. Using each mutation signal point as the center, water level data is extracted forward and backward as the confirmation interval to determine whether there is a continuous mutation trend, and mutation peak points with complete trends are selected. The effective mutation peaks are arranged in chronological order, and adjacent peaks with time intervals not exceeding a preset range are grouped together to form a continuous spike pin band. For each spike pin band, calculate the maximum mutation amplitude, average mutation rate, total duration, and minimum peak interval, and record its start and end times, time interval, and internal peak point number.

3. The big data governance and analysis method applied to a water conservancy data center according to claim 2, characterized in that, The steps for constructing a spark fragment pool are as follows: The data is extended forward by a predetermined time based on the start time of each spike pin band, and backward by a predetermined time based on the end time, to form a continuous data interval. Each data interval is divided into multiple extremely short time segments with fixed lengths, continuous start and end times, and no overlap between segments. Complete original pipeline data is retained within each extremely short time segment. All extremely short time segments are pieced together in chronological order to form a continuous data set covering an extended time interval, thus creating a spark fragment pool; Feature statistics are performed on all extremely short time segments in the spark fragment pool to generate structured descriptive information.

4. The big data governance and analysis method applied to a water conservancy data center according to claim 3, characterized in that, The steps to create a fragment passport are as follows: Traverse each extremely short fragment in the spark fragment pool and extract the time identification field and hydrological statistical feature field to construct an identification field set; A structured fragment passport is generated for each very short fragment based on the set of identification fields; Map each segment passport back to the original stream data, and append the segment number and the position marker field within the segment to the original data fields; A unified passport index table is established for all passport fragments, which are then categorized, sorted, and have their core identification information stored according to the spark fragment pool number.

5. The big data governance and analysis method applied to a water conservancy data center according to claim 4, characterized in that, The process of opening a corridor without compression is as follows: Extract the extremely short data segments with attached passport identifiers as the target dataset for protection and then merge and strip out the regular data stream; In the data archiving structure, uncompressible corridors are defined for the protected target data set, and dedicated storage paths are set according to time, event level, and source pool number; Implement a tiered archiving strategy based on the intensity of mutation characteristics, data scarcity, and correspondence with historical events, and configure storage protection measures corresponding to the archiving level. Establish a data mapping table to record the storage path, number of backups, retention period, and subsequent application status of each data segment, and realize controllable retrieval of data.

6. The big data governance and analysis method applied to a water conservancy data center according to claim 5, characterized in that, The hierarchical archiving strategy classifies data segments with high mutation intensity, high data scarcity, and correlation with historical emergencies as core mutation fragments. It adopts a dual-copy redundant storage method, storing them in different physical storage paths, and periodically performs integrity verification operations.

7. The big data governance and analysis method applied to a water conservancy data center according to claim 5, characterized in that, The process of forming a continuous trigger trajectory string is as follows: Extract data segments containing fragmented passports from the uncompressed corridor, construct a passport time list sorted in ascending order of start time, and build an initial index structure to trigger trajectory construction; Based on the passport time list, determine whether adjacent segments meet the concatenation condition and filter out segments that can be continuously constructed into trajectories; Connect the segments that meet the concatenation conditions in sequence to construct the basic trigger trajectory segment and record the time range and trend characteristics of the trajectory segment; The basic trigger trajectory segments are judged for time continuity and trend compatibility and aggregated into trigger trajectory strings; Generate a precursor guide chain identifier for the trigger trajectory string and establish its mapping relationship in the original data.

8. The big data governance and analysis method applied to a water conservancy data center according to claim 1, characterized in that, The data compression threshold is dynamically adjusted based on the trigger trajectory string to activate the tidal threshold balance wheel: the compression threshold is relaxed when the trajectory string is in the rising phase and tightened when the trajectory string is in the falling phase, so as to drive the archive path to perform adaptive control and complete the long-term fidelity storage of extreme signals. The steps are as follows: Based on the trigger trajectory string, all segment passports are extracted and the trend phase attributes of the trajectory string, namely the rising phase, falling phase, and oscillation phase, are identified in chronological order. The compression threshold adjustment strategy is set according to the trend phase of the trajectory string. The compression threshold is relaxed in the rising phase, tightened in the falling phase, and a neutral compression threshold is maintained in the oscillation phase. Map the compression threshold adjustment strategy to the data archiving channel and write the compression strategy corresponding to the execution trend stage into the data of each segment; The compression control results of the tidal threshold balance wheel mechanism are periodically evaluated, and the compression threshold adjustment parameters are fine-tuned based on the evaluation results. Create a compression strategy execution file for each trigger trajectory string and record the history of compression threshold changes, trend stage switching time points, compression effect indicators, and archive path numbers.