Data processing method of seabed-based monitoring equipment
Through a multi-stage data processing process, including singular value decomposition, tensor decomposition and deep-sea environmental perception attention network model, the problem of difficult to distinguish environmental changes and equipment drift in seabed monitoring data in traditional technology is solved, and the accuracy and reliability of monitoring data are significantly improved.
Patent Information
- Application Number
- CN202510668123.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Traditional seabed-based monitoring equipment is difficult to effectively distinguish data drift from environmental changes in long-term work, resulting in the impact of monitoring accuracy and reliability.
A multi-stage data processing process is adopted, including normalization processing, singular value decomposition, sliding time domain window segmentation, CUDA parallel processing flow, tensor decomposition and deep-sea environment perceived attention network model, to build a seabed environment state tensor model and establish a data drift correction model.
Effectively distinguishing environmental changes from equipment drift, improving the accuracy and reliability of long-term monitoring data, and supporting stable monitoring of seabed environments.
Smart Images

Figure CN120217209A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electrical digital data processing, and more particularly, relates to a data processing method for seabed-based monitoring devices. Background Art
[0002] Seabed-based monitoring devices are widely used in fields such as ocean science research, seabed resource exploration, and marine environmental protection. Traditional seabed-based monitoring technologies mainly rely on multi-sensor arrays to collect seabed environmental parameters, obtain raw data through data acquisition units, and then perform basic processing such as noise reduction and filtering through signal processing modules. These devices usually use simple outlier detection and statistical filtering methods to process data and can provide reliable monitoring results in relatively stable environments.
[0003] However, traditional technologies face serious challenges in long-term seabed monitoring. Due to the complex and variable seabed environment, sensors will age and drift during long-term operation, resulting in a gradual shift of the measurement baseline. At the same time, it is difficult to distinguish the data anomalies caused by natural changes such as turbulence and tides in the seabed environment from sensor drift. Existing data processing methods often misjudge natural environmental changes as equipment failures or misinterpret sensor drift as environmental anomalies, seriously affecting monitoring accuracy and reliability.
[0004] Even more intractable is that traditional technologies lack effective means to simultaneously process the spatio-temporal correlation and long-term drift of multi-source data, especially in the seabed environment with limited computing resources, where real-time and efficient processing of a large amount of monitoring data cannot be achieved. The environmental changes and equipment drift in seabed data are intertwined to form a complex mixed pattern, and traditional single processing methods are difficult to solve this core technical problem, seriously restricting the long-term stable application of seabed monitoring technologies. That is to say, there is a technical problem in the prior art that it is difficult to effectively distinguish the data drift and environmental change characteristics during the long-term operation of seabed-based monitoring devices. Summary of the Invention
[0005] In view of this, the present invention provides a data processing method for seabed-based monitoring devices, which can solve the technical problem in the prior art that it is difficult to effectively distinguish the data drift and environmental change characteristics during the long-term operation of seabed-based monitoring devices.
[0006] The present invention is implemented as follows: The present invention provides a data processing method for seabed-based monitoring equipment, including: performing normalization processing on the data collected by the seabed-based monitoring equipment and eliminating abnormal data to determine an effective data set; using the singular value decomposition method to perform dimensionality reduction processing on the effective data set, extracting key features and generating a feature matrix; implementing a sliding time-domain window segmentation on the feature matrix to form multiple groups of time series data blocks; starting a CUDA parallel processing stream to construct a disorder matrix and a drift matrix; extracting disorder pattern feature vectors and drift principal components in the CUDA parallel processing stream; extracting abnormal data features in the CPU processing stream to generate an abnormal feature matrix; using tensor decomposition technology to fuse the disorder pattern feature vectors, drift principal components, and abnormal feature matrix to construct a seabed environmental state tensor model; constructing a data drift correction model based on a residual neural network to correct and compensate real-time monitoring data; and using a hierarchical clustering algorithm to classify the processed data to establish a seabed state evaluation criterion.
[0007] Among them, the singular value decomposition method specifically organizes the seabed monitoring data into a matrix form, and by calculating eigenvalues and eigenvectors, decomposes the data matrix into the product of three sub-matrices, retains the eigenvectors corresponding to the larger singular values, and discards noise and redundant information.
[0008] Among them, the disorder matrix specifically refers to a feature matrix constructed by calculating the volatility, uncertainty, and disorder degree of seabed monitoring data in the time and space dimensions, and is used to quantitatively describe the turbulence, perturbation, and nonlinear dynamic processes in the seabed environment.
[0009] Among them, the drift matrix specifically refers to a matrix constructed by comparing the deviation degree between the real-time readings of sensors and the historical reference values. Each element in the matrix represents the drift amount of the corresponding sensor at the corresponding time point, and is used to track and analyze the performance changes of sensors and the long-term evolution trend of the environment.
[0010] Among them, an adaptive spectrum optimization function is used in the process of constructing the disorder matrix, which is dynamically adjusted according to the frequency characteristic differences of the monitoring data under different seabed environmental conditions. By dynamically adjusting the analysis window length and frequency resolution, the accurate capture of data characteristics under different seabed environments is realized.
[0011] Among them, a pre-trained deep-sea environmental perception attention network model is used in the data drift correction process. This model combines the Transformer architecture and the recurrent neural network structure, and introduces a multi-scale drift perception attention mechanism designed for the characteristics of seabed data, which can identify and compensate for complex drift patterns in seabed monitoring data at different time scales.
[0012] Among them, the specific structure of the deep - sea environment perception attention network model is a hybrid architecture that combines a multi - layer bidirectional recurrent neural network and a multi - head self - attention mechanism. The bottom layer uses a convolutional neural network to extract features from multi - source sensor data. The middle layer uses six - layer Transformer encoders to extract temporal features and cross - sensor correlation features. The top layer uses a fully - connected network with skip connections for data drift estimation and compensation.
[0013] Among them, the deep - sea environment perception attention network model introduces a sparse attention mechanism for seabed feature perception in the Transformer encoder. The sparse attention mechanism dynamically adjusts the attention range according to the physical characteristics of the seabed environment, and can simultaneously focus on short - term fluctuations and long - term drift patterns. The entire model adopts a hierarchical design, and processing modules are designed respectively for data characteristics under different depths and environmental conditions.
[0014] Among them, during the pre - training process of the deep - sea environment perception attention network model, long - term monitoring data collected from the global seabed monitoring network is used. Data segments containing complete environmental change cycles and sensor drift phenomena are screened, and a physical oceanography model is used to generate synthetic data to enhance the coverage of the dataset, forming a comprehensive training dataset covering various typical seabed environments such as shallow - sea areas, deep - sea plains, trench areas, and hydrothermal areas.
[0015] Among them, the seabed - based monitoring device mainly consists of a sensor array system, a data acquisition unit, a signal processing module, an energy supply system, a communication transmission module, and a protective shell. The sensor array system includes a pressure sensor, a temperature sensor, a flow velocity sensor, a seismic sensor, and a chemical substance detector.
[0016] Compared with the prior art, the present invention provides a data - processing method for a seabed - based monitoring device. The present invention proposes a data - processing method for a seabed - based monitoring device. Through a multi - stage data - processing flow, it effectively solves the problem that it is difficult to distinguish between environmental changes and device drift during long - term monitoring. This method first normalizes the original data and eliminates outliers, and uses singular value decomposition to extract key features; then uses an adaptive sliding window to segment the time series, constructs a disorder matrix and a drift matrix; finally, through tensor decomposition and the deep - sea environment perception attention network model, it accurately distinguishes between environmental changes and device drift.
[0017] The present invention effectively overcomes the limitations of traditional technologies. Through the coordinated work of the CUDA parallel processing stream and the CPU processing stream, it realizes the efficient processing of large - scale seabed monitoring data; the multi - dimensional tensor model successfully captures the high - order correlations between different sensor data; the pre - trained deep - sea environment perception attention network can identify and compensate for complex drift patterns on multiple time scales, effectively distinguishing between natural environmental changes and device drift.
[0018] Through the organic combination of the above technical means, the present invention successfully solves the core technical problem of difficult distinction between environmental changes and equipment drift in seabed-based monitoring data, significantly improves the accuracy and reliability of long-term monitoring data, provides a strong guarantee for the long-term stable monitoring of the seabed environment, and has important practical value for marine scientific research and seabed resource development. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flowchart of the method of the present invention.
[0020] Figure 2 It is a schematic diagram of the composition of the seabed-based monitoring equipment in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0022] As Figure 1 shown, it is a flowchart of a data processing method for a seabed-based monitoring equipment provided by the present invention, and this method includes the following steps: S01. Perform normalization processing on the data collected by the seabed-based monitoring equipment, and use an outlier detection algorithm to identify and remove abnormal data outside a predetermined threshold range to determine a valid data set; S02. Based on the valid data set, use the singular value decomposition method to perform dimensionality reduction processing on the original data matrix, eliminate data redundancy, extract key features and generate a feature matrix; S03. Implement sliding time-domain window segmentation for the feature matrix, adaptively adjust the window width according to the characteristics of seabed environmental changes, and form multiple groups of time series data blocks; S04. Start a CUDA parallel processing stream to receive the time series data blocks, calculate the disorder degree index within each data block, and construct a disorder matrix; at the same time, count the degree of deviation of the sensor readings from the reference value and construct a drift matrix; S05. Perform singular value decomposition on the disorder matrix in the CUDA parallel processing stream to extract disorder mode feature vectors; perform principal component analysis on the drift matrix to extract drift principal components; S06. Extract abnormal data features in the CPU processing stream to generate an abnormal feature matrix, and align the abnormal feature matrix with the disorder mode feature vectors and the drift principal components; S07. Use tensor decomposition technology to fuse the disorder mode feature vectors, the drift principal components and the abnormal feature matrix, construct a seabed environmental state tensor model, and extract the coupling relationship between environmental factors and monitoring data; S08. Build a data drift correction model based on a residual neural network. Input the drift principal components, train the network to learn the drift patterns in the long time series, and correct and compensate the real-time monitoring data. S09. Use a hierarchical clustering algorithm to classify the processed data, establish a seabed state evaluation criterion, and generate a multi-level seabed state evaluation result. Among them, the singular value decomposition method specifically organizes the seabed monitoring data into a matrix form, decomposes the data matrix into the product of three sub-matrices by calculating eigenvalues and eigenvectors, retains the eigenvectors corresponding to the larger singular values, and discards the noise and redundant information.
[0023] Among them, the sliding time-domain window segmentation specifically divides the continuous time series data into a series of overlapping data segments according to a preset time length, and there is a certain proportion of overlap between adjacent windows, which is used to capture the time correlation and gradual change characteristics in the data.
[0024] Among them, the tensor decomposition technology specifically represents the multi-dimensional data as a high-order tensor, and decomposes the tensor into the outer product of multiple low-dimensional factors through multilinear algebra operations, which can retain the high-order correlation between data and discover potential data structures.
[0025] Among them, the CUDA parallel processing stream specifically refers to a parallel computing process implemented by using the Compute Unified Device Architecture technology on the graphics processing unit. By distributing large-scale matrix operations to thousands of computing cores to execute simultaneously, efficient and intensive computing is achieved.
[0026] Among them, data drift specifically refers to the phenomenon of slow deviation of the measurement reference line of seabed-based monitoring equipment during long-term operation, which is caused by sensor aging, environmental changes, or the characteristics of the equipment itself, and affects the data accuracy.
[0027] Among them, the disorder matrix specifically refers to a feature matrix constructed by calculating the volatility, uncertainty, and disorder degree of seabed monitoring data in the time and space dimensions, which is used to quantitatively describe the turbulence, perturbation, and nonlinear dynamic processes in the seabed environment.
[0028] Among them, the drift matrix specifically refers to a matrix constructed by comparing the deviation degree between the real-time readings of the sensor and the historical reference value. Each element in the matrix represents the drift amount of the corresponding sensor at the corresponding time point, which is used to track and analyze the performance change of the sensor and the long-term evolution trend of the environment.
[0029] The seabed-based monitoring device mainly consists of a sensor array system, a data acquisition unit, a signal processing module, an energy supply system, a communication and transmission module, and a protective housing. The sensor array system includes a pressure sensor, a temperature sensor, a flow velocity sensor, a seismic sensor, and a chemical substance detector, which conduct all-round monitoring of the seabed environment through the collaborative work of multi-source sensors. The data acquisition unit is responsible for sampling and preliminary processing of the signals from each sensor. The signal processing module contains a digital signal processor and an embedded computing unit for executing data processing algorithms. The energy supply system uses a combination of lithium batteries and seawater energy harvesting devices to provide long-term energy for the device. The communication and transmission module is responsible for transmitting the processed data to the sea surface or shore-based stations through acoustic communication or optical cable networks. The entire system is protected by a waterproof and pressure-resistant titanium alloy housing to ensure stable operation in the high-pressure deep-sea environment.
[0030] The adaptive spectrum optimization function is used to optimize the construction process of the disorder matrix in S04, and dynamically adjusts according to the frequency characteristic differences of the monitoring data under different seabed environmental conditions. The inputs include a time series data block, an environmental type identification parameter, historical spectrum baseline data, a signal-to-noise ratio threshold, and a frequency range limit parameter. The output is an optimized disorder matrix, which includes the main frequency components, the distribution of disorder degree indicators, and the frequency drift quantization indicators. This function realizes the accurate capture of data characteristics under different seabed environments by dynamically adjusting the analysis window length and frequency resolution, and is especially suitable for processing seabed monitoring data with long-term slow drift.
[0031] The pre-trained deep-sea environment perception attention network model is used to optimize the data drift correction process in S08. This model integrates the Transformer architecture and the recurrent neural network structure, and introduces a multi-scale drift perception attention mechanism designed for the characteristics of seabed data, which can identify and compensate for complex drift patterns in seabed monitoring data at different time scales. The attention weight allocation parameters in the model need to be determined according to three key parameters: the sensor drift rate of the seabed monitoring device, the environmental change rate, and the data acquisition frequency.
[0032] The specific structure of the deep-sea environment perception attention network model is a hybrid architecture that combines a multi-layer bidirectional recurrent neural network and a multi-head self-attention mechanism. The bottom layer uses a convolutional neural network to extract features from multi-source sensor data. The middle layer uses six layers of Transformer encoders to extract temporal features and cross-sensor correlation features. The top layer uses a fully connected network with skip connections for data drift estimation and compensation. A sparse attention mechanism for seabed characteristics perception is introduced into the Transformer encoder. The sparse attention mechanism dynamically adjusts the attention range according to the physical characteristics of the seabed environment, and can simultaneously focus on short-term fluctuations and long-term drift patterns. The entire model adopts a hierarchical design, and processing modules are designed respectively for data characteristics under different depths and environmental conditions.
[0033] The steps for establishing the training dataset during the pre-training process of the deep-sea environment perception attention network model specifically include collecting long-term monitoring data from the global seabed monitoring network, screening data segments that contain complete environmental change cycles and sensor drift phenomena, performing expert annotation on each data segment to mark the normal environmental change and sensor drift parts, generating synthetic data using physical oceanography models to enhance the dataset coverage, mixing real data and synthetic data in a ratio of 4:1 to construct the training set and validation set, constructing sub-datasets for transfer learning for different seabed environment types respectively, and finally forming a comprehensive training dataset covering various typical seabed environments such as shallow sea areas, deep-sea plains, trench areas, and hydrothermal areas.
[0034] The steps for pre-training the deep-sea environment perception attention network model specifically include first performing initialization training on synthetic data to learn basic temporal patterns, then performing supervised learning on the global dataset to master the general seabed environmental change laws, then performing fine-tuning training on various types of seabed environment data to enhance the model's adaptability in the environment, using contrastive learning methods to train the model to distinguish environmental changes and sensor drift, introducing a curriculum learning strategy to gradually transition from simple patterns to complex drift patterns, using multi-task learning to optimize both the drift detection and compensation tasks simultaneously, and finally using knowledge distillation technology to compress the large model into a lightweight model suitable for deployment on the limited computing resources of seabed monitoring devices. The entire pre-training process is executed in parallel on multiple graphics processing unit servers using a distributed computing framework to ensure that the model can fully learn the complex patterns in long-term temporal data.
[0035] The specific implementation manners of the above steps are described in detail below. The specific implementation manner of step S01 is to first perform normalization processing on the original data collected by the seabed-based monitoring equipment, unify the data with different dimensions to the range of 0 to 1. Specifically, the min-max normalization method is adopted, that is, for each sensor data sequence, the normalized value is calculated as the original value minus the minimum value divided by the difference between the maximum value and the minimum value, where the minimum value and the maximum value are the minimum value and the maximum value of the sensor data sequence respectively. Then, an improved local outlier factor algorithm is used to detect outliers. This algorithm identifies outliers by calculating the ratio of the local density of a sample point to that of its k nearest neighbor sample points. The value of k is adaptively selected according to the scale of the dataset, generally taking 5% to 10% of the total number of samples. The outlier determination threshold is set to 2.5, that is, when the local outlier factor of a certain data point is greater than 2.5, it is determined as an outlier and removed. The purpose of this step is to eliminate noise and outliers in the data, ensure the data quality for subsequent analysis and processing, and improve the reliability and accuracy of data processing.
[0036] The specific implementation manner of step S02 is to organize the effective dataset obtained in step S01 into a matrix form, where the rows represent time points and the columns represent the measured values of different sensors. Then, singular value decomposition is applied to the matrix, and the original matrix is decomposed into the product of three matrices, including two orthogonal matrices and a diagonal matrix. The elements on the diagonal matrix are singular values. According to the magnitudes of the singular values, the first r singular values and the corresponding eigenvectors with a cumulative contribution rate reaching 95% are selected to form the reduced-dimensional feature matrix. In practice, the value of r is usually selected such that the energy of the retained singular values accounts for more than 95% of the total energy. For typical seabed monitoring data, the value of r is generally between 10 and 20. The purpose of this step is to reduce the data dimension, eliminate redundant information and noise in the data, extract the key information reflecting the main characteristics of the seabed environment, and provide a more compact and effective data representation for subsequent processing.
[0037] The specific implementation of step S03 is to perform time-domain segmentation on the feature matrix generated in step S02 using the adaptive sliding window method. First, the basic window width is set to 1 hour, and an adaptive adjustment coefficient is set according to the characteristics of seabed environment changes. This coefficient depends on the environmental change rate and data acquisition frequency, and generally ranges from 0.5 to 2.0. The variance change rate of real-time monitoring data is monitored. When the variance change rate exceeds a preset threshold (usually 20%), the window width is adjusted according to the formula of the basic window width multiplied by one plus the adjustment coefficient multiplied by the change rate, ensuring that the window width is reduced when the environment changes violently and enlarged when the environment is relatively stable. The overlap rate between adjacent windows is set to 50% to ensure data continuity and temporal correlation. Through this adaptive window segmentation method, the feature matrix is divided into multiple groups of time series data blocks, and each data block contains all sensor feature data within a certain period of time. The purpose of this step is to adapt to the dynamic change characteristics of the seabed environment, ensure that subtle change information can be captured when the environment changes violently, and reduce data redundancy and improve processing efficiency when the environment is relatively stable.
[0038] The specific implementation of step S04 is to start the CUDA parallel processing stream and allocate computing resources to process the multiple groups of time series data blocks generated in step S03. For each data block, calculate the disorder degree indicators, including entropy value, Lyapunov exponent, and complexity, etc. Specifically, the sample entropy algorithm is used to calculate the complexity and irregularity of the time series, the embedding dimension is set to 2, and the similarity tolerance is set to 0.2 times the standard deviation. At the same time, calculate the Lyapunov exponent to evaluate the chaotic degree of the sequence, the reconstruction delay is selected as 1 / 4 period, and the reconstruction dimension is determined by the false nearest neighbor method. Organize these disorder degree indicators into a disorder matrix, where each element in the matrix represents the disorder degree indicator of a certain sensor within a certain time window. In parallel, calculate the deviation degree of each sensor reading relative to the historical reference value and construct a drift matrix, where each element represents the drift amount of a certain sensor within a certain time window. The drift amount is calculated using the exponentially weighted moving average method, the smoothing factor is set to 0.1, and the threshold is set to 3 times the standard deviation. This step uses the CUDA parallel computing architecture to distribute large-scale matrix operations to thousands of computing cores for simultaneous execution, significantly improving processing efficiency. The purpose is to quantitatively describe the turbulence, perturbation, and sensor drift in the seabed environment and provide a basis for subsequent analysis.
[0039] The specific implementation of step S05 is to perform singular value decomposition on the disorder matrix constructed in step S04 in the CUDA parallel processing stream, decomposing the matrix into the product of three matrices. Analyze the singular value distribution, and select the right singular vectors corresponding to the first k largest singular values as the disorder pattern feature vectors. The value of k is selected such that the cumulative contribution rate reaches 90%. At the same time, perform principal component analysis on the drift matrix, calculate the covariance matrix as the transpose of the drift matrix multiplied by the drift matrix and then divided by the number of samples minus 1, where the number of samples is the number of time windows. Solve the characteristic equation to obtain the eigenvalues and eigenvectors, and select the principal components with an explained variance of 85% as the drift principal components. This step is executed in parallel on the GPU, making full use of the high-performance computing capabilities of the CUDA architecture. The purpose is to extract the main patterns and features from the disorder matrix and the drift matrix, reduce the data dimension, and extract the key information reflecting the disorder state of the seabed environment and the sensor drift trend.
[0040] The specific implementation of step S06 is to extract features from the abnormal data identified in step S01 in the CPU processing stream. First, group the abnormal data according to time windows, and calculate the statistical features of the abnormal data within each window, including frequency, amplitude, spatial distribution, and time distribution, etc. Use the kernel density estimation method to analyze the distribution characteristics of the abnormal data, and select the bandwidth parameter h as 0.1. Then extract the abnormal pattern features, apply the random forest algorithm to identify the main types and features of the abnormal data, set the number of trees to 100, and the minimum number of samples in the leaf nodes to 5. Organize the extracted abnormal features into an abnormal feature matrix A, and align A with the disorder pattern feature vectors and drift principal components extracted in step S05 in the time dimension through timestamp alignment to generate a unified time index, ensuring the temporal consistency of the three types of feature data. The purpose of this step is to extract the environmental information contained in the abnormal data and align it with the disorder pattern and drift principal components in the time dimension, providing a basis for subsequent data fusion.
[0041] The specific implementation of step S07 is to fuse the disorder pattern feature vectors, drift principal components, and abnormal feature matrix obtained in steps S05 and S06 using tensor decomposition technology. First, construct a third-order tensor X, whose three dimensions correspond to time, sensor type, and feature type (disorder, drift, abnormal) respectively. Then apply Tucker decomposition to decompose the tensor X into the product of a core tensor G and three factor matrices A, B, and C, that is, X≈G×1A×2B×3C, where × Denotes the tensor-matrix product along the n-th dimension. The dimension selection of the core tensor is based on empirical rules, generally taking 30% - 50% of the original size of each dimension. By analyzing the element sizes of the core tensor and their corresponding factor vectors, the coupling relationships between different features are identified, and a seabed environmental state tensor model is constructed. This model can express the complex associations among environmental factors, sensor characteristics, and data anomalies, providing a theoretical basis for multi-dimensional information fusion in seabed state assessment. The purpose of this step is to discover the hidden multi-dimensional association patterns in the data through high-order tensor decomposition and construct a tensor model that comprehensively reflects the seabed environmental state.
[0042] The specific implementation of step S08 is to construct a data drift correction model based on a residual neural network. The network structure includes 1 input layer, 4 residual blocks, and 1 output layer. Each residual block consists of 2 convolutional layers and 1 shortcut connection. The input layer receives the drift principal components extracted in step S05, and the output layer generates drift correction coefficients. The network training uses the Adam optimizer, with the initial learning rate set to 0.001 and adjusted using the cosine annealing strategy. The loss function uses a combination of mean squared error and L1 regularization, with the regularization coefficient set to 0.0001. The training data selects the data segments with known drift patterns in the historical monitoring data and divides them into a training set and a validation set according to a ratio of 8:2. During the training process, when the validation loss no longer decreases for 5 consecutive epochs, the early stopping mechanism is triggered. The trained model is applied to the real-time monitoring data, and correction coefficients are generated according to the identified drift patterns to correct and compensate the data. The purpose of this step is to learn the long-term drift patterns in the seabed monitoring data, establish an effective drift correction model, and improve the accuracy and reliability of long-term monitoring data.
[0043] The specific implementation of step S09 is to classify the data processed in the previous steps using a hierarchical clustering algorithm. First, the distance metric between data points is defined, and the Mahalanobis distance is used to calculate the similarity between samples, considering the covariance structure of the data distribution. Then, the Ward minimum variance method is used as the clustering criterion, which selects the two clusters with the smallest increase in within-class variance for merging at each step. The clustering process starts from individual samples and gradually merges the most similar clusters to finally form a hierarchical structure. The quality of different clustering results is evaluated by calculating the silhouette coefficient and the Davies-Bouldin index to determine the optimal number of clusters, which is generally between 4 and 8. According to physical oceanography knowledge and combined with expert experience, physical meanings are assigned to each cluster, and a seabed state assessment standard is established, which is divided into four levels: normal, slightly abnormal, moderately abnormal, and severely abnormal. Finally, a seabed state assessment report is generated based on the clustering results of the real-time monitoring data. The purpose of this step is to scientifically classify the processed data, establish an objective seabed state assessment standard, and provide decision-making support for seabed environmental monitoring and early warning.
[0044] The structure of the deep - sea environment perception attention network model adopts a hybrid architecture that combines a multi - layer bidirectional recurrent neural network and a multi - head self - attention mechanism. The bottom layer consists of 3 convolutional blocks, each convolutional block contains 2 one - dimensional convolutional layers and 1 max - pooling layer. The convolutional kernel sizes are 3 and 5 respectively, and the number of channels doubles layer by layer starting from 32. The middle layer is composed of 6 layers of Transformer encoders. Each layer contains a multi - head self - attention sub - layer and a feed - forward neural network sub - layer. The number of attention heads is set to 8, and the hidden layer dimension is 512. The top layer is composed of 4 layers of fully - connected networks, and the number of neurons is 256, 128, 64, and 32 in sequence. The activation function uses ReLU, and skip connections are added between the first 3 layers of the top layer to alleviate the problem of gradient disappearance. A sparse attention mechanism for seabed feature perception is introduced in the Transformer encoder. This mechanism dynamically adjusts the attention distribution according to the physical characteristics of the seabed environment, and the attention sparsity parameter is set to 0.7. The entire model adopts a hierarchical design, and processing modules are designed respectively for data characteristics under different depths and environmental conditions. Shallow - sea areas, deep - sea plains, trench areas, and hydrothermal areas correspond to different attention modules, and the module selection is controlled by the environmental type parameter.
[0045] The construction process of the training data set for the deep - sea environment perception attention network model first collects long - term monitoring data from the public data in the global seabed monitoring network, including the monitoring site data of typical regions such as the Mid - Atlantic Ridge, the East Pacific Rise, and the Mariana Trench, with a time span of not less than 3 years. Clean and pre - process the collected original data to remove data segments with obvious errors and a large number of missing values. Then, screen data segments that contain a complete environmental change cycle and sensor drift phenomenon, and the length of the data segment is at least 1 month. Invite more than 5 oceanography experts to annotate each data segment to distinguish between normal environmental changes and sensor drift parts, and the annotation consistency requirement reaches a Cohen's Kappa coefficient of not less than 0.8. Based on the physical oceanography model and the existing annotated data, generate synthetic data that is 10 times the amount of the original data to enhance the coverage of the data set. Mix the real data and the synthetic data in a ratio of 4:1, and divide them into a training set, a validation set, and a test set in a ratio of 7:2:1. Construct sub - data sets for typical seabed environments such as shallow - sea areas (water depth < 200 meters), deep - sea plains (water depth 200 - 4000 meters), trench areas (water depth > 4000 meters), and hydrothermal areas. Each sub - data set contains at least 100 valid data segments. Finally, form a comprehensive training data set covering a variety of typical seabed environments, with a total data volume of not less than 10TB, ensuring that the model can learn the data characteristics and drift laws under different seabed environments.
[0046] The following details the mathematical models or calculation processes involved in the present invention.
[0047] In step S01, the normalization processing formula for the data collected by the seabed-based monitoring device is expressed as follows: ; Wherein, is the normalized data value; is the original data value; is the minimum value of the sensor data sequence; is the maximum value of the sensor data sequence.
[0048] The outlier detection uses the Local Outlier Factor (LOF) algorithm, and its calculation formula is: ; Wherein, is the local outlier factor of point ; is the nearest neighbor point set of point ; is the local reachability density of point ; is the number of nearest neighbor points.
[0049] The calculation formula for the local reachability density is: ; Wherein, is the reachability distance from point to point , which is defined as: ; Wherein, is the distance from point to its th nearest neighbor point; is the Euclidean distance between point and point .
[0050] In step S02, the matrix representation of Singular Value Decomposition (SVD) is: ; Wherein, is the original data matrix, is the number of time points, is the number of sensors; and are orthogonal matrices respectively; is a diagonal matrix, and the elements on the diagonal are singular values.
[0051] The calculation formula for the feature matrix after dimensionality reduction is: ; In the formula, includes the first columns of ; includes the diagonal matrix composed of the first singular values of ; includes the first columns of . The selection of satisfies: In the formula, is the th singular value.
[0052] In step S03, the calculation formula for the adaptive sliding window width is: ; In the formula, is the actual window width; is the basic window width, set to 1 hour; is the adjustment coefficient, and its value range is 0.5 to 2.0; is the data variance change rate, and its calculation formula is: ; In the formula, is the data variance at the current moment; is the data variance at the previous moment.
[0053] In step S04, the calculation of the disorder degree index includes the calculation formula for sample entropy: ; In the formula, is the sample entropy value; is the embedding dimension, and its value is 2; is the similarity tolerance, and its value is 0.2 times the data standard deviation; is the data length; is the number of pattern pairs matched in dimension ; is the number of pattern pairs matched in dimension .
[0054] The calculation process of the Lyapunov exponent first requires phase space reconstruction: ; In the formula, is the reconstructed state vector; is the th point of the time series; The reconstruction delay is 1 / 4 of the period; The reconstruction dimension is determined by the false nearest neighbor method.
[0055] The calculation formula of the Lyapunov exponent is: ; In the formula, is the Lyapunov exponent; is the th time point; is the distance between adjacent orbits in the phase space at time
[0056] The drift is calculated using the exponentially weighted moving average method, and the formula is: ; In the formula, is the smoothed value at time is the observed value at time is the smoothing factor, with a value of 0.1; is the smoothed value at time
[0057] The drift is defined as: ; In the formula, is the drift at time is the observed value at time is the smoothed value at time
[0058] In step S05, perform singular value decomposition on the disorder matrix : ; In the formula, is the disorder matrix, is the number of time windows, is the number of sensors; and are orthogonal matrices respectively; is a diagonal matrix, and the elements on the diagonal are singular values.
[0059] The selection of the disorder mode eigenvector satisfies: ; In the formula, is the th singular value; is the number of selected eigenvectors.
[0060] Perform principal component analysis on the drift matrix The covariance matrix calculation formula is: ; In the formula, is the covariance matrix; is the drift matrix, is the number of samples, is the number of sensors.
[0061] The selection of the principal components satisfies: ; In the formula, is the covariance matrix of the th eigenvalue; is the number of selected principal components.
[0062] In step S06, kernel density estimation is used to analyze the distribution characteristics of abnormal data, and its calculation formula is: ; In the formula, is the kernel density estimate value; is the number of samples; is the bandwidth parameter, with a value of 0.1; is the kernel function, usually the Gaussian kernel, expressed as: ; In the formula, is the standardized distance.
[0063] In addition, in step S06, the random forest algorithm is used to identify the main types and characteristics of abnormal data, and its Gini impurity calculation formula is: ; In the formula, is the Gini impurity of the dataset ; is the number of classes; is the proportion of samples of the th class in the dataset.
[0064] The calculation formula for information gain is: ; In the formula, is the information gain of the feature ; is the number of values of the feature ; is the sample subset where the feature takes the value of ; and are the number of samples in the subset and the total dataset, respectively.
[0065] In step S07, the tensor decomposition technique applies Tucker decomposition, expressed as: ; In the formula, is the original third-order tensor, is the size of the time dimension, is the size of the sensor type dimension, is the size of the feature type dimension; is the core tensor, , , ; , , are the factor matrices; represents the tensor-matrix product along the th dimension.
[0066] The formula for dimension selection of the core tensor is: ; ; ; In the formula, represents rounding up; , , are adjustment parameters, and their value range is 0 to 0.2 times the original dimension size.
[0067] In step S08, the loss function of the residual neural network is: ; In the formula, is the total loss; is the mean squared error, and its calculation formula is: ; In the formula, is the number of samples; is the true value; is the predicted value.
[0068] is the L1 regularization term, and its calculation formula is: ; In the formula, is the number of model parameters; is the th model parameter; is the regularization coefficient, with a value of 0.0001.
[0069] The learning rate is adjusted using the cosine annealing strategy, and the formula is: ; In the formula, is the learning rate for the th round; is the minimum learning rate, with a value of 0.00001; is the maximum learning rate, with a value of 0.001; is the total number of rounds; is the current round.
[0070] In step S09, the Mahalanobis distance calculation formula is: ; In the formula, is the Mahalanobis distance between samples and ; is the covariance matrix of the data.
[0071] The merging criterion of Ward's minimum variance method is: ; In the formula, is the variance increment after the merger of clusters and cluster ; and are the number of samples in clusters and cluster respectively; and are the mean vectors of clusters and cluster respectively; represents the Euclidean norm.
[0072] The silhouette coefficient calculation formula is: ;
[0073] In the formula, is the silhouette coefficient of sample ; is the average distance between sample and other samples in the same cluster; is the average distance between sample and the nearest non - same - cluster.
[0074] The Davies - Bouldin index calculation formula is: ; In the formula, is the Davies-Bouldin index; is the number of clusters; is the cluster average distance from the samples of the cluster to the cluster center; is the cluster center of the cluster and the cluster distance between the centers.
[0075] When selecting the optimal number of clusters, the silhouette coefficient and the Davies-Bouldin index are comprehensively considered. The number of clusters with the largest silhouette coefficient and the smallest Davies-Bouldin index is taken, and at the same time, the knowledge of physical oceanography is combined for verification to ensure that the clustering results have practical physical significance.
[0076] The local outlier factor algorithm detects outliers by comparing the density differences between sample points and their local neighborhoods. Compared with traditional distance-based or statistical methods, it is more suitable for processing data with complex distributions. This algorithm takes into account the local distribution characteristics of the data and can detect sample points that are not obvious globally but are abnormal in the local environment, especially suitable for various complex fluctuations and abnormal phenomena existing in the seabed environment.
[0077] Singular value decomposition is a powerful matrix decomposition technique that realizes dimensionality reduction and denoising by extracting the main patterns of the data. In the processing of seabed monitoring data, SVD can effectively separate signals and noise and extract the key information reflecting the main characteristics of the seabed environment. The selection of singular values is based on the cumulative contribution rate of energy, ensuring that the main information in the data is retained while redundant and noisy information is removed.
[0078] The adaptive sliding window method dynamically adjusts the window width according to the characteristics of data changes. It uses a smaller window when the environment changes violently to capture subtle changes, and a larger window when the environment is relatively stable to reduce redundancy. The innovation of this method lies in associating the window width with the change rate of data variance, realizing the adaptive analysis of the dynamic changes of the seabed environment.
[0079] Sample entropy and Lyapunov exponent are important indicators for quantifying the complexity and chaos degree of time series. The larger the sample entropy, the higher the complexity and irregularity of the time series; a positive Lyapunov exponent indicates that the system has chaotic characteristics, and the larger the value, the stronger the sensitivity of the system to the initial conditions. The combination of these two indicators can comprehensively describe the turbulent and disordered state in the seabed environment.
[0080] The exponentially weighted moving average method can smooth short-term fluctuations and reflect long-term trends when calculating the drift amount. The selection of the smoothing factor balances the response speed to new data and the smoothing effect. This method is particularly suitable for dealing with the slow drift phenomenon existing in seabed monitoring data.
[0081] Principal component analysis realizes data dimensionality reduction and feature extraction by finding the main variation directions of the data. When dealing with the drift matrix, PCA can identify the main drift patterns, reducing the data dimensionality while retaining key information. The selection of principal components is based on the proportion of explained variance, ensuring that the main variations in the data are captured.
[0082] Tucker decomposition is a high-order tensor decomposition technique that can handle multi-dimensional data and retain high-order correlations. In the seabed environmental state modeling, Tucker decomposition can simultaneously consider the correlations in three dimensions of time, sensor type, and feature type, realizing the fusion analysis of multi-source data. The dimension selection of the core tensor is based on the proportion of the original dimensions, balancing the model complexity and expressive power.
[0083] The residual neural network adopts a loss function combining MSE and L1 regularization in the drift correction model. The MSE term ensures the prediction accuracy of the model, and the L1 regularization term promotes the sparsity and generalization ability of the model. The cosine annealing learning rate adjustment strategy can use a larger learning rate at the beginning of training to quickly approach the optimal solution, and a smaller learning rate at the end of training for fine-tuning, improving the training efficiency and performance of the model.
[0084] Mahalanobis distance takes into account the covariance structure of the data when calculating sample similarity, can handle the correlation between features, and is more suitable for dealing with multi-variable data compared to Euclidean distance. Ward's minimum variance method pursues a merging strategy with the minimum within-class variance in hierarchical clustering, which is beneficial to forming a compact clustering structure. The silhouette coefficient and Davies-Bouldin index, as clustering evaluation metrics, evaluate the clustering quality from different perspectives, comprehensively considering the within-class compactness and between-class separability, providing an objective basis for determining the optimal number of clusters.
[0085] Specifically, the principle of the present invention is as follows: The technical principle of the present invention is based on a multi-level data processing and model fusion strategy. By combining feature extraction, multi-dimensional analysis, and deep learning of seabed monitoring data, the effective distinction between environmental changes and equipment drift is realized. First, the present invention uses the singular value decomposition method to perform dimensionality reduction on the original data. By retaining the eigenvectors with larger singular values, the main features of the data are effectively extracted, while noise and redundant information are eliminated, laying a foundation for subsequent processing.
[0086] Secondly, the present invention innovatively proposes the concepts of the disorder matrix and the drift matrix, and realizes efficient calculation through CUDA parallel processing technology. The disorder matrix quantitatively describes the non-linear dynamic processes such as turbulence and perturbation in the seabed environment, reflecting the natural change characteristics of the environment; while the drift matrix tracks and analyzes the change trend of sensor performance by comparing the deviation degree between the real-time readings of the sensor and the historical reference values. These two matrices characterize the data features from different dimensions, providing a theoretical basis for distinguishing environmental changes and equipment drift.
[0087] The core of the present invention lies in using tensor decomposition technology to fuse the disordered pattern feature vectors, drift principal components, and anomaly feature matrices, and constructing a seabed environmental state tensor model. Compared with traditional matrix analysis methods, tensor decomposition can retain the high-order correlations between data and capture the structural features in multi-dimensional data more comprehensively. At the same time, the deep-sea environmental perception attention network model designed by the present invention combines the Transformer architecture and the recurrent neural network structure, and through the multi-scale drift perception attention mechanism, can simultaneously focus on short-term fluctuations and long-term drift patterns, effectively learning and compensating for complex sensor drifts.
[0088] By dynamically adjusting the analysis window length and frequency resolution through an adaptive spectrum optimization function, the present invention can accurately capture the data characteristics under different seabed environments; the pre-trained deep-sea environmental perception attention network distinguishes environmental changes from sensor drifts through contrastive learning, and achieves efficient processing under limited computing resources. This multi-level and multi-dimensional data processing method theoretically ensures the effective distinction between environmental changes and equipment drifts, and solves the problem of the accuracy of long-term monitoring data.
[0089] A specific embodiment 1 of the present invention is provided below, and the specific implementation manners of each step in this embodiment 1 are described in detail as follows.
[0090] The specific implementation manner of step S01 is to perform normalization processing on the original data collected by the seabed-based monitoring equipment, and uniformly convert the data with different dimensions into the interval [0, 1]. Specifically, the min-max normalization method is adopted. For each sensor data sequence x, the normalized value is calculated as follows: ; In the formula, is the normalized data value; is the original data value; is the minimum value of this sensor data sequence; is the maximum value of this sensor data sequence.
[0091] Then, an improved local outlier factor algorithm is used to detect outliers. This algorithm identifies outliers by calculating the local density ratio between a sample point and its k-nearest neighbor sample points. The calculation formula is: ; In the formula, is the local outlier factor of point ; is the nearest neighbor point set of point ; is the local reachability density of point ; is the number of nearest neighbor points.
[0092] The formula for local reachability density is as follows: ; In the formula, is the reachable distance from point to point , which is defined as: ; In the formula, is the distance from point to its th nearest neighbor point; is the Euclidean distance between point and point .
[0093] The value of k is adaptively selected according to the scale of the data set, generally taking 5% - 10% of the total number of samples. The outlier determination threshold is set to 2.5, that is, when the local outlier factor of a certain data point is greater than 2.5, it is determined as an outlier and removed. The purpose of this step is to eliminate noise and outliers in the data, ensure the quality of the data for subsequent analysis and processing, and improve the reliability and accuracy of data processing.
[0094] The specific implementation of step S02 is to organize the effective data set obtained in step S01 into a matrix form X, where the rows represent time points and the columns represent the measurement values of different sensors. Then, apply singular value decomposition to matrix X: ; In the formula, is the original data matrix, is the number of time points, is the number of sensors; and are orthogonal matrices respectively; is a diagonal matrix, and the elements on the diagonal are singular values.
[0095] Sort according to the magnitudes of the singular values, and select the first r singular values and the corresponding eigenvectors with a cumulative contribution rate reaching 95% to form the reduced - dimensional feature matrix: ; In the formula, contains the first columns of ; contains the diagonal matrix composed of the first singular values of ; contains the first columns of . The selection of ; In the formula, is the th singular value. In practice, the r value is usually selected such that the energy of the retained singular values accounts for more than 95% of the total energy. For typical seabed monitoring data, the r value is generally between 10 and 20. The purpose of this step is to reduce the data dimension, eliminate redundant information and noise in the data, extract key information reflecting the main characteristics of the seabed environment, and provide a more compact and effective data representation for subsequent processing.
[0096] The specific implementation of step S03 is to perform time-domain segmentation on the feature matrix generated in step S02 using an adaptive sliding window method. First, set the basic window width to 1 hour, and set an adaptive adjustment coefficient α according to the characteristics of seabed environment changes. This coefficient depends on the environmental change rate and data acquisition frequency, and generally ranges from 0.5 to 2.0. Monitor the variance change rate of real-time data. The adjustment formula for the window width is: ; In the formula, is the actual window width; is the basic window width, set to 1 hour; is the adjustment coefficient, with a value range of 0.5 to 2.0; is the data variance change rate, and the calculation formula is: ; In the formula, is the data variance at the current moment; is the data variance at the previous moment. When the variance change rate exceeds a preset threshold (usually 20%), the window width is adjusted according to the above formula to ensure that the window width is reduced when the environment changes violently and expanded when the environment is relatively stable. The overlap rate between adjacent windows is set to 50% to ensure data continuity and temporal correlation. Through this adaptive window segmentation method, the feature matrix is divided into multiple groups of time series data blocks, and each data block contains all sensor feature data within a certain period of time. The purpose of this step is to adapt to the dynamic change characteristics of the seabed environment, ensure that subtle change information can be captured when the environment changes violently, and reduce data redundancy and improve processing efficiency when the environment is relatively stable.
[0097] The specific implementation of step S04 is to start a CUDA parallel processing stream and allocate computing resources to process the multiple groups of time series data blocks generated in step S03. For each data block, calculate the disorder degree indicators, including entropy value, Lyapunov exponent, and complexity, etc. Specifically, use the sample entropy algorithm to calculate the complexity and irregularity of the time series: ; In the formula, is the sample entropy value; is the embedding dimension, with a value of 2; is the similarity tolerance, with a value of 0.2 times the standard deviation of the data; is the data length; is the number of pattern pairs matched at dimension ; is the number of pattern pairs matched at dimension .
[0098] At the same time, the Lyapunov exponent is calculated to evaluate the chaos degree of the sequence. First, phase space reconstruction is performed: ; In the formula, is the reconstructed state vector; is the th point of the time series; is the reconstruction delay, with a value of 1 / 4 of the period; is the reconstruction dimension, determined by the false nearest neighbor method.
[0099] Then the Lyapunov exponent is calculated: ; In the formula, is the Lyapunov exponent; is the th time point; is the distance between adjacent orbits in the phase space at time
[0100] These disorder indexes are organized into a disorder matrix M. Each element in the matrix represents the disorder index of the jth sensor in the ith time window. In parallel, the deviation degree of each sensor reading relative to the historical reference value is calculated using the exponentially weighted moving average method: ; In the formula, is the smoothed value at time ; is the observed value at time is the smoothing factor, with a value of 0.1; is the smoothed value at time
[0101] The drift amount is defined as: ; In the formula, is the drift amount at time is the observed value at time is the smoothed value at time
[0102] Construct a drift matrix D, where each element in the matrix represents the drift amount of the j-th sensor in the i-th time window. The threshold is set to 3 times the standard deviation. This step uses the CUDA parallel computing architecture to distribute large-scale matrix operations to thousands of computing cores for simultaneous execution, significantly improving the processing efficiency. The purpose is to quantitatively describe the turbulence, perturbation, and sensor drift in the seabed environment and provide a basis for subsequent analysis.
[0103] The specific implementation of step S05 is to perform singular value decomposition on the disorder matrix M constructed in step S04 in the CUDA parallel processing stream: ; In the formula, is the disorder matrix, is the number of time windows, is the number of sensors; and are orthogonal matrices respectively; is a diagonal matrix, and the elements on the diagonal are singular values.
[0104] Analyze the singular value distribution, and select the right singular vectors corresponding to the first k largest singular values as the disorder pattern feature vectors. The value of k is selected to satisfy: ; In the formula, is the -th singular value; is the number of selected feature vectors.
[0105] At the same time, perform principal component analysis on the drift matrix D and calculate the covariance matrix: ; In the formula, is the covariance matrix; is the drift matrix, is the number of samples, is the number of sensors.
[0106] Solve the characteristic equation to obtain the eigenvalues and the eigenvectors , and select the first p principal components with an explained variance of 85% as the drift principal components. The selection satisfies: ; In the formula, is the covariance matrix of the Eigenvalues; is the number of selected principal components.
[0107] This step is executed in parallel on the GPU, making full use of the high-performance computing capabilities of the CUDA architecture. The purpose is to extract the main patterns and features from the disorder matrix and drift matrix, reduce the data dimension, and extract the key information reflecting the disorder state of the seabed environment and the drift trend of the sensor.
[0108] The specific implementation of step S06 is to extract the features of the abnormal data identified in step S01 in the CPU processing stream. First, group the abnormal data according to the time window, and calculate the statistical features of the abnormal data within each window, including frequency, amplitude, spatial distribution, and time distribution, etc. Use the kernel density estimation method to analyze the distribution characteristics of the abnormal data: ; where is the kernel density estimate value; is the number of samples; is the bandwidth parameter, with a value of 0.1; is the kernel function, usually the Gaussian kernel, expressed as: ; where is the standardized distance.
[0109] Then extract the abnormal pattern features, and apply the random forest algorithm to identify the main types and features of the abnormal data. The Gini impurity calculation formula is: ; where is the Gini impurity of the dataset ; is the number of classes; is the proportion of samples of the th class in the dataset.
[0110] The information gain calculation formula is: ; where is the information gain of the feature ; is the number of values of the feature ; is the sample subset where the feature takes the value ; and are the number of samples of the subset and the total dataset respectively.
[0111] Set the number of trees in the random forest to 100 and the minimum number of samples in a leaf node to 5. Organize the extracted abnormal features into an abnormal feature matrix A, and align A with the disorder pattern feature vector and drift principal components extracted in step S05 in the time dimension through timestamp alignment to generate a unified time index, ensuring the temporal consistency of the three types of feature data. The purpose of this step is to extract the environmental information contained in the abnormal data and align it with the disorder pattern and drift principal components in the time dimension, providing a basis for subsequent data fusion.
[0112] The specific implementation of step S07 is to fuse the disorder pattern feature vector, drift principal components, and abnormal feature matrix obtained in steps S05 and S06 using tensor decomposition technology. First, construct a third-order tensor X, whose three dimensions correspond to time, sensor type, and feature type (disorder, drift, abnormal) respectively. Then apply Tucker decomposition to decompose the tensor X into the product of a core tensor G and three factor matrices A, B, and C: ; where is the original third-order tensor, is the size of the time dimension, is the size of the sensor type dimension, is the size of the feature type dimension; is the core tensor, , , ; , , are factor matrices; denotes the tensor-matrix product along the th dimension.
[0113] The formula for selecting the dimensions of the core tensor is: ; ; ; where denotes rounding up; , , are adjustment parameters, and their value ranges are 0 to 0.2 times the original dimension size. By analyzing the element sizes of the core tensor and their corresponding factor vectors, identify the coupling relationships between different features and construct a seabed environmental state tensor model. This model can express the complex associations among environmental factors, sensor characteristics, and data anomalies, providing a theoretical basis for multi-dimensional information fusion in seabed state assessment. The purpose of this step is to discover the hidden multi-dimensional association patterns in the data through high-order tensor decomposition and construct a tensor model that comprehensively reflects the seabed environmental state.
[0114] The specific implementation of step S08 is to construct a data drift correction model based on a residual neural network. The network structure includes 1 input layer, 4 residual blocks, and 1 output layer. Each residual block consists of 2 convolutional layers and 1 shortcut connection. The input layer receives the drift principal components extracted in step S05, and the output layer generates drift correction coefficients. The network is trained using the Adam optimizer with an initial learning rate of 0.001, which is adjusted using the cosine annealing strategy: ; In the formula, is the learning rate for the th round; is the minimum learning rate, with a value of 0.00001; is the maximum learning rate, with a value of 0.001; is the total number of rounds; is the current round.
[0115] The loss function uses a combination of mean squared error and L1 regularization: ; In the formula, is the total loss; is the mean squared error, and its calculation formula is: ; In the formula, is the number of samples; is the true value; is the predicted value.
[0116] is the L1 regularization term, and its calculation formula is: ; In the formula, is the number of model parameters; is the th model parameter; is the regularization coefficient, with a value of 0.0001.
[0117] The training data is selected from the data segments with known drift patterns in the historical monitoring data and divided into a training set and a validation set in a ratio of 8:2. During the training process, when the validation loss does not decrease for 5 consecutive rounds, the early stopping mechanism is triggered. The trained model is applied to the real-time monitoring data, and correction coefficients are generated according to the identified drift patterns to correct and compensate the data. The purpose of this step is to learn the long-term drift law in the seabed monitoring data, establish an effective drift correction model, and improve the accuracy and reliability of the long-term monitoring data.
[0118] The specific implementation of step S09 is to classify the data processed in the previous steps using the hierarchical clustering algorithm. First, define the distance metric between data points, and calculate the similarity between samples using the Mahalanobis distance: ; In the formula, is the Mahalanobis distance between samples and ; is the covariance matrix of the data.
[0119] Then, use the Ward minimum variance method as the clustering criterion. This method selects the two clusters with the smallest increase in within-class variance for merging at each step: ; In the formula, is the variance increment after the merger of cluster and cluster ; and are the number of samples in cluster and cluster respectively; and are the mean vectors of cluster and cluster respectively; represents the Euclidean norm.
[0120] The clustering process starts from individual samples, gradually merges the most similar clusters, and finally forms a hierarchical structure. Evaluate the quality of different clustering results by calculating the silhouette coefficient and the Davies-Bouldin index, and determine the optimal number of clusters. The formula for calculating the silhouette coefficient is: ; In the formula, is the silhouette coefficient of sample ; is the average distance between sample and other samples in the same cluster; is the average distance between sample and the nearest non-same cluster.
[0121] The formula for calculating the Davies-Bouldin index is: ; In the formula, is the Davies-Bouldin index; is the number of clusters; is the average distance from the samples in cluster to the cluster center; is the center of cluster and cluster The distance between the centers.
[0122] The optimal number of clusters is generally between 4 and 8. According to the knowledge of physical oceanography and combined with expert experience, physical meanings are assigned to each cluster, an evaluation criterion for seabed conditions is established, and it is divided into four levels: normal, slightly abnormal, moderately abnormal, and severely abnormal. Finally, based on the clustering results of real-time monitoring data, a seabed condition evaluation report is generated. The purpose of this step is to scientifically classify the processed data, establish an objective evaluation criterion for seabed conditions, and provide decision-making support for seabed environmental monitoring and early warning.
[0123] Optionally, the seabed-based monitoring device adopts a combination of a multi-source sensor array system, a high-performance data processing module, and an efficient energy supply system to achieve all-round monitoring and data processing of the seabed environment. The data processing and analysis methods in the above steps fully consider the characteristics of the seabed environment and the characteristics of monitoring data, and construct a complete set of seabed environmental monitoring data processing methods through links such as normalization processing, anomaly detection, feature extraction, multi-dimensional data fusion, drift correction, and state evaluation. This method can effectively process multi-source heterogeneous data collected by seabed-based monitoring devices, identify environmental changes and equipment drift, provide accurate and reliable seabed condition evaluation results, and provide data support and decision-making basis for marine scientific research and seabed resource development.
[0124] Furthermore, as Figure 2 shown, the seabed-based monitoring device in Embodiment 1 of the present invention adopts a modular design. The main structure is composed of a titanium alloy shell, which is flat and cylindrical, with a diameter of 75 cm and a height of 30 cm, and the overall weight is about 120 kg. The device shell adopts a multi-layer composite structure, with an inner layer of electromagnetic shielding material, a middle layer of high-strength heat insulation material, and an outer layer of corrosion-resistant titanium alloy, which can withstand the high-pressure environment of a water depth of 6000 meters. The bottom is designed with an adjustable support foot system, which ensures the horizontal stability of the device on an uneven seabed through hydraulic control. The sensor array system is evenly distributed around the device periphery, mainly including 8 high-precision pressure sensors, 12 distributed temperature sensors, 4 three-dimensional acoustic Doppler current meters, 6 three-axis seismic sensors, and various chemical substance detectors, covering an area with a radius of 20 meters around the device. The sensors adopt a modular interface design and can be flexibly configured and replaced according to the characteristics of different sea areas.
[0125] The data acquisition unit is located in the center of the equipment. It uses a 32-bit high-performance microcontroller as the core and is equipped with a high-precision analog-to-digital converter. The sampling rate can be dynamically adjusted according to monitoring needs, supporting up to 2,000 samples per second. It also has a 256GB solid-state storage capacity, which can continuously store 3-6 months of complete monitoring data. The signal processing module integrates a low-power FPGA and two DSP chips to form a three-level processing architecture. The front-end FPGA is responsible for signal filtering and data preprocessing, the mid-end DSP performs complex algorithm calculations, and the back-end embedded computing unit is responsible for data management and decision analysis. This architectural design enables the device to process a large amount of sensor data in real time in the seabed environment, execute complex algorithm processes from S01 to S09, and especially complete large-scale matrix calculation tasks under CUDA acceleration, to achieve accurate evaluation and prediction of the seabed environment status.
[0126] The energy supply system adopts an innovative hybrid energy design. The main body is a high-energy-density lithium-ion battery pack with a total capacity of 600Wh. It is also equipped with a seawater temperature difference energy collector and a micro-water flow power generation device. Under normal marine conditions, it can provide about 15-20W of continuous energy replenishment, extending the working life of the equipment to 1-2 years. To cope with extreme situations, the device is also equipped with an intelligent power consumption management system that can automatically adjust the sampling frequency and data processing depth according to the environmental status and power level, and switch to low-power mode when energy is tight, retaining only the core monitoring function.
[0127] The communication transmission module includes two independent systems: one is a high-speed acoustic communication system with an operating frequency in the range of 18-22kHz, a data transmission rate of up to 10kbps, and an effective communication distance of about 3 kilometers; the other is an optoelectronic composite cable system, which is used to establish a high-speed data link with a nearby relay station or scientific research platform, with a transmission rate of up to 100Mbps. The two systems work together to ensure the reliability of data transmission under various sea conditions. The transmission content mainly includes the processed environmental status assessment results, key characteristic indicators, and system self-diagnosis information. The original data that needs to be analyzed in depth can be uploaded in batches after receiving the command.
[0128] The software system adopts a layered architecture. The bottom layer is a real-time operating system that provides precise time synchronization and task scheduling; the middle layer is a data processing engine that implements the entire algorithm flow from S01 to S09; the top layer is an intelligent decision-making system that is responsible for monitoring data anomaly pattern recognition, environmental status assessment, and early warning generation.
[0129] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: When conducting deep-sea environmental research in the northern slope area of the South China Sea, researchers deployed a set of seabed-based monitoring equipment. The equipment is located at a water depth of about 1200 meters and is equipped with various sensors such as pressure sensors, temperature sensors, flow velocity sensors, seismic sensors, and chemical substance detectors. The monitoring equipment continuously operates for 18 months and collects data at a frequency of once every 30 minutes, aiming to observe seabed environmental changes, cold seep activities, and potential geological disasters. Due to the obvious influence of the monsoon in the South China Sea and the existence of complex water mass exchange and internal wave activities, the monitoring data shows obvious seasonal changes and random disturbance characteristics. At the same time, the sensors showed varying degrees of drift under long-term working conditions, affecting the accuracy of the data. The researchers used the data processing method of the present invention to systematically process and analyze the collected seabed monitoring data.
[0130] First, perform normalization processing and outlier detection on the original data collected by the seabed-based monitoring equipment. The continuous monitoring for 18 months generated approximately 26,280 data points at time points (18 months × 30 days × 24 hours × 2 times / hour), and each time point contains the readings of 5 sensors, totaling 131,400 data points. Apply the min-max normalization method to unify all sensor data into the interval [0, 1], and then use the local outlier factor algorithm for outlier detection. Set the k value to 6% of the total number of samples (about 7884 nearest neighbor points), and the outlier determination threshold to 2.5. As shown in Table 1, the statistical results of outlier detection for different sensors are as follows: Table 1 Statistical table of outlier data detection results for each sensor in the South China Sea
[0131] Outlier detection found that the outlier proportion of the flow velocity sensor is the highest, reaching 5.70%, mainly due to the frequent internal wave activities in the South China Sea, which generate a large number of short-term flow velocity peaks when internal waves pass by. The chemical substance detector has the second highest outlier rate, reaching 3.50%, which is related to the intermittent release of chemical substances such as methane during cold seep activities during the observation period.
[0132] Next, the researchers applied singular value decomposition to the effective data set for dimensionality reduction. Organize the effective data into a matrix form of 24782×5 (take the effective time point data common to all sensors), and obtain the singular value results as shown in Table 2 through SVD decomposition: Table 2 Table of singular value decomposition results of data in the South China Sea
[0133] According to the principle that the cumulative energy accounts for 95%, the first three singular values and the corresponding eigenvectors are selected to construct the reduced feature matrix, reducing the data dimension from 5 dimensions to 3 dimensions, reducing the data volume by 40% while retaining the main information.
[0134] The third step is to implement adaptive sliding time domain window segmentation for the feature matrix after dimensionality reduction. The initial basic window width is set to 2 hours, and the adaptive adjustment coefficient α is set to 1.2. According to the characteristics of the seabed environment in the South China Sea, special attention is paid to the semi-diurnal periodic changes caused by internal tides and internal waves, while taking into account the seasonal changes caused by the monsoon. By monitoring the rate of change of data variance, when the rate of change exceeds the threshold of 25%, the window width is dynamically adjusted. As shown in Table 3, the window width adjustment at different stages of environmental change: Table 3 Adaptive window width adjustment in the South China Sea
[0135] The window segmentation results show that during the prevailing period of the South China Sea summer monsoon, due to frequent typhoon activities and enhanced internal waves, the environment changes dramatically, the data variance change rate is as high as 58.3%, and the window width is automatically adjusted to the minimum value of 0.60 hours, generating the most data blocks; while in the transition period of spring and autumn, the environment is relatively stable, and the window width is expanded to about 2.5 hours, effectively reducing data redundancy.
[0136] In the fourth step, start the CUDA parallel processing stream to receive the time series data blocks and calculate the disorder index and drift in each data block. Using the CUDA 11.4 parallel computing architecture on the NVIDIA RTX 3090 GPU, 18,830 data blocks were distributed to 10,496 CUDA cores for simultaneous processing. The sample entropy (embedding dimension m = 2, similarity tolerance r = 0.2 × σ) and Lyapunov index (reconstruction delay τ = 8 hours, reconstruction dimension d = 4) were calculated, and the disorder matrix was constructed. At the same time, the exponentially weighted moving average method (smoothing factor α = 0.1) was used to calculate the drift of each sensor and construct the drift matrix. Table 4 shows the average values of the disorder index in different seasons: Table 4 Statistics of disorder index in different seasons in the South China Sea
[0137] In the fifth step, SVD decomposition is performed on the disorder matrix in the CUDA parallel processing flow to extract the disorder mode feature vector; principal component analysis is performed on the drift matrix to extract the drift principal component. The analysis results are shown in Table 5: Table 5 Results of principal component analysis of turbulence patterns and drift in the South China Sea
[0138] The results show that the chaotic pattern can be described by 3 eigenvectors, with the cumulative explained variance reaching 91.28%; while the drift characteristics can be represented by 2 principal components, with the cumulative explained variance reaching 88.46%. Parallel computing significantly shortens the processing time, and all calculations are completed in only 8.9 seconds.
[0139] Step 6: Extract abnormal data features in the CPU processing stream. Analyze the 4336 abnormal data points identified in Step 1, use the kernel density estimation method (bandwidth parameter h = 0.1) to analyze the distribution characteristics of the abnormal data, and use the random forest algorithm (number of trees = 100, minimum number of samples in leaf nodes = 5) to identify the main types and characteristics of the abnormal data. The types and distributions of the extracted abnormal features are shown in Table 6: Table 6 Analysis results of abnormal data features in the South China Sea
[0140] Step 7: Use tensor decomposition technology to fuse the chaotic pattern eigenvectors, drift principal components, and abnormal feature matrices. Construct a third-order tensor with dimensions [18830×5×3], and apply Tucker decomposition to decompose the tensor into the product of a core tensor and three factor matrices. According to the formula, the dimension of the core tensor is calculated to be [5649×2×1], 𝛥 is set to 0.1×I, 𝛥 and 𝛥ᵣ are both set to 0. By analyzing the element sizes of the core tensor and their corresponding factor vectors, the coupling relationships between different features in the South China Sea are identified, especially the correlation between internal wave activities and chemical substance releases, and a complete seabed environmental state tensor model is constructed.
[0141] Step 8: Build a data drift correction model based on a residual neural network. The network structure includes 1 input layer, 4 residual blocks, and 1 output layer. Use the Adam optimizer to train the network, with an initial learning rate of 0.001, and use the cosine annealing strategy for adjustment. The loss function uses a combination of mean squared error and L1 regularization, with a regularization coefficient λ = 0.0001. Select the known drift pattern data in the historical data for training, and divide the training set and validation set in a ratio of 8:2. After 350 training rounds, the model converges, and the drift correction effects of each sensor are shown in Table 7: Table 7 Drift correction effects of each sensor in the South China Sea
[0142] In the ninth step, the hierarchical clustering algorithm is used to classify the processed data. The Mahalanobis distance is used to calculate the similarity between samples, and the Ward minimum variance method is used as the clustering criterion. By calculating the silhouette coefficient and Davies-Bouldin index, the quality of different clustering results is evaluated, and the optimal number of clusters is determined to be 5. According to the oceanographic characteristics of the South China Sea, physical meanings are assigned to each cluster, and an evaluation standard for seabed conditions is established. The clustering results are shown in Table 8: Table 8 Clustering Results of Seabed Condition Assessment in the South China Sea
[0143] Traditional methods for processing South China Sea seabed monitoring data mainly rely on fixed threshold filtering and simple statistical analysis, making it difficult to cope with the complex and changeable environmental characteristics of the South China Sea. Especially when dealing with long-term time series data, it is unable to effectively identify and correct sensor drift, resulting in a gradual decline in data quality. At the same time, traditional methods lack the ability to fuse multi-source heterogeneous data and are difficult to discover complex associations between different environmental factors. In addition, the computational efficiency is low, making it difficult to support real-time analysis of large-scale data, and it usually takes several days to complete all data processing.
[0144] Compared with traditional methods, the data processing method of the present invention shows significant advantages in the application in the South China Sea. First, through the adaptive window segmentation technology, it can dynamically adjust the analysis window according to the characteristics of monsoon alternation and internal wave activities in the South China Sea, improving the ability to capture short-term environmental changes. Second, the CUDA parallel processing technology reduces the data processing time from 72 hours of traditional methods to 3.5 hours, with an efficiency improvement of about 20 times. Third, the tensor decomposition technology realizes the extraction of the correlation pattern between internal waves and cold seeps unique to the South China Sea, revealing the coupling relationship between environmental factors that was difficult to discover in the past. Fourth, the drift correction model constructed by the residual neural network reduces the average sensor drift rate from 3.84% to 0.67%, and effectively corrects the large drift (6.83%) of the chemical substance detector in particular. Fifth, the state evaluation method based on hierarchical clustering identifies five typical seabed states in the South China Sea, providing a scientific basis for resource exploration and environmental monitoring in the South China Sea. Compared with traditional methods, the present invention achieves high precision, high efficiency, and high integration in the processing of South China Sea seabed monitoring data, providing strong data support for deep-sea research in the South China Sea.
[0145] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 9, 10, and 11 as follows.
[0146] Table 9 Variable Explanation Table (Part 1)
[0147] Table 10 Variable Explanation Table (Part 2)
[0148] Table 11 Variable Explanation Table (Part III)
[0149] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A data processing method for seabed-based monitoring equipment, characterized in that, Including: Performing normalization processing on the data collected by the seabed-based monitoring equipment and eliminating abnormal data to determine the effective data set; Performing dimensionality reduction processing on the effective data set using the singular value decomposition method to extract key features and generate a feature matrix; Implementing sliding time-domain window segmentation on the feature matrix to form multiple groups of time series data blocks; Starting a CUDA parallel processing stream to construct a disorder matrix and a drift matrix; extracting disorder pattern feature vectors and drift principal components in the CUDA parallel processing stream; extracting abnormal data features in the CPU processing stream to generate an abnormal feature matrix; using tensor decomposition technology to fuse the disorder pattern feature vectors, drift principal components, and abnormal feature matrix to construct a seabed environmental state tensor model; constructing a data drift correction model to correct and compensate the real-time monitoring data; using a hierarchical clustering algorithm to classify the processed data to establish a seabed state evaluation criterion.
2. The data processing method of the seabed-based monitoring device according to claim 1, wherein, The singular value decomposition method specifically organizes the seabed monitoring data into a matrix form, and through calculating eigenvalues and eigenvectors, decomposes the data matrix into the product of three sub-matrices.
3. The data processing method of the seabed-based monitoring device according to claim 2, characterized in that, The disorder matrix specifically refers to a feature matrix constructed by calculating the volatility, uncertainty, and disorder degree of the seabed monitoring data in the time and space dimensions, and is used to quantitatively describe the turbulence, perturbation, and nonlinear dynamic processes in the seabed environment.
4. The data processing method of the seabed-based monitoring device according to claim 3, characterized in that The drift matrix specifically refers to a matrix constructed by comparing the deviation degree of the real-time readings of the sensors with the historical reference values. Each element in the matrix represents the drift amount of the corresponding sensor at the corresponding time point, and is used to track and analyze the performance changes of the sensors and the long-term evolution trend of the environment.
5. The data processing method of the seabed-based monitoring device according to claim 4, characterized in that An adaptive spectrum optimization function is used in the process of constructing the disorder matrix, which is dynamically adjusted according to the frequency characteristic differences of the monitoring data under different seabed environmental conditions. By dynamically adjusting the analysis window length and frequency resolution, the accurate capture of the data characteristics under different seabed environments is realized.
6. The data processing method of the seabed-based monitoring device according to claim 5, characterized in that, A pre-trained deep-sea environmental perception attention network model is used in the data drift correction process. This model integrates the Transformer architecture and the recurrent neural network structure, and introduces a multi-scale drift perception attention mechanism designed for the characteristics of seabed data, and is used to identify and compensate the complex drift patterns in the seabed monitoring data at different time scales.
7. The data processing method of the seabed-based monitoring device according to claim 6, characterized in that The specific structure of the deep-sea environmental perception attention network model is a hybrid architecture combining a multi-layer bidirectional recurrent neural network and a multi-head self-attention mechanism. The bottom layer uses a convolutional neural network to extract features from multi-source sensor data, the middle layer uses six-layer Transformer encoders to extract temporal features and cross-sensor correlation features, and the top layer uses a fully connected network with skip connections for data drift estimation and compensation.
8. The data processing method of the seabed-based monitoring device according to claim 7, characterized in that, The deep-sea environmental perception attention network model introduces a seabed characteristic-aware sparse attention mechanism in the Transformer encoder. The sparse attention mechanism dynamically adjusts the attention range according to the physical characteristics of the seabed environment, and is used to simultaneously focus on short-term fluctuations and long-term drift patterns. The whole model adopts a hierarchical design, and processing modules are respectively designed for the data characteristics under different depths and environmental conditions.
9. The data processing method of the seabed-based monitoring device according to claim 8, wherein During the pre-training process of the deep-sea environmental perception attention network model, long-term monitoring data collected from the global seabed monitoring network is used to screen data segments that contain complete environmental change cycles and sensor drift phenomena, and a physical oceanography model is used to generate synthetic data to enhance the coverage of the dataset, forming a comprehensive training dataset covering various typical seabed environments such as shallow sea areas, deep-sea plains, trench areas, and hydrothermal areas.
10. The data processing method of the seabed-based monitoring device according to claim 9, characterized in that, The seabed-based monitoring device mainly consists of a sensor array system, a data acquisition unit, a signal processing module, an energy supply system, a communication transmission module, and a protective housing. The sensor array system includes a pressure sensor, a temperature sensor, a flow velocity sensor, a seismic sensor, and a chemical substance detector.
Citation Information
Patent Citations
Determination method of marine ecological environment damage causal-relationships
CN108492007A
High-resolution temperature field reconstruction system based on optimized sonic sensor array
CN117571152A
Salinity drift correction method and system for buoy observation data
CN119249061A
Sea level rise influence risk assessment method, medium and system
CN119379018A
Marine water environment abnormal state secondary and derivative probability modeling and estimating method
CN119720046A
Cited By
Airborne navigation data cleaning method based on multi-source sensing data
CN120407552A
Coking coal detection chamber online monitoring method and system
CN120783474A
Automatic processing method for real-time observation data of ocean station
CN121479292A
A method for automatic processing of real-time observation data of a marine station
CN121479292B