Energy-saving transmission method of buoy monitoring equipment

Through the method of classifying marine data and deep learning compression, the contradiction between energy consumption and data integrity in data transmission of marine buoy equipment is solved, and efficient and energy-saving data transmission is achieved, which extends the working time of the equipment and maintains data quality.

CN120342552APending Publication Date: 2025-07-18青岛道万科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818234.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional marine buoy monitoring equipment has a balance between energy consumption and data integrity in data transmission. Existing methods usually either sacrifice data integrity or have high energy consumption, making it difficult to achieve efficient energy utilization while ensuring data integrity.

Method used

The AdaClust clustering algorithm and HyperLogLog technology are used to classify ocean data, set up differentiated time windows, feature extraction and compression through similarity matching and OceanFormer model, and combined with low-power transmission protocols, data compression and efficient transmission are achieved.

Benefits of technology

On the premise of ensuring data integrity, significantly reduce data transmission volume and energy consumption, extend the working time of float equipment, and improve data transmission efficiency and equipment life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120342552A_ABST
    Figure CN120342552A_ABST
Patent Text Reader

Abstract

The invention provides an energy-saving transmission method of buoy monitoring equipment, which belongs to the technical field of electric digital data processing, and comprises the following steps of: classifying ocean data, dividing the data into large, medium and small change sections, and setting corresponding time windows; segmenting the data based on a time window, calculating the similarity between a window data segment and a classic data segment, replacing the high-similarity data with an identifier, and calling a marine data optimization function for dimensionality reduction of the low-similarity data; secondly, feature extraction and compression are carried out through a pre-trained OceanFormer model, a multi-head attention mechanism is adopted by OceanFormer, and dynamic adjustment can be carried out according to data characteristics; finally, the processed data are transmitted to a receiving terminal through a low-power-consumption protocol, the receiving terminal restores the identifier according to data packet header information and reconstructs complete ocean data, and balance of energy consumption and data integrity is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric digital data processing. Specifically, it relates to an energy-saving transmission method for buoy monitoring devices. Background Art

[0002] Marine environment monitoring is the basis of marine scientific research. As an important marine observation platform, buoy monitoring devices are responsible for the long-term, continuous, and real-time collection of marine data. Traditional marine buoy monitoring systems collect and transmit data at fixed time intervals. This method is simple and direct but lacks pertinence. In practical applications, marine environmental parameters such as temperature, salinity, pH value, etc. exhibit different variation characteristics at different times, sometimes changing violently and sometimes relatively stably.

[0003] However, traditional data transmission methods use the same collection frequency and transmission strategy for all data, ignoring the differences in data variation characteristics. This results in a large amount of redundant data when the data changes slowly, and may miss key information when the data changes violently. This one-size-fits-all transmission strategy not only wastes limited energy resources, reduces the working life of buoy devices, but also increases the data transmission burden and storage pressure.

[0004] Currently, the industry has tried to reduce energy consumption through various data compression algorithms and low-power transmission protocols. However, these methods usually either sacrifice data integrity for energy conservation or maintain data integrity but still have high energy consumption. How to achieve efficient utilization of energy while ensuring the integrity of marine data monitoring has become the core technical problem faced by marine buoy monitoring devices. Summary of the Invention

[0005] In view of this, the present invention provides an energy-saving transmission method for buoy monitoring devices, which can solve the balance technical problem between energy consumption and data integrity in the data transmission process of existing marine buoy monitoring devices.

[0006] The present invention is implemented as follows: The present invention provides an energy-saving transmission method for a buoy monitoring device, including: comprehensively analyzing the collected ocean data, using a clustering algorithm combined with a cardinality estimation technique to perform massive clustering and grouping on the ocean data, and dividing the ocean data into large-change segments, medium-change segments, and small-change segments; setting different time window lengths for different change segments; performing segmented processing on the ocean data of different change segments based on the time window to form multiple window data segments; calculating the similarity between the window data segments and the classical data segments to generate a similarity matrix; setting a similarity threshold, and when the similarity is higher than the threshold, replacing the original window data segment with the unique identifier of the classical data segment; for the window data segments with similarity lower than the threshold, calling an ocean data optimization function to perform data dimensionality reduction processing; inputting the optimized data segments into a pre-trained deep neural network model for feature extraction and compression, combining the replaced unique identifier with the processed data segments to form a data packet to be transmitted, and adding data packet header information; sending the data packet to be transmitted to the receiving terminal through a low-power transmission protocol, and the receiving terminal restores the data and reconstructs the complete ocean data according to the data packet header information.

[0007] Among them, the clustering algorithm is the AdaClust clustering algorithm, and the cardinality estimation technique is the HyperLogLog technique. Among them, the AdaClust clustering algorithm is an improved algorithm that adds an adaptive clustering center number determination mechanism on the basis of the K-means clustering. The traditional K-means requires specifying the clustering number K in advance, while AdaClust automatically determines the optimal clustering number through the data distribution characteristics; Among them, the large-change segment refers to the data interval where the change amplitude of the ocean data exceeds a preset first threshold within a unit time; the medium-change segment refers to the data interval where the change amplitude of the ocean data is between the preset first threshold and the preset second threshold within a unit time; the small-change segment refers to the data interval where the change amplitude of the ocean data is lower than the preset second threshold within a unit time.

[0008] Among them, the time window length for the large-change segment is set to the first length, the time window length for the medium-change segment is set to the second length, and the time window length for the small-change segment is set to the third length.

[0009] Among them, the first length refers to a time window length of 5 to 10 minutes, which is used to capture the characteristics of the ocean data in the large-change segment; the second length refers to a time window length of 15 to 30 minutes, which is used to capture the characteristics of the ocean data in the medium-change segment; the third length refers to a time window length of 45 to 60 minutes, which is used to process the ocean data in the small-change segment.

[0010] Among them, each element in the similarity matrix represents the similarity value between a pair of window data segments and the classical data segments.

[0011] Among them, the ocean data optimization function is used to perform dimensionality reduction and key feature extraction on the window data segment. The inputs include the window data segment, the time window length, the data change type, the similarity value, and the frequency feature. The output is the optimized data segment that retains the key features after dimensionality reduction.

[0012] Among them, the deep neural network model is the OceanFormer model. The parameters of the multi-head attention mechanism in the OceanFormer model are dynamically adjusted according to the time window length, the data change type, and the similarity threshold.

[0013] Among them, the specific structure of the OceanFormer model is a deep neural network based on the Transformer architecture, which includes two major parts: an encoder and a decoder. The encoder is composed of six layers of multi-head self-attention layers and a feed-forward neural network. Each layer of the multi-head attention mechanism contains eight attention heads. The number of attention heads is dynamically adjusted according to the time window length, and the sparsity of the attention weight matrix is adaptively adjusted according to the data change type.

[0014] Among them, the decoder uses the cross-attention mechanism to fuse the information output by the encoder with the information of the existing classical data segment at the receiving terminal, and ensures the stability of the model through residual connection and layer normalization. The end of the OceanFormer model uses a fully connected layer to map the features to the compressed space.

[0015] Compared with the prior art, the energy-saving transmission method of a buoy monitoring device provided by the present invention proposes an adaptive energy-saving transmission method based on data characteristics. By dividing the ocean data into large, medium, and small change segments and setting corresponding time window lengths, a differential processing strategy for data with different change degrees is realized. This method uses the AdaClust clustering algorithm and the HyperLogLog technology to efficiently complete the classification of massive data, and greatly reduces the redundant data transmission volume through similarity matching and identifier replacement.

[0016] By using the OceanFormer model to extract and compress the features of data with low similarity, the present invention realizes the "on-demand compression" of data, that is, the data segments with large changes retain more details, and the data segments with small changes are compressed to a higher degree, so as to reduce the data transmission volume on the premise of ensuring information integrity. At the receiving end, relying on the pre-trained OceanFormer decoder can accurately reconstruct the original data to ensure the accuracy of data analysis.

[0017] The present invention successfully solves the balance problem between energy consumption and data integrity in the data transmission process of ocean buoy monitoring devices, enables the buoy devices to achieve continuous monitoring for a longer time with lower energy consumption, and at the same time ensures the data quality required for scientific research, which has important application value for long-term monitoring of the ocean environment. Description of the Drawings

[0018] Figure 1 This is a flowchart of the method of the present invention. Detailed Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0020] As Figure 1 shown, it is a flowchart of an energy-saving transmission method for a buoy monitoring device provided by the present invention. This method includes the following steps: S01. Comprehensively analyze the collected ocean data, and use the AdaClust clustering algorithm combined with the HyperLogLog technology to perform massive clustering and grouping on the ocean data, and divide the ocean data into large change segments, medium change segments, and small change segments; S02. Set the time window length to the first length for the large change segment, set the time window length to the second length for the medium change segment, and set the time window length to the third length for the small change segment; S03. Perform segmented processing on the ocean data based on the time window to form multiple window data segments, and each window data segment contains all the collected data between the window start time and the window end time; S04. Calculate the similarity between each window data segment and the classical data segments stored in the database to generate a similarity matrix, and each element in the similarity matrix represents the similarity value of a pair of the window data segment and the classical data segment; S05. Set a similarity threshold. When the similarity between the window data segment and the classical data segment is higher than the similarity threshold, use the unique identifier of the classical data segment to replace the original window data segment; S06. For the window data segments with similarity lower than the similarity threshold, call the ocean data optimization function to perform data dimensionality reduction processing. The input includes the window data segment, the time window length, the data change type, the similarity value, and the frequency characteristics, and the output is the optimized data segment; S07. Input the optimized data segment into the pre-trained OceanFormer model for feature extraction and compression. The parameters of the multi-head attention mechanism in the OceanFormer model are dynamically adjusted according to the time window length, the data change type, and the similarity threshold; S08. Combine the replaced unique identifier with the data segment processed by the OceanFormer model to form a data packet to be transmitted, and add data packet header information, where the data packet header information includes the data change type and the time window length; S09. Send the data packet to be transmitted to the receiving terminal through a low-power transmission protocol. The receiving terminal restores the unique identifier to the corresponding classical data segment according to the data packet header information, and reconstructs the complete ocean data through the decoder of the OceanFormer model.

[0021] Among them, the AdaClust clustering algorithm is an adaptive mass data clustering method that realizes efficient clustering of mass data by dynamically adjusting the number of clustering centers.

[0022] Among them, the HyperLogLog technology is a probabilistic algorithm for estimating the cardinality of mass data, which occupies extremely little memory space in data analysis to process large-scale data sets.

[0023] Among them, the large change segment refers to the data interval where the change amplitude of the ocean data exceeds a preset first threshold within a unit time.

[0024] Among them, the medium change segment refers to the data interval where the change amplitude of the ocean data is between a preset first threshold and a preset second threshold within a unit time.

[0025] Among them, the small change segment refers to the data interval where the change amplitude of the ocean data is lower than a preset second threshold within a unit time.

[0026] Among them, the data change types include three types: the large change segment, the medium change segment, and the small change segment.

[0027] Among them, the first length refers to the time window length of 5 to 10 minutes, which is used to capture the ocean data characteristics of the large change segment.

[0028] Among them, the second length refers to the time window length of 15 to 30 minutes, which is used to capture the ocean data characteristics of the medium change segment.

[0029] Among them, the third length refers to the time window length of 45 to 60 minutes, which is used to process the ocean data of the small change segment.

[0030] Among them, the frequency feature refers to the frequency domain feature vector obtained by performing spectral analysis on the window data segment.

[0031] Among them, the ocean data optimization function is used to perform dimensionality reduction and key feature extraction on the window data segment. The inputs include the window data segment, the time window length, the data change type, the similarity value, and the frequency feature, and the output is the optimized data segment that retains key features after dimensionality reduction.

[0032] Among them, the specific structure of the OceanFormer model is a deep neural network based on the Transformer architecture, which includes two major parts: an encoder and a decoder. The encoder consists of six layers of multi-head self-attention layers and a feed-forward neural network. Each layer of the multi-head attention mechanism contains eight attention heads, and the number of attention heads is dynamically adjusted according to the time window length. The attention weight matrix adaptively adjusts the sparsity according to the data change type. The decoder uses a cross-attention mechanism to fuse the information output by the encoder with the information of the existing classical data segment of the receiving terminal, and ensures the stability of the model through residual connection and layer normalization. The end of the OceanFormer model uses a fully connected layer to map features to the compressed space.

[0033] Among them, the steps for establishing the training data set during the pre-training process of the OceanFormer model specifically include collecting the ocean monitoring data for up to three years from buoy devices in multiple sea areas around the world, cleaning and standardizing the ocean monitoring data, and dividing it into a training set and a validation set according to the data change type. The training set contains one million pieces of data for each of the large change segment, the medium change segment, and the small change segment, and the validation set contains one hundred thousand pieces of data for each of the large change segment, the medium change segment, and the small change segment. Feature engineering modules are constructed for different time window lengths to extract time-frequency domain features, and the mapping relationship between the original data and the compressed data is established as a supervision signal to construct a self-supervised learning task to enhance the feature extraction ability of the OceanFormer model.

[0034] Among them, the steps for pre-training the OceanFormer model specifically include first performing self-supervised pre-training on a large scale of ocean data to learn the internal representation of the data, using a masked autoencoder structure to randomly mask some features of the input data and training the OceanFormer model to recover, and then performing supervised fine-tuning on the sea area data to train the OceanFormer model to adapt to the ocean characteristics of different regions. During the fine-tuning process, a dynamic batch size and learning rate scheduling strategy are used to ensure the training stability, a multi-task learning framework is used to optimize the reconstruction loss and the compression rate at the same time, the generalization performance of the OceanFormer model is evaluated on the validation set and the optimal checkpoint is selected, and finally, the knowledge of the large model is transferred to the lightweight model through knowledge distillation technology to reduce the deployment resource requirements. The entire pre-training process is completed on a distributed computing cluster to accelerate the training process.

[0035] The specific implementation manners of the above steps are described in detail below. The specific implementation manner of step S01 is to comprehensively analyze and process the collected marine data. First, the original marine data is preprocessed, including outlier detection and processing, missing value filling, and data standardization. The preprocessed data is used as the input of the AdaClust clustering algorithm. The AdaClust clustering algorithm realizes efficient clustering of massive data by dynamically adjusting the number of clustering centers. The algorithm first randomly selects multiple initial clustering centers, then calculates the optimal number of clustering centers based on the data distribution density, and then performs iterative clustering assignment on the data points. During this process, the HyperLogLog technology is combined to estimate the cardinality of massive data. The HyperLogLog technology processes massive data sets in an extremely small memory space through probability statistics methods, enabling the system to effectively process massive buoy monitoring data under limited resource conditions. After clustering, the marine data is divided into large change segments, medium change segments, and small change segments according to the data change amplitude. The change amplitude is based on the average change rate of adjacent sampling points within one hour. The preset first threshold for the large change segment is a change rate of more than 15% per hour, and the preset second threshold for the small change segment is a change rate of less than 3% per hour. Those in between are classified as medium change segments. The purpose of this step is to reasonably classify the marine data and lay a foundation for subsequent targeted processing.

[0036] The specific implementation manner of step S02 is to set corresponding time window lengths for different types of changing marine data. For the large change segment, the time window length is set to the first length, specifically in the range of 5 to 10 minutes, and generally 8 minutes is taken as the default value. This setting can capture rapidly changing marine environmental parameters in a timely manner. For the medium change segment, the time window length is set to the second length, specifically in the range of 15 to 30 minutes, and generally 20 minutes is taken as the default value. This setting takes into account both the data representativeness and the transmission efficiency. For the small change segment, the time window length is set to the third length, specifically in the range of 45 to 60 minutes, and generally 50 minutes is taken as the default value. This setting is suitable for slowly changing marine parameters and can significantly reduce the data transmission volume. The setting of the time window length will also be fine-tuned according to different sea area characteristics, seasonal changes, and monitoring index types. The optimal window length is calculated through an adaptive algorithm, which establishes a window length optimization function based on the historical data change rate and the energy consumption model. The purpose of this step is to implement a differential processing strategy for data with different degrees of change and balance data accuracy and transmission energy consumption.

[0037] The specific implementation of step S03 is to segment the ocean data based on a set time window. First, the continuous ocean data is divided according to the time window length corresponding to the change type based on the data timestamp, forming multiple window data segments. Each window data segment contains all the collected data from the start time to the end time of the window. The sliding window technique is used in the division process, and there is no overlap between adjacent windows to ensure data integrity and avoid redundancy. For each window data segment, its basic statistical features are calculated, including statistics such as mean, variance, maximum value, minimum value, median, etc. At the same time, time series features such as trendiness, periodicity, autocorrelation coefficient, etc. are extracted and these features are attached to the metadata of the window data segment for subsequent similarity calculation. In addition, the quality of each window data segment is evaluated, abnormal data points are marked, and the data quality score is recorded. The purpose of this step is to reasonably segment the continuous ocean data in the time dimension to facilitate subsequent processing and analysis of each data segment and provide a basis for similarity calculation.

[0038] The specific implementation of step S04 is to calculate the similarity between the window data segment and the classic data segment and construct a similarity matrix. First, a set of classic data segments of the same type and length as the current window data segment are extracted from the database. Classic data segments refer to representative ocean data patterns in history, and these classic data segments are either expert-annotated or extracted from historical data through unsupervised learning methods. Then, a multi-dimensional similarity calculation method is used to evaluate the similarity between each window data segment and the classic data segment. This method comprehensively considers multiple dimensions such as Euclidean distance, dynamic time warping distance, frequency domain similarity, and statistical feature similarity. For time series data, the dynamic time warping algorithm can handle the alignment problem between sequences of different lengths. This algorithm allows non-linear stretching on the time axis, enabling two time series to find the best matching path. The calculated similarity value ranges from 0 to 1, and the closer the value is to 1, the higher the similarity. Finally, a similarity matrix is generated, and each element in the matrix represents the similarity value between a pair of window data segments and classic data segments. The purpose of this step is to find the association between the current monitored data and historical typical patterns and provide a basis for subsequent data compression and efficient transmission.

[0039] The specific implementation of step S05 is to set a similarity threshold and perform data replacement. According to the characteristics of different sea areas, different seasons, and different monitoring parameters, an appropriate similarity threshold is set. Generally, the similarity threshold is set between 0.85 and 0.95. For key monitoring parameters such as tsunami warning-related data, the threshold can be increased to above 0.95. For general environmental parameters such as conventional temperature monitoring, the threshold can be appropriately reduced to around 0.85. When the similarity between the window data segment and a certain classic data segment is higher than the set threshold, the replacement mechanism is triggered, and the unique identifier of the classic data segment is used to replace the original window data segment. This identifier is usually a 32-bit or 64-bit hash value, which occupies extremely small storage space. After replacement, only the identifier needs to be transmitted in the transmission data packet instead of the complete data. The receiving end extracts the corresponding classic data segment from the local database according to the identifier for restoration, greatly reducing the amount of data to be transmitted. For cases where the similarity is close but does not reach the threshold, the system will also record the different parts and evaluate their importance. If the differences do not affect the data analysis results, the system will automatically increase the similarity score to facilitate replacement. The purpose of this step is to significantly reduce the amount of data to be transmitted, improve the transmission efficiency, and reduce energy consumption through data pattern recognition and replacement.

[0040] The specific implementation of step S06 is to call the ocean data optimization function for data dimensionality reduction processing for the window data segment with similarity lower than the threshold. This optimization function first receives input parameters such as the window data segment, time window length, data change type, similarity value, and frequency characteristics, and then adaptively selects the most suitable dimensionality reduction algorithm according to these parameters. For large change segment data, the locally linear embedding algorithm that preserves local features is adopted; for medium change segment data, the t-SNE algorithm that balances global and local features is adopted; for small change segment data, the principal component analysis algorithm is adopted to extract the main change features. During the dimensionality reduction process, the dimensionality reduction ratio is dynamically adjusted according to the data change type. The large change segment retains 25% to 30% of the original data dimensions, the medium change segment retains 15% to 20% of the dimensions, and the small change segment only retains 5% to 10% of the dimensions. At the same time, based on the frequency feature analysis results, the key frequency band information is retained, and multi-scale features are extracted through wavelet transform and important coefficients are screened. Finally, by integrating the time domain and frequency domain features, an optimized data segment with a higher information density is generated. Although the volume of this data segment is significantly reduced, the key features of the original data are retained. The purpose of this step is to reduce the amount of data through dimensionality reduction technology while ensuring the effectiveness of the data, and reduce energy consumption for subsequent transmission.

[0041] The specific implementation of step S07 is to input the optimized data segment into the pre-trained OceanFormer model for feature extraction and compression. The OceanFormer model is a deep neural network based on the Transformer architecture, consisting of two major parts: an encoder and a decoder. The encoder is composed of six layers of multi-head self-attention layers and feed-forward neural networks. Each layer of the multi-head attention mechanism contains eight attention heads. In the model processing, first, the input data is converted into a high-dimensional vector representation through the embedding layer, and then temporal information is added through position encoding. Then, the data sequentially passes through multiple layers of self-attention mechanisms, and the parameters of each attention mechanism are dynamically adjusted according to the characteristics of the current processed window data segment. For large change segment data, the number of attention heads is increased to 12 to strengthen the attention to local features; for small change segment data, the number of attention heads is reduced to 4 to reduce the computational complexity. The attention weight matrix adaptively adjusts the sparsity according to the data change type. A lower sparsity is adopted for large change segment data to ensure the complete capture of information, and a higher sparsity is adopted for small change segment data to further compress the data. The feature vector processed by the encoder then passes through a compression module, which combines adaptive quantization technology and Huffman coding to achieve efficient compression. The purpose of this step is to use a deep learning model to extract the deep features of ocean data and achieve a high compression ratio data representation, laying a foundation for low-power transmission.

[0042] The specific implementation of step S08 is to combine the processed data to form a data packet to be transmitted. First, the replaced unique identifier is integrated with the data segment processed by the OceanFormer model. The identifier data is placed at the front end of the data packet, and the compressed data is placed at the back end, with a delimiter marker set between them. Then, packet header information is added, including key information such as data change type, time window length, data acquisition timestamp, data quality marker, checksum, etc. The packet header adopts a compact design, with a typical size of 64 bytes, where the data change type occupies 2 bits, the time window length occupies 6 bits, the timestamp occupies 32 bits, the quality marker occupies 8 bits, and the checksum occupies 16 bits. The overall structure of the data packet follows the hierarchical design principle, facilitating layer-by-layer parsing at the receiving end. In addition, a metadata segment is added to the data packet to record the key parameters in the data processing process, such as the adopted dimensionality reduction ratio, compression parameters, etc. These information are crucial for the receiving end to reconstruct the original data. A cyclic redundancy checksum is finally appended to the data packet to ensure the data integrity during transmission. The purpose of this step is to encapsulate the processed data in a standard format to ensure that the receiving end can correctly parse and restore the data.

[0043] The specific implementation of step S09 is to send the data packet to be transmitted to the receiving terminal through a low-power transmission protocol. First, evaluate the current network environment and select the most suitable low-power transmission protocol. In the offshore area, the LoRaWAN protocol is preferred. This protocol operates in an unlicensed frequency band and has the characteristics of long-distance and low power consumption, with a transmission distance of up to 10 kilometers. In the open sea area, the satellite communication protocol such as the Iridium short message service is preferred. Although this service has a relatively high power consumption, it covers the global sea area. Before transmission, the data packet is fragmented, and the size of each fragment is dynamically adjusted according to the maximum transmission unit of the selected protocol, generally controlled between 100 and 250 bytes. The transmission adopts an adaptive power control strategy, dynamically adjusting the transmission power according to the channel quality. When the channel quality is good, the power is reduced to save energy, and when the channel quality is poor, the power is increased to ensure the transmission success rate. After receiving the data packet, the receiving terminal first verifies the integrity, and then identifies the data characteristics according to the information in the data packet header. For the data packet containing the unique identifier, the corresponding classic data segment is extracted from the local database. For the data segment processed by the OceanFormer model, it is reconstructed through the decoder of the OceanFormer model. The decoder uses the cross-attention mechanism to fuse the information output by the encoder with the information of the existing classic data segments at the receiving terminal, and ensures the model stability through residual connection and layer normalization, and finally reconstructs the complete ocean data. The purpose of this step is to achieve efficient data transmission through low-power transmission technology and complete data reconstruction at the receiving end, realizing the entire energy-saving transmission process.

[0044] The specific structure of the OceanFormer model optimizes ocean data based on the standard Transformer architecture. Its encoder consists of six layers of multi-head self-attention layers and a feed-forward neural network. Each self-attention mechanism in each layer contains eight attention heads, with each attention head having a dimension of 64, collectively forming a 512-dimensional feature vector space. The encoder input layer uses a temporal embedding method to convert the original ocean data into a high-dimensional representation, and at the same time introduces learnable positional encoding to retain the temporal relationship of the data. After each multi-head attention layer, there is a feed-forward neural network, which consists of two fully connected layers. The dimension of the first layer is 2048, and the GELU activation function is used. The second layer restores the dimension to 512, which is the same as the output dimension of the attention layer. Residual connections and layer normalization are used between the attention layer and the feed-forward layer to ensure stable signal transmission and smooth gradient flow. Considering the periodic characteristics of ocean data, the OceanFormer model introduces frequency-enhanced positional encoding, which combines the characteristics of Fourier transform and can better capture the periodic fluctuations of ocean data. The parameters of the multi-head attention of the model are dynamically adjusted according to the length of the time window. For long-window data, more attention heads are used to capture long-term dependencies, while for short-window data, the number of attention heads is correspondingly reduced to improve computational efficiency. The sparsity of the attention weight matrix is adjusted according to the type of data change. A dense attention matrix is used for data segments with large changes, and a sparse attention matrix is used for data segments with small changes to achieve reasonable allocation of computing resources. The decoder part uses a cross-attention mechanism to fuse the encoder output with the information of the classic data segment at the receiving terminal, and gradually reconstructs the original data features through three layers of cross-attention layers. Finally, a fully connected layer is used at the end of the OceanFormer model to map the features to a compressed space, and the output dimension of this layer can be dynamically adjusted according to the compression requirements, usually controlled between 10% and 30% of the original data dimension.

[0045] The specific implementation of establishing the training dataset for the OceanFormer model is as follows. First, ocean monitoring data for three years is collected from buoy monitoring networks in multiple seas around the world, covering the main seas to be observed. The monitoring parameters include key indicators such as seawater temperature, salinity, flow velocity, wave height, and air pressure. The sampling frequency ranges from once per minute to once per hour. The raw data is cleaned and standardized, including removing obvious outliers, filling missing values, unifying the sampling frequency, and standardizing the numerical range. Data cleaning uses anomaly detection algorithms based on statistics and domain knowledge to mark and process data points outside the normal range. Missing values are filled using linear interpolation, spline interpolation, or time series prediction models according to the data type. Standardization processing unifies data of different scales into the [-1, 1] interval to facilitate model learning. The cleaned data is divided into large change segments, medium change segments, and small change segments according to the data change type. Each type is further divided into a training set and a validation set. The training set contains one million data points of each type, and the validation set contains one hundred thousand data points of each type. Feature engineering modules are constructed for different time window lengths to extract time domain features and frequency domain features. Time domain features include statistics such as mean, variance, kurtosis, and skewness. Frequency domain features are extracted by the fast Fourier transform to obtain the power spectral density and main frequency components. A mapping relationship between the raw data and the compressed data is established as a supervision signal, and the compressed data is obtained through dimensionality reduction techniques such as autoencoders or principal component analysis. In addition, self-supervised learning tasks are constructed to enhance the model's feature extraction ability, such as randomly masking part of the ocean data and training the model to predict the masked part, or predicting the data change trend at future time points. The training dataset constructed by these methods can comprehensively reflect the complex change patterns of the ocean environment and provide sufficient learning materials for the OceanFormer model.

[0046] The following details the mathematical models or calculation processes involved in the present invention.

[0047] In step S01, the AdaClust clustering algorithm and the HyperLogLog technique are used for clustering and grouping ocean data. The AdaClust clustering algorithm is an improved algorithm that adds an adaptive clustering center number determination mechanism based on the K-means clustering. The traditional K-means requires specifying the clustering number K in advance, while AdaClust automatically determines the optimal clustering number through the data distribution characteristics. The core calculation process of the AdaClust clustering algorithm can be expressed as follows: ; In the formula, is the Euclidean distance from the data point to the clustering center ; is the -th dimensional feature value of the data point is the clustering center of the dimensional eigenvalue; is the feature dimension.

[0048] The dynamic adjustment of the number of clustering centers is calculated based on the data distribution density: ; In the formula, is the optimal number of clustering centers; is the total number of data points; is the adjustment coefficient, and its value range is 1.0 to 2.5; is the data point and the Euclidean distance between them.

[0049] The HyperLogLog technology is used for the cardinality estimation of massive data, and its calculation formula is: ; In the formula, is the cardinality estimation value; is the number of registers, usually taking the value of ; is the value of the th register; is the correction coefficient. When is the case, ; when is the case, ; when is the case, .

[0050] The calculation formula of the data change rate is: ; In the formula, is the average change rate of data per unit time; is the number of sampling points per unit time; is the value of the th sampling point.

[0051] In step S02, the adaptive adjustment algorithm for the time window length is based on the historical data change rate and the energy consumption model: ; In the formula, is the optimal window length (in minutes); is the average data change rate; is the energy consumption per unit data transmission (in joules / byte); is the season factor, with a range of 0 to 1; is the proportionality coefficient, with a value range of 5 to 15; is the change rate influence index, and its value range is 0.3 - 0.7; is the energy consumption influence index, and its value range is 0.1 - 0.3; is the seasonal adjustment coefficient, and its value range is 0.1 - 0.5.

[0052] In step S04, the multi - dimensional similarity calculation method evaluates the similarity between the window data segment and the classical data segment. The Euclidean distance calculation formula is: ; In the formula, is the Euclidean distance between data segments and ; is the th data point of data segment is the th data point of data segment is the number of data points.

[0053] The dynamic time warping distance calculation formula is: ; In the formula, is the dynamic time warping distance between data segments and ; is the squared distance of the th matching point pair on the optimal matching path; is the length of the optimal matching path.

[0054] The optimal matching path is solved by the dynamic programming algorithm: ; In the formula, is the cumulative distance from the starting point to point ; is the distance between point and ; the initial condition is .

[0055] The frequency - domain similarity calculation formula is: ; In the formula, is the frequency - domain similarity between data segments and ; is the th component of the spectrum of data segment is the th component; is the number of spectral components.

[0056] The comprehensive similarity calculation formula is: ; In the formula, is the comprehensive similarity, and its value range is 0 to 1; is the maximum possible value of the Euclidean distance; is the maximum possible value of the dynamic time warping distance; is the statistical feature similarity; , , , are the weights of each part respectively, satisfying , generally takes the value of 0.3, takes the value of 0.4, takes the value of 0.2, takes the value of 0.1.

[0057] The statistical feature similarity calculation formula is: ; In the formula, is the statistical feature similarity; , , , are the mean, standard deviation, skewness and kurtosis of the data segment respectively; , , , are the mean, standard deviation, skewness and kurtosis of the data segment respectively; , , , are the standardization factors of the corresponding statistics respectively.

[0058] In step S06, the ocean data optimization function performs data dimensionality reduction processing. For the large change segment data, the locally linear embedding algorithm is adopted: ; In the formula, is the data point after dimensionality reduction; is the th nearest neighbor representation coefficient of the data point in the original space; is the total number of data points; the constraint conditions are and .

[0059] In step S07, the multi-head self-attention mechanism of the OceanFormer model can be expressed as: ; In the formula, , , are the query matrix, key matrix, and value matrix respectively; is the dimension of the key vector; is the sparse mask matrix, which is used to control the sparsity of the attention weights, , when , it means to mask the attention weight at this position.

[0060] The dynamic adjustment formula for the number of attention heads is: ; In the formula, is the actual number of attention heads used; is the basic number of attention heads, with a value of 8; is the current time window length (in minutes); and are the minimum and maximum time window lengths, which are 5 and 60 respectively; is the numerical encoding of the data change type, where a large change segment is 2, a medium change segment is 1, and a small change segment is 0; and are the minimum and maximum data change type encodings, which are 0 and 2 respectively; and are adjustment coefficients, with values of 0.3 and 0.2 respectively.

[0061] The adaptive adjustment formula for attention sparsity is: ; In the formula, is the sparsity, expressed as the proportion of masked elements in the attention matrix; is the basic sparsity, with a value of 0.3; and are adjustment coefficients, with values of 0.1 and 0.15 respectively.

[0062] In step S08, the packet header information structure can be represented as a bit field: , ; In the formula, is the data change type field, which occupies 2 bits, and the values 00, 01, and 10 represent small change, medium change, and large change respectively; is the time window length field, which occupies 6 bits and can represent window lengths from 0 to 63 minutes; is a timestamp field, occupying 32 bits, and adopting the UNIX timestamp format; is a data quality flag field, occupying 8 bits. The upper 4 bits represent the data integrity level, and the lower 4 bits represent the data accuracy level; is a checksum field, occupying 16 bits, and is generated by using the CRC16 algorithm.

[0063] The calculation formula for the cyclic redundancy checksum is: ; In the formula, is a 16-bit cyclic redundancy checksum; is the data to be verified; is the generating polynomial, adopting the CRC-16-CCITT standard, , and the corresponding hexadecimal value is 0x1021.

[0064] In step S09, the calculation formula for the adaptive power control strategy is: ; In the formula, is the actual transmission power (dBm); and are the minimum and maximum transmission powers respectively, and the value ranges are 10~14dBm and 20~30dBm respectively; is the current channel signal-to-noise ratio (dB); and are the minimum and maximum signal-to-noise ratio thresholds respectively, and the values are 5dB and 25dB respectively; is the adjustment exponent, and the value range is 0.5~2.0. When , ; when , .

[0065] Optionally, the similarity matrix mentioned in step S04 can be expressed as: ; In the formula, is the similarity matrix; is the th window data segment and the th classical data segment similarity value; is the number of window data segments; is the number of classical data segments.

[0066] Optionally, for the t-SNE algorithm (applicable to medium-varying segments) mentioned in step S06: ; ; ; ; wherein, is the conditional probability of the nearest neighbor of point in the high-dimensional space; is the conditional probability of the nearest neighbor of point ; the standard deviation of the Gaussian distribution, which is determined by binary search; is the symmetrized joint probability; is the total number of data points; is the similarity between point and point in the low-dimensional space; is the KL divergence between distributions and , which is used as the objective function of t-SNE; and are the low-dimensional representations after dimensionality reduction.

[0067] Optionally, for the principal component analysis algorithm (applicable to small change segments) mentioned in step S06: ; ; ; wherein, is the data covariance matrix; is the total number of data points; is the original data point; is the mean of all data points; and are the th eigenvalue and eigenvector of the covariance matrix, respectively; is the projection matrix composed of the first eigenvectors; is the data point after dimensionality reduction.

[0068] Optionally, wavelet transform is used for multi-scale feature extraction (mentioned in step S06): ; wherein, is the wavelet coefficient of the function at scale and position ; is the wavelet mother function; denotes the complex conjugate of

[0069] Optionally, the position encoding calculation formula mentioned in step S07: ; ; In the formula, For location , the dimension is The positional encoding value of is the feature dimension of the model, and its value is 512.

[0070] Optional, frequency-enhanced position coding (for the periodicity of ocean data): ; In the formula, It is frequency enhanced position coding; is the standard position code; The data-dominant frequency; is the initial phase, Evenly distributed within the range; is the enhancement coefficient, and its value range is 0.1~0.5.

[0071] The design of these equations takes into account the particularity of ocean data and the energy limitations of buoy equipment. The AdaClust algorithm achieves efficient classification of massive data by dynamically adjusting the number of cluster centers; the multi-dimensional similarity calculation comprehensively considers the similarity of time domain, frequency domain and statistical features, improving the accuracy of data matching; the dynamic adjustment of the attention mechanism enables the OceanFormer model to adaptively allocate computing resources according to data characteristics; the adaptive power control strategy dynamically adjusts the transmission power according to the channel quality, maximizing energy conservation while ensuring transmission quality. The overall solution achieves efficient and energy-saving transmission of buoy monitoring equipment data through the organic combination of multiple technologies such as data segmentation, similarity matching, deep learning compression and low-power transmission.

[0072] Specifically, the principle of the present invention is: the core principle of the present invention is to introduce the "data-driven" concept into the field of ocean monitoring data transmission. Ocean monitoring data is often relatively stable most of the time and changes little; however, there may be a situation where the data suddenly changes in a short period of time. The formulation of the transmission strategy is guided by an in-depth analysis of the data characteristics. First, the present invention uses the AdaClust clustering algorithm and HyperLogLog technology to efficiently group massive ocean data. This combination can process large-scale data sets under limited memory conditions, ensuring the efficiency of classification. Based on the classification of data change amplitude, the present invention establishes an adaptive time window mechanism. For data with drastic changes, a shorter time window (5-10 minutes) is used to capture fast-changing features, while for data with slow changes, a longer time window (45-60 minutes) is used to reduce redundant sampling, thereby achieving a balance between data integrity and transmission volume.

[0073] Secondly, the present invention innovatively introduces a similarity comparison mechanism. By constructing a similarity matrix between the window data segment and the classical data segment, the recognition of repetitive data is achieved. When the similarity is higher than the threshold, only the unique identifier of the classical data segment is transmitted. In this way, the original data is transformed into lightweight identifiers, significantly reducing the amount of transmitted data. For data with low similarity, the ocean data optimization function and the OceanFormer model are used for in-depth processing. The multi-head attention mechanism of the OceanFormer model can dynamically adjust parameters according to the data characteristics, capture the long-term and short-term dependencies and periodic patterns in the ocean data, and achieve efficient data compression.

[0074] Finally, through pre-training with a deep neural network based on the Transformer architecture, the present invention ensures that the model has sufficient generalization ability to handle data characteristics under different sea areas and different environmental conditions. The application of knowledge distillation technology enables the model to operate efficiently on resource-constrained buoy devices. This multi-level method combining data characteristic analysis, similarity matching, and deep learning ensures the efficiency and integrity of data transmission in principle and solves the contradiction between energy consumption and data quality.

[0075] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0076] The specific implementation of step S01 is to comprehensively analyze and process the collected ocean data. First, the original ocean data is preprocessed, including outlier detection and processing, missing value filling, and data standardization. The preprocessed data is used as the input of the AdaClust clustering algorithm. The AdaClust clustering algorithm realizes efficient clustering of massive data by dynamically adjusting the number of clustering centers. Its core calculation process is the calculation of Euclidean distance: ; In the formula, is the Euclidean distance from the data point to the clustering center ; is the th dimensional feature value of the data point is the th dimensional feature value of the clustering center is the feature dimension.

[0077] The dynamic adjustment of the number of clustering centers is based on the calculation of data distribution density: ; In the formula, is the optimal number of clustering centers; is the total number of data points; is the adjustment coefficient, with a value range of 1.0 to 2.5; is the data point and the Euclidean distance between them.

[0078] This algorithm first randomly selects multiple initial clustering centers, then calculates the optimal number of clustering centers based on the data distribution density, and then performs iterative clustering assignment on the data points. During this process, the HyperLogLog technology is combined to estimate the cardinality of massive data. The HyperLogLog technology processes massive data sets in a very small memory space through probability statistics methods, and its calculation formula is: ; In the formula, is the cardinality estimate value; is the number of registers, usually taking the value of ; is the value of the th register; is the correction coefficient. When , ; when , ; when , .

[0079] After completing the clustering, the ocean data is divided into large change segments, medium change segments, and small change segments according to the data change amplitude. The change rate calculation formula is: ; In the formula, is the average change rate of the data per unit time; is the number of sampling points per unit time; is the value of the th sampling point.

[0080] The change amplitude is based on the average change rate of adjacent sampling points within one hour. The preset first threshold for the large change segment is a change rate of more than 15% per hour, and the preset second threshold for the small change segment is a change rate of less than 3% per hour. Those in between are classified as medium change segments. The purpose of this step is to reasonably classify the ocean data and lay a foundation for subsequent targeted processing.

[0081] The specific implementation of step S02 is to set corresponding time window lengths for different types of changing ocean data. For large change segments, the time window length is set to the first length, with a specific range of 5 to 10 minutes, and generally 8 minutes is taken as the default value; for medium change segments, the time window length is set to the second length, with a specific range of 15 to 30 minutes, and generally 20 minutes is taken as the default value; for small change segments, the time window length is set to the third length, with a specific range of 45 to 60 minutes, and generally 50 minutes is taken as the default value. The adaptive adjustment algorithm for the time window length is based on the historical data change rate and the energy consumption model: ; In the formula, is the optimal window length (in minutes); is the average data change rate; is the energy consumption per unit data transmission (in joules / byte); is the season factor, with a range of 0 to 1; is the proportionality coefficient, with a value range of 5 to 15; is the change rate influence index, with a value range of 0.3 to 0.7; is the energy consumption influence index, with a value range of 0.1 to 0.3; is the season adjustment coefficient, with a value range of 0.1 to 0.5. The purpose of this step is to implement a differential processing strategy for data with different degrees of change, and balance data accuracy and transmission energy consumption.

[0082] The specific implementation of step S03 is the same as the foregoing, and will not be elaborated here.

[0083] The specific implementation of step S04 is to calculate the similarity between the window data segment and the classic data segment, and construct a similarity matrix. First, extract from the database the set of classic data segments of the same type and length as the current window data segment. Classic data segments refer to representative ocean data patterns in history, and these classic data segments are either marked by experts or extracted from historical data through unsupervised learning methods. Then, use a multi-dimensional similarity calculation method to evaluate the similarity between each window data segment and the classic data segment, including Euclidean distance calculation: ; In the formula, is the Euclidean distance between data segments and ; is the th data point of data segment ; is the th data point of data segment ; is the number of data points.

[0084] Dynamic time warping distance calculation: ; In the formula, is the dynamic time warping distance between data segments and ; is the squared distance of the th matching point pair on the optimal matching path; is the length of the optimal matching path.

[0085] The optimal matching path is solved by the dynamic programming algorithm: ; In the formula, is the cumulative distance from the starting point to point ; is the distance between point and ; The initial condition is .

[0086] Frequency domain similarity calculation: ; In the formula, is the frequency domain similarity between data segments and ; is the th component of the spectrum of data segment ; is the th component of the spectrum of data segment ; is the number of spectrum components.

[0087] Comprehensive similarity calculation: ; In the formula, is the comprehensive similarity, and its value range is 0 to 1; is the maximum possible value of the Euclidean distance; is the maximum possible value of the dynamic time warping distance; is the statistical feature similarity; , , , are the weights of each part respectively, satisfying , generally takes the value of 0.3, takes the value of 0.4, takes the value of 0.2, takes the value of 0.1.

[0088] Calculation of statistical feature similarity: ; In the formula, is the statistical feature similarity; , , , are the mean, standard deviation, skewness, and kurtosis of the data segment respectively; , , , are the mean, standard deviation, skewness, and kurtosis of the data segment respectively; , , , are the standardization factors of the corresponding statistics.

[0089] The calculated similarity value ranges from 0 to 1. The closer the value is to 1, the higher the similarity. Finally, a similarity matrix is generated: ; In the formula, is the similarity matrix; is the similarity value between the -th window data segment and the -th classical data segment; is the number of window data segments; is the number of classical data segments. The purpose of this step is to find the correlation between the current monitored data and the historical typical patterns, providing a basis for subsequent data compression and efficient transmission.

[0090] The specific implementation of step S05 is the same as the foregoing and will not be elaborated here.

[0091] The specific implementation of step S06 is to call the ocean data optimization function for data dimensionality reduction processing for the window data segments with similarity lower than the threshold. This optimization function first receives input parameters such as the window data segment, time window length, data change type, similarity value, and frequency characteristics, and then adaptively selects the most suitable dimensionality reduction algorithm according to these parameters. For the data in the large change segment, the locally linear embedding algorithm that preserves local features is adopted: ; In the formula, is the data point after dimensionality reduction; is the -th nearest neighbor representation coefficient of the data point in the original space; is the total number of data points; the constraint condition is and 。

[0092] For the data in the changing segment, the t-SNE algorithm is adopted: ; ; ; ; In the formula, is the point in the high-dimensional space is the point is the conditional probability of the neighbor of the point; is the standard deviation of the Gaussian distribution, which is determined by binary search; is the symmetrized joint probability; is the total number of data points; is the point in the low-dimensional space and the point is the similarity between; is the distribution and is the KL divergence between, as the objective function of t-SNE; and are the low-dimensional representations after dimensionality reduction.

[0093] For the data in the small changing segment, the principal component analysis algorithm is adopted: ; ; ; In the formula, is the data covariance matrix; is the total number of data points; is the original data point; is the mean of all data points; and are the th eigenvalue and eigenvector of the covariance matrix respectively; is the projection matrix composed of the first eigenvectors; is the data point after dimensionality reduction.

[0094] During the dimensionality reduction process, wavelet transform is also used for multi-scale feature extraction: ; In the formula, is the wavelet coefficient of the function at the scale and the position ; is the wavelet mother function; denote the complex conjugate of

[0095] During the dimensionality reduction process, the dimensionality reduction ratio is dynamically adjusted according to the type of data change. For large change segments, 25% to 30% of the original data dimensions are retained; for medium change segments, 15% to 20% of the dimensions are retained; and for small change segments, only 5% to 10% of the dimensions are retained. The purpose of this step is to reduce the data volume through dimensionality reduction technology while ensuring the effectiveness of the data, thereby reducing energy consumption for subsequent transmission.

[0096] The specific implementation of step S07 is to input the optimized data segment into the pre-trained OceanFormer model for feature extraction and compression. The OceanFormer model is a deep neural network based on the Transformer architecture, consisting of two main parts: an encoder and a decoder. The encoder is composed of six layers of multi-head self-attention layers and feed-forward neural networks, and each layer of the multi-head attention mechanism contains eight attention heads. The multi-head self-attention mechanism of the model can be expressed as: ; In the formula, , , are the query matrix, key matrix, and value matrix respectively; is the dimension of the key vector; is the sparse mask matrix, used to control the sparsity of the attention weights, , when it means masking the attention weight at this position.

[0097] The positional encoding calculation of the OceanFormer model is: ; ; In the formula, is the positional encoding value at position with dimension ; is the feature dimension of the model, with a value of 512.

[0098] In view of the periodic characteristics of ocean data, frequency-enhanced positional encoding is introduced: ; In the formula, is the frequency-enhanced positional encoding; is the standard positional encoding; is the data-dominant frequency; is the initial phase, uniformly distributed within the range of ; is the enhancement coefficient, with a value range of 0.1 to 0.5.

[0099] The dynamic adjustment formula for the number of attention heads is as follows: ; In the formula, is the number of attention heads actually used; is the basic number of attention heads, with a value of 8; is the current time window length (in minutes); and are the minimum and maximum time window lengths, which are 5 and 60 respectively; is the numerical encoding of the data change type. The large change segment is 2, the medium change segment is 1, and the small change segment is 0; and are the minimum and maximum data change type encodings, which are 0 and 2 respectively; and are adjustment coefficients, with values of 0.3 and 0.2 respectively.

[0100] The adaptive adjustment formula for attention sparsity is as follows: ; In the formula, is the sparsity, expressed as the proportion of masked elements in the attention matrix; is the basic sparsity, with a value of 0.3; and are adjustment coefficients, with values of 0.1 and 0.15 respectively.

[0101] The purpose of this step is to use the deep learning model to extract the deep features of ocean data and achieve a high compression ratio data representation, laying a foundation for low-power transmission.

[0102] The specific implementation of step S08 is to combine the processed data to form a data packet to be transmitted. First, integrate the replaced unique identifier with the data segment processed by the OceanFormer model. The identifier data is placed at the front end of the data packet, and the compressed data is placed at the back end, with a delimiter marker set between them. Then, add the data packet header information, including key information such as data change type, time window length, data acquisition timestamp, data quality marker, checksum, etc. The data packet header structure can be represented as a bit field: , ; In the formula, is the data change type field, occupying 2 bits. The values 00, 01, and 10 represent small change, medium change, and large change respectively; is the time window length field, occupying 6 bits, and can represent window lengths from 0 to 63 minutes; It is a timestamp field, occupying 32 bits and adopting the UNIX timestamp format; It is a data quality flag field, occupying 8 bits. The upper 4 bits represent the data integrity level, and the lower 4 bits represent the data accuracy level; It is a checksum field, occupying 16 bits and generated by the CRC16 algorithm.

[0103] The calculation formula for the cyclic redundancy checksum is: ; In the formula, is the 16-bit cyclic redundancy checksum; is the data to be checked; is the generating polynomial, adopting the CRC-16-CCITT standard, , and the corresponding hexadecimal value is 0x1021.

[0104] The purpose of this step is to encapsulate the processed data in a standard format to ensure that the receiving end can correctly parse and restore the data.

[0105] The specific implementation of step S09 is to send the data packet to be transmitted to the receiving terminal through a low-power transmission protocol. First, evaluate the current network environment and select the most suitable low-power transmission protocol. In the nearshore area, the LoRaWAN protocol is preferred. This protocol operates in an unlicensed frequency band and has the characteristics of long distance and low power consumption, with a transmission distance of up to 10 kilometers. In the open sea area, a satellite communication protocol such as the Iridium short message service is preferred. Although this service has a relatively high power consumption, it covers the global sea area. Before transmission, the data packet is fragmented, and the size of each fragment is dynamically adjusted according to the maximum transmission unit of the selected protocol, generally controlled between 100 and 250 bytes. The transmission adopts an adaptive power control strategy, and its calculation formula is: ; In the formula, is the actual transmission power (dBm); and are the minimum and maximum transmission powers respectively, and their value ranges are 10 - 14 dBm and 20 - 30 dBm respectively; is the current channel signal-to-noise ratio (dB); and are the minimum and maximum signal-to-noise ratio thresholds respectively, with values of 5 dB and 25 dB respectively; is the adjustment exponent, and its value range is 0.5 - 2.0. When , ; when , .

[0106] After receiving a data packet, the receiving terminal first verifies its integrity, and then identifies the data characteristics based on the information in the packet header. For a data packet containing a unique identifier, the corresponding classical data segment is extracted from the local database; for a data segment processed by the OceanFormer model, it is reconstructed through the decoder of the OceanFormer model. The decoder uses the cross-attention mechanism to fuse the information output by the encoder with the information of the existing classical data segments in the receiving terminal, and ensures the stability of the model through residual connections and layer normalization, and finally reconstructs the complete ocean data. The purpose of this step is to achieve efficient data transmission through low-power transmission technology, and complete data reconstruction at the receiving end to realize the entire energy-saving transmission process.

[0107] To better understand and implement the present invention, the following provides an embodiment 2 of a specific application scenario of the present invention: A marine research team deployed a set of buoy monitoring networks in the northern waters of a certain sea area. The network consists of 20 buoy stations, and each buoy is equipped with various sensors such as temperature, salinity, dissolved oxygen, pH value, and chlorophyll concentration, which are responsible for real-time monitoring of changes in marine environmental parameters. Since the buoy location is far from land, the energy supply mainly relies on the combination of solar panels and storage batteries. The energy is limited and the replenishment is difficult. The traditional data transmission method consumes high energy, seriously affecting the continuous working time of the buoy. The research team decided to apply the energy-saving transmission method of the present invention to solve this problem.

[0108] First, the research team conducted a comprehensive analysis of the collected marine data, and used the AdaClust clustering algorithm combined with the HyperLogLog technology to conduct massive clustering and grouping of the marine data. Taking the temperature data as an example, the monitoring data in the past 3 months totaled about 4.35 million. Through the analysis of the AdaClust clustering algorithm, the calculation result of the optimal number of clustering centers was 17. The clustering analysis shows that the temperature data changes violently during the typhoon passing period, belonging to the large change segment, the temperature changes within the daily tidal cycle belong to the medium change segment, and the temperature changes under most stable weather conditions belong to the small change segment. According to the calculation results of the change amplitude, as shown in Table 1: Table 1 Statistical characteristics of temperature data changes in different sea area environments

[0109] For marine data of different change types, the research team set corresponding time window lengths. For data in the large change segment, the time window length was set to 8 minutes; for data in the medium change segment, the time window length was set to 20 minutes; for data in the small change segment, the time window length was set to 50 minutes. In actual applications, the system will also make adaptive adjustments according to the current sea area characteristics and seasonal changes. For example, during the typhoon-prone period in summer, the window length of the large change segment will be automatically shortened to 6 minutes to capture more detailed changes.

[0110] Based on the set time window, the system segmented the ocean data, forming multiple window data segments. Taking the temperature data of a buoy station on March 15th as an example, a total of 96 window data segments were formed on that day, including 15 large change segments, 42 medium change segments, and 39 small change segments. The system calculated the basic statistical characteristics of each window data segment, as shown in Table 2: Table 2 Statistical Characteristics of Window Data Segments of Different Change Types

[0111] The system extracted a set of classic data segments of the same type and length as the current window data segment from the database, and calculated the similarity between each window data segment and the classic data segments. The classic data segment library contains a total of 583 classic temperature data patterns, which were extracted through long-term analysis of historical data. For the window data segment T0315-001, the system calculated its similarity with all classic patterns of small change segments, as shown in Table 3: Table 3 Calculation Results of the Similarity between Window Data Segment T0315-001 and Classic Data Segments

[0112] A similarity threshold of 0.90 was set. When the similarity between the window data segment and the classic data segment is higher than this threshold, the system replaces the original window data segment with the unique identifier of the classic data segment. According to the calculation results in Table 3, the comprehensive similarity between the window data segment T0315-001 and the classic data segment TP-S-042 is 0.939, exceeding the set threshold. Therefore, the original data can be replaced with the unique identifier of TP-S-042. Similarly, among the 96 window data segments on that day, 67 window data segments found classic data segments with similarity higher than the threshold and were replaced, with a replacement rate of 69.8%.

[0113] For the 29 window data segments with similarity lower than the threshold, the system calls the ocean data optimization function for data dimensionality reduction processing. Taking the large change segment window data segment T0315-064 as an example, the original number of data points is 240 (sampled every 2 seconds within an 8-minute window). After dimensionality reduction through the locally linear embedding algorithm, 27.5% of the dimensions of the original data are retained, that is, 66 data points. For medium change segments and small change segments, 17.8% and 8.5% of the dimensions are retained respectively. Table 4 shows the effect of the dimensionality reduction processing: Table 4 Effect of Dimensionality Reduction Processing for Data of Different Change Types

[0114] The downsampled data segments are input into the pre-trained OceanFormer model for feature extraction and compression. The OceanFormer model is a deep neural network based on the Transformer architecture, consisting of an encoder and a decoder. For large-variation segment data, the number of attention heads is increased to 12; for medium-variation segment data, the default 8 attention heads are maintained; for small-variation segment data, it is reduced to 4 attention heads. Table 5 shows the dynamic adjustment of the attention mechanism parameters: Table 5 Dynamic Adjustment of the Attention Mechanism Parameters of the OceanFormer Model

[0115] After being processed by the OceanFormer model, the data is further compressed, and the size of the finally formed data packet is significantly reduced. Table 6 shows the data volume changes at different processing stages: Table 6 Statistical Data Volume Changes at Each Stage of Data Processing

[0116] The system combines the processed data to form data packets to be transmitted. The header information of the data packet adopts a compact design, with a size of 64 bytes, including key information such as data change type, time window length, data acquisition timestamp, data quality flag, checksum, etc. For the 96 data packets formed on the same day, the total data volume is 14.2KB, while the original data volume is 335.6KB, and the compression ratio reaches 23.6:1.

[0117] The data packets to be transmitted are sent to the receiving terminal through the LoRaWAN low-power transmission protocol. The transmission adopts an adaptive power control strategy, dynamically adjusting the transmission power according to the channel quality, saving energy to the greatest extent while ensuring the transmission success rate. Table 7 shows the relationship between the transmission power and the distance: Table 7 Transmission Power Adjustment under Different Distance Conditions

[0118] After receiving the data packets, the receiving terminal first verifies the integrity, and then identifies the data features according to the header information of the data packets. For the data packets containing unique identifiers, the corresponding classic data segments are extracted from the local database; for the data segments processed by the OceanFormer model, they are reconstructed through the decoder of the OceanFormer model. The average error between the reconstructed data and the original data is only 0.028°C, meeting the accuracy requirements of ocean scientific research.

[0119] Traditional marine buoy data transmission methods mainly adopt a fixed-period, full-volume transmission strategy, that is, all monitored data are collected and transmitted at preset fixed time intervals (usually 15 minutes or 30 minutes), regardless of the data change characteristics, resulting in low energy utilization efficiency. Some improved methods adopt a threshold-triggered transmission mechanism, that is, data is transmitted only when the monitored parameter change exceeds a preset threshold. However, this method may not transmit data for a long time when the data changes slowly, resulting in data loss and affecting the continuity of research. There are also some methods that reduce the amount of transmitted data through simple data compression algorithms, but the compression rate is limited and key information may be lost.

[0120] Compared with traditional methods, the energy-saving transmission method of the present invention has brought significant improvements: First, the marine data is classified by an adaptive clustering algorithm, and different time windows are set for different change types, realizing the intelligent adjustment of data collection and transmission strategies; Second, a data replacement mechanism based on similarity is introduced, and classic data patterns are used to greatly reduce the amount of transmitted data; Third, a deep learning model OceanFormer is used to extract the deep features of marine data, realizing a high compression ratio data representation; Fourth, through a low-power transmission protocol and an adaptive power control strategy, the energy consumption is minimized while ensuring the data transmission quality. Practical applications show that after adopting the energy-saving transmission method of the present invention, the average working time of the buoy device has been extended from the original 63 days to 185 days, an increase of 193.7%, the total amount of data transmitted has been reduced by 95.8%, and at the same time, the data reconstruction accuracy has been maintained within an acceptable range, meeting the needs of marine scientific research.

[0121] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 8, 9, and 10 below.

[0122] Table 8 Variable Explanation Table (Part 1)

[0123] Table 9 Variable Explanation Table (Part 2)

[0124] Table 10 Variable Explanation Table (Part 3)

[0125] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. An energy-saving transmission method for a buoy monitoring device, characterized in that, Including: Conduct a comprehensive analysis of the collected ocean data. Use a clustering algorithm combined with a cardinality estimation technique to perform massive clustering and grouping on the ocean data, and divide the ocean data into large-change segments, medium-change segments, and small-change segments. Set different time window lengths for different change segments. Based on the time window, perform segmented processing on the ocean data of different change segments to form multiple window data segments. Calculate the similarity between the window data segments and the classic data segments to generate a similarity matrix. Set a similarity threshold. When the similarity is higher than the threshold, use the unique identifier of the classic data segment to replace the original window data segment. For the window data segments with similarity lower than the threshold, call the ocean data optimization function to perform data dimensionality reduction processing. Input the optimized data segments into a pre-trained deep neural network model for feature extraction and compression. Combine the replaced unique identifier with the processed data segments to form a data packet to be transmitted, and add packet header information. Send the data packet to be transmitted to the receiving terminal through a low-power transmission protocol. The receiving terminal restores the data and reconstructs the complete ocean data according to the packet header information.

2. The energy-saving transmission method of the buoy monitoring device according to claim 1, characterized in that, The clustering algorithm is the AdaClust clustering algorithm, and the cardinality estimation technique is the HyperLogLog technique.

3. The energy-saving transmission method of the buoy monitoring device according to claim 2, wherein The large-change segment refers to the data interval where the change amplitude of the ocean data exceeds the preset first threshold within a unit time. The medium-change segment refers to the data interval where the change amplitude of the ocean data is between the preset first threshold and the preset second threshold within a unit time. The small-change segment refers to the data interval where the change amplitude of the ocean data is lower than the preset second threshold within a unit time.

4. The energy-saving transmission method of the buoy monitoring device according to claim 3, characterized in that Set the time window length for the large-change segment to be the first length, the time window length for the medium-change segment to be the second length, and the time window length for the small-change segment to be the third length.

5. The energy-saving transmission method of the buoy monitoring device according to claim 4, characterized in that, The first length refers to a time window length of 5 to 10 minutes, which is used to capture the characteristics of the ocean data in the large-change segment. The second length refers to a time window length of 15 to 30 minutes, which is used to capture the characteristics of the ocean data in the medium-change segment. The third length refers to a time window length of 45 to 60 minutes, which is used to process the ocean data in the small-change segment.

6. The energy-saving transmission method of the buoy monitoring device according to claim 5, characterized in that, Each element in the similarity matrix represents the similarity value between a pair of window data segments and the classic data segments.

7. The energy-saving transmission method of the buoy monitoring device according to claim 6, characterized in that, The ocean data optimization function is used to perform dimensionality reduction and key feature extraction processing on the window data segments. The input includes the window data segments, time window length, data change type, similarity value, and frequency characteristics, and the output is the optimized data segment that retains key features after dimensionality reduction.

8. The energy-saving transmission method of the buoy monitoring device according to claim 7, characterized in that, The deep neural network model is the OceanFormer model, and the parameters of the multi-head attention mechanism in the OceanFormer model are dynamically adjusted according to the time window length, data change type, and similarity threshold.

9. The energy-saving transmission method of the buoy monitoring device according to claim 8, characterized in that, The specific structure of the OceanFormer model is a deep neural network based on the Transformer architecture, which includes two major parts: an encoder and a decoder. The encoder consists of six layers of multi-head self-attention layers and a feed-forward neural network. Each layer of the multi-head attention mechanism contains eight attention heads. The number of attention heads is dynamically adjusted according to the time window length, and the attention weight matrix adaptively adjusts the sparsity according to the data change type.

10. The energy-saving transmission method of the buoy monitoring device according to claim 9, characterized in that, The decoder uses the cross-attention mechanism to fuse the information output by the encoder with the information of the existing classical data segments at the receiving terminal, and ensures the stability of the model through residual connections and layer normalization. The end of the OceanFormer model uses a fully connected layer to map the features to the compression space.