A deep-sea buoy multi-physical field data compression, cleaning and collaborative fusion system

Through the collaborative work of the deep-sea wireless sensing gateway, edge compression and server cleaning, and multi-physics fusion module, the problems of data pollution and inconsistent multi-parameter observations in deep-sea marine environmental monitoring have been solved, achieving efficient and reliable data acquisition and fusion, and generating a high-precision marine data field.

CN122496790APending Publication Date: 2026-07-31CCCC SECOND HARBOR ENGINEERING CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CCCC SECOND HARBOR ENGINEERING CO LTD
Filing Date
2026-04-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing marine environmental monitoring systems face problems such as data contamination and inconsistent multi-parameter observations in deep-sea areas. Traditional data cleaning methods cannot meet real-time requirements, and single-parameter observations are insufficient to comprehensively depict the overall state of the ocean. Existing fusion methods lack fluid dynamic constraints, resulting in insufficient data quality and continuity.

Method used

The system employs a deep-sea wireless sensing gateway hardware platform for multi-physical quantity data acquisition, an edge-end data compression module for structured compression, a server-side data noise cleaning module for intelligent noise identification and suppression, and a multi-physical field collaborative data fusion module for spatiotemporal registration and dynamic assimilation fusion to generate a spatiotemporally continuous optimal ocean data field.

Benefits of technology

It has achieved continuous and stable acquisition of multi-physical data in harsh marine environments, reduced data transmission pressure, improved data quality and channel utilization, and generated a logically self-consistent high-precision marine data field, providing reliable data support for deep-sea scientific research and disaster early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496790A_ABST
    Figure CN122496790A_ABST
Patent Text Reader

Abstract

This invention provides a system for compressing, cleaning, and collaboratively fusing multiphysics field data from deep-sea buoys, relating to the field of data management technology. The system includes a deep-sea wireless sensing gateway hardware platform, an edge-end data compression module, a server-end data noise cleaning module, and a multiphysics field collaborative data fusion module. By constructing an "edge-cloud collaborative" system architecture, the hardware acquisition and edge compression at the deep-sea buoy end are decoupled from the data cleaning and multiphysics field fusion at the shore-based or cloud-based end into collaborative functional modules, forming a closed-loop chain from data perception to the generation of high-value information. This invention overcomes the limitations of single-parameter observation and provides high-precision, high-reliability core data support for deep-sea scientific research, engineering applications, and disaster early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and more specifically, to a system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys. Background Technology

[0002] As the most resource-rich and environmentally complex region on Earth, the acquisition of dynamic environmental information from the deep sea has become a crucial support for safeguarding maritime rights, resource development, and climate change research. Constructing an all-weather, high-precision, three-dimensional deep-sea observation network to accurately grasp data on multiple physical quantities such as hydrostatic pressure, wave characteristics, and stratified current velocities is an important foundation for achieving ocean transparency. However, current marine environmental monitoring systems still face the following technical bottlenecks in practical engineering applications: Ocean buoys, drifting on the sea surface for extended periods, are inevitably affected by factors such as wave action, biofouling, and electrical noise from sensors. Consequently, the raw data they collect inevitably contains background Gaussian noise and non-physical impulse outliers. Traditional data cleaning methods, such as moving average filtering or median filtering, are essentially low-pass filters. While smoothing noise, they often obscure high-frequency signal details with real physical meaning and cause significant phase lag in the data, making them unsuitable for the real-time requirements of marine disaster prevention and mitigation. Standard Kalman filtering, relying on a fixed noise covariance parameter, struggles to adapt to dynamic changes in sea state levels and is prone to divergence when subjected to large outlier impacts, resulting in poor robustness in data cleaning.

[0003] The deep-sea dynamic environment is a complex system composed of pressure, wave, and velocity fields coupled together, making it difficult to comprehensively depict the overall ocean state from observational data of a single parameter. However, current ocean data applications largely remain in a discrete state of "single-point, single-parameter" analysis. Data such as hydrostatic pressure, wave characteristics, and stratified current velocities often originate from different sensors or observation platforms, resulting in "island effect" problems such as inconsistent spatiotemporal benchmarks, asynchronous sampling frequencies, and discontinuous data coverage. Existing fusion methods are mostly based on purely mathematical statistical models such as Kriging interpolation or inverse distance weighting, lacking the constraints of fluid dynamic physical mechanisms. The resulting fused fields often violate the laws of mass or momentum conservation, leading to scientific inconsistency and failing to accurately reflect the overall evolution of the deep-sea dynamic environment. Summary of the Invention

[0004] The problem solved by this invention is one or more of the aforementioned related technical problems.

[0005] To address the aforementioned issues, this invention provides a system for compressing, cleaning, and collaboratively fusing multi-physics field data from deep-sea buoys, comprising: a deep-sea wireless sensing gateway hardware platform, an edge-end data compression module, a server-end data noise cleaning module, and a multi-physics field collaborative data fusion module. The deep-sea wireless sensing gateway hardware platform is deployed on deep-sea buoys or underwater moorings to collect raw data of multiple physical quantities in the deep-sea environment. The edge data compression module is used to perform structured compression and frame-level encapsulation of the original data of the multiple physical quantities based on physical constraints to generate compressed data packets; The server-side data noise cleaning module is used to receive and decompress the compressed data packet to restore the original observation data, and to identify and suppress background noise and non-physical outliers in the original observation data in real time based on the improved Kalman filter algorithm, and output the cleaned observation data. The multiphysics collaborative data fusion module is used to perform spatiotemporal registration and dynamic assimilation fusion on the cleaned observation data based on a preset physical constraint system, so as to generate a spatiotemporally continuous optimal ocean data field.

[0006] Optionally, the edge data compression module is specifically applied to: The original data of the multiple physical quantities are identified and labeled to obtain different types of multiple physical quantity data; Based on the different types of multi-physical quantity data, the corresponding multi-level prediction models are used to process them to obtain the corresponding final prediction values. A joint residual sequence is constructed based on the final predicted value and the corresponding observation value in the original data of the multi-physical quantities, and the joint residual sequence is divided and recombined at multiple scales. Based on the statistical distribution characteristics of the recombined joint residual sequence, a lossless encoding method is adaptively selected locally at the edge node of the buoy for encoding; and according to the preset fixed frame length constraint, the encoded data stream is organized and encapsulated at the frame level to generate the compressed data packet.

[0007] Optionally, the multi-level prediction model is constructed based on physical constraints and the correlation of multiple physical quantities. The step of processing according to the corresponding multi-level prediction model to obtain the corresponding final predicted value includes: Based on the first prediction level in the multi-level prediction model, a basic prediction value is generated according to the multi-physical quantity data, and the basic prediction value is corrected based on the second prediction level in the multi-level prediction model to generate the final prediction value of the current sampling point.

[0008] Optionally, the server-side data noise cleaning module is specifically applied to: The compressed data packet is decompressed to obtain the original observation data, and the original observation data is preprocessed to construct a standardized observation vector. A state-space model is constructed based on the physical characteristics of the marine dynamic environment. The observation vector and the optimal state estimate of the previous time step are used as inputs. Kalman filtering is performed to update the time and generate the state prediction value and prediction error covariance of the current sampling point. Based on the deviation between the current state prediction value and the corresponding observation value in the observation vector, the information of the current sampling point is obtained; and based on the corresponding prediction error covariance and the preset observation noise covariance, the theoretical standard deviation of the information is obtained. A strong radial constraint threshold is constructed based on the current sampling point's information and the theoretical standard deviation. A preset multiple of the theoretical standard deviation is used as a threshold. The current sampling point's information amplitude is compared with the threshold to obtain a comparison result. Kalman filtering measurement updates are performed based on the comparison result to obtain the optimal state estimate at the current time, which is used as the cleaned observation data.

[0009] Optionally, the step of updating the measurement using Kalman filtering based on the comparison result to obtain the optimal state estimate at the current time, as the cleaned observation data, includes: If the innovation amplitude of the current sampling point does not exceed the threshold, it is determined that the observation value corresponding to the current sampling point belongs to the background noise. Based on the Sage-Husa adaptive filtering algorithm, the system noise covariance and the observation noise covariance are iteratively corrected in real time with the innovation of the current sampling point as input, so as to obtain the updated current system noise covariance and the updated current observation noise covariance. A forgetting factor is introduced to weight the historical data. If the innovation amplitude of the current sampling point exceeds the threshold, it is determined that the observation value corresponding to the current sampling point belongs to a non-physical impulse field value. The update of the system noise covariance and the observation noise covariance is frozen, the system noise covariance and the observation noise covariance at the previous moment are kept unchanged, and the Kalman gain is suppressed. Based on the updated or frozen current system noise covariance, current observation noise covariance, information of the current sampling point, and the state prediction value, Kalman filtering measurement updates are performed to obtain the optimal state estimate at the current moment, which is used as the cleaned observation data.

[0010] Optionally, the multiphysics collaborative data fusion module is specifically applied to: Different types of data in the cleaned observation data are preprocessed separately to obtain corresponding standardized ocean data; The standardized marine data is unified in terms of spatiotemporal reference, and a grid background field is generated through a unified spatiotemporal registration and grid construction process. A unified state vector is constructed based on the grid background field. Under the guidance of the preset physical constraint system, a unified dynamic data assimilation loop is executed to obtain the optimized unified state vector. The dynamic data assimilation loop is repeated until the current unified state vector satisfies the convergence condition, thereby obtaining the optimal ocean data field.

[0011] Optionally, the dynamic data assimilation loop includes: Using the current state value in the unified state vector as input, the state evolution is predicted based on the physical model through ensemble Kalman filtering to obtain the predicted state vector; then, the predicted state vector is updated according to the real-time observation data of the corresponding spatiotemporal nodes in the grid background field to obtain the updated state vector; the core parameters in the updated state vector are corrected for prediction errors through a long short-term memory network model to obtain the optimized unified state vector.

[0012] Optionally, the preset physical constraint system includes at least one of the following: hydrostatic equilibrium equation, simplified shallow water wave equation, geostrophic equilibrium equation, and mass conservation continuity equation.

[0013] Optionally, the adaptive selection of lossless coding method includes selecting an appropriate coding method from fixed-length bit-width coding, variable-length bit-width coding, or entropy coding based on at least one of the maximum amplitude, zero-value ratio, and dispersion index of the joint residual sequence, as well as the current computing resource or power consumption status of the edge node.

[0014] Optionally, the core parameters include wave height, wave velocity, and pressure.

[0015] The beneficial effects of the deep-sea buoy multiphysics data compression, cleaning, and collaborative fusion system of the present invention are: By constructing an "edge-cloud collaborative" system architecture, the hardware acquisition and edge compression at the deep-sea buoy end are decoupled from the data cleaning and multi-field fusion at the shore-based or cloud-based end, forming a closed-loop end-to-end system from data perception to high-value information generation. At the hardware level, the system ensures continuous and stable acquisition of raw data of multiple physical quantities in harsh marine environments through a highly reliable and low-power perception gateway. At the edge end, structured compression based on physical constraints significantly reduces the data transmission pressure for narrowband communications such as BeiDou satellites, improving channel utilization and reducing transmission power consumption. At the server end, an improved Kalman filter algorithm intelligently identifies and suppresses noise in the decompressed raw observation data, effectively solving the data contamination problem under complex sea conditions and improving data quality. Finally, by introducing multi-physics field collaborative fusion with a pre-defined physical constraint system, the cleaned multi-source heterogeneous data is integrated into a spatiotemporally continuous and logically self-consistent optimal marine data field, breaking through the limitations of single-parameter observation and providing high-precision and high-reliability core data support for deep-sea scientific research, engineering applications, and disaster early warning. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of a deep-sea buoy multiphysics data compression, cleaning, and collaborative fusion system according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the data processing flow of an edge-end data compression module according to an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the data processing flow of a server-side data noise cleaning module according to an embodiment of the present invention. Figure 4 This is a schematic diagram illustrating the data processing flow of a multiphysics collaborative data fusion module according to an embodiment of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0018] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0019] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0020] It should be noted that the one or more modifications mentioned in this invention are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as one or more.

[0021] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0022] Currently, deep-sea observation equipment is typically deployed in open waters far from the mainland, exposed to extreme environments of high salt spray corrosion, strong ultraviolet radiation, and severe mechanical vibration. This places extremely high demands on the high reliability and low power consumption survivability of the sensing hardware. Most existing data acquisition terminals lack sophisticated time-sharing power supply strategies and milliwatt-level sleep mechanisms, resulting in high static power consumption. Furthermore, they lack specific anti-corrosion designs and multi-level lightning protection for the marine environment, making them highly susceptible to downtime and damage during prolonged periods of rain or thunderstorms, severely impacting data continuity.

[0023] In deep-sea areas far from shore-based cellular network coverage, data backhaul primarily relies on BeiDou satellite short message communication. Compared to terrestrial broadband, BeiDou communication is characterized by extremely narrow bandwidth, strictly limited single message length, controlled communication frequency, and high costs. Existing data transmission schemes mostly employ general lossless compression algorithms or simple data sampling, failing to fully consider the continuity, periodicity, and correlation of physical properties such as hydrostatic pressure, waves, and current velocity, resulting in low compression ratios and inefficient payload utilization. Furthermore, the lack of dedicated optimization for the fixed frame length of BeiDou short messages often leads to a mismatch between the compressed data length and the message payload, resulting in a large number of padding bits and wasting valuable satellite channel resources.

[0024] In summary, there is an urgent need in this field for an edge-cloud collaborative solution that can comprehensively address hardware reliability, transmission efficiency, data quality, and multi-field integration, in order to overcome existing technological bottlenecks and provide solid technical support for the development and security of deep-sea resources.

[0025] In response to the problems existing in the aforementioned related technologies, such as Figure 1 As shown in the figure, the present invention provides a deep-sea buoy multi-physics field data compression, cleaning and collaborative fusion system, including: a deep-sea wireless sensing gateway hardware platform, an edge data compression module, a server data noise cleaning module and a multi-physics field collaborative data fusion module; The deep-sea wireless sensing gateway hardware platform is deployed on deep-sea buoys or underwater moorings to collect raw data of multiple physical quantities in the deep-sea environment.

[0026] Specifically, the deep-sea wireless sensing gateway hardware platform, serving as the data acquisition front-end of the entire system, is deployed on buoys or submersibles in deep-sea areas. This platform integrates multi-protocol sensor interfaces and simultaneously mounts different types of heterogeneous sensors, including hydrostatic pressure gauges for measuring water pressure, wave sensors for acquiring parameters such as wave height and wave velocity, and acoustic Doppler current profilers for observing seawater current velocities at different depths.

[0027] During actual operation, the platform automatically wakes up according to a preset acquisition cycle, simultaneously or sequentially acquiring corresponding raw signals through various sensor interfaces: the hydrostatic pressure gauge acquires pressure time-series values ​​at different depths in the deep sea, the wave sensor acquires the undulation intensity and propagation rate parameters of sea surface waves, and the current profiler acquires the horizontal and vertical velocity components of each standard depth layer from the surface to the deep sea. The acquired raw data of multiple physical quantities are collected in the form of digital signals and sent to the platform. After initial timestamping, the data is temporarily stored in the local storage unit, forming a multi-source heterogeneous data set to be processed for subsequent use by the edge data compression module. The entire acquisition process operates automatically under unattended conditions. The platform's high-protection structural design adapts to the deep-sea environment of high salt spray, high humidity and heat, and severe vibration, ensuring continuous and stable acquisition of raw data.

[0028] By deploying a deep-sea wireless sensing gateway hardware platform, a one-stop, highly reliable acquisition of raw data on multiple physical quantities such as hydrostatic pressure, wave characteristics, and stratified current velocity was achieved in extremely harsh marine environments. This solved the data fragmentation problem caused by the poor environmental adaptability and single interface of traditional acquisition terminals, which could not simultaneously mount multiple heterogeneous sensors. As the physical carrier of edge computing and data acquisition, the platform decouples the acquisition of raw data from subsequent compression processing, providing a stable and unified data source for the entire system. This lays the foundation for data perception and the generation of high-value information, significantly improving the engineering feasibility and data integrity of long-term unattended deep-sea observation operations.

[0029] The edge data compression module is used to perform structured compression and frame-level encapsulation of the original data of the multiple physical quantities based on physical constraints to generate compressed data packets.

[0030] Specifically, the edge data compression module runs within the deep-sea wireless sensing gateway hardware platform, serving as an intermediate processing link connecting data acquisition and remote transmission. Once the gateway hardware platform completes the acquisition of raw data for multiple physical quantities and temporarily stores it locally, this module is triggered to perform structured compression processing on the raw data based on physical constraints.

[0031] Taking a typical data acquisition task as an example, the module first acquires a raw data set containing hydrostatic pressure time series values, wave characteristic parameters (such as significant wave height and wave velocity), and layered velocity profiles (including u, v, and w components). Addressing the continuous nature of the hydrostatic pressure data, the periodicity of the wave data, and the vertical correlation of the velocity data, the module introduces corresponding physical constraints. For example, it uses the continuous variation law of hydrostatic pressure to constrain the prediction direction of the pressure data, the conservation of wave energy to constrain the joint variation trend of wave height and wave velocity, and the fluid continuity of adjacent depth layers to constrain the structural morphology of the velocity profile. Guided by these physical constraints, the module performs structured analysis and removes redundant information from the raw data, extracting key information that characterizes the core changing features of the raw data. Subsequently, the module organizes and encapsulates the processed key information at the frame level according to the fixed frame length requirements of BeiDou satellite short messages, forming a compressed data packet adapted for satellite channel transmission, which is stored in the transmission buffer awaiting access by the communication module. The entire compression process is completed independently locally at the buoy edge node, without interaction with a remote server.

[0032] By deploying this data compression module within the edge node of the buoy, this application achieves efficient and lossless compression and adaptive encapsulation of raw data containing multiple physical quantities under the strong constraint of BeiDou satellite communication bandwidth. It fully utilizes the inherent physical laws and structural redundancy of marine data, significantly improving compression efficiency and payload utilization. Compared with general compression algorithms, it can increase the amount of effective information carried in a single communication by several times. At the same time, the compression process is completed independently locally at the data source end, without relying on remote computing resources, avoiding the additional power consumption caused by frequent communication, and effectively extending the endurance of the buoy system in energy-constrained environments. In addition, through a special encapsulation design for the fixed frame length of BeiDou short messages, it eliminates the padding waste caused by data length mismatch, reduces the risk of packet loss and retransmission, and provides key technical support for the reliable backhaul of deep-sea monitoring data.

[0033] The server-side data noise cleaning module is used to receive and decompress the compressed data packet to restore the original observation data, and to identify and suppress background noise and non-physical outliers in the original observation data in real time based on the improved Kalman filter algorithm, and output the cleaned observation data.

[0034] Specifically, the server-side data noise cleaning module is deployed in a shore-based data center or cloud server, serving as an intermediate processing link connecting data transmission and data fusion. When the compressed data packets sent by the edge terminal are transmitted to the server via satellite or cellular network, the module first decompresses the received compressed data packets to completely restore the original observation data collected by the buoy, including hydrostatic pressure time series data, wave characteristic parameter data, and stratified velocity profile data.

[0035] Due to the complexity of the deep-sea environment, these restored raw observation data typically contain two types of interference: one is background Gaussian noise caused by wind and wave disturbances, sensor electronic noise, etc., which manifests as small-amplitude, high-frequency random fluctuations; the other is non-physical pulse outliers caused by biological impacts, sensor momentary malfunctions, etc., which manifest as large-amplitude, abrupt abnormal peaks. To address this mixed noise scenario, the module employs an improved Kalman filter algorithm: the algorithm first constructs a state-space model based on the physical characteristics of the marine dynamic environment to theoretically predict the current data state; then, it compares and analyzes the theoretical predictions with the actual observations. Through the algorithm's internal adaptive discrimination mechanism, it intelligently distinguishes whether the current deviation originates from normal background noise or abnormal pulse outliers. For background noise, the algorithm dynamically adjusts the filtering parameters to adapt to the real-time changing noise level; for pulse outliers, the algorithm automatically masks their influence to prevent dirty data from contaminating the filtering process. After the above processing, the module outputs cleaned observation data that has filtered out noise interference while fully preserving the true physical trends, for subsequent use by the multiphysics collaborative data fusion module. The entire process is completed online in real time on the server side, supporting batch processing of continuous data streams.

[0036] By deploying a server-side data noise cleaning module, high-precision and intelligent cleaning of raw observation data transmitted from deep sea was achieved, effectively solving the data pollution problem caused by the coexistence of background noise and sudden outliers under complex sea conditions. The improved Kalman filter algorithm used in the module can accurately identify and suppress non-physical impulse outliers while filtering out high-frequency random noise, avoiding parameter divergence or data distortion caused by outlier interference in traditional filtering methods, and significantly improving the signal-to-noise ratio and physical fidelity of the output data. The cleaned observation data not only retains the integrity of real physical characteristics such as tidal cycle and gradual change in current velocity, but also eliminates the impact of abnormal changes on subsequent fusion analysis, providing a reliable data foundation for constructing a high-confidence deep sea environmental data field, and significantly enhancing the data quality control capability of the entire system in complex marine environments.

[0037] The multiphysics collaborative data fusion module is used to perform spatiotemporal registration and dynamic assimilation fusion on the cleaned observation data based on a preset physical constraint system, so as to generate a spatiotemporally continuous optimal ocean data field.

[0038] Specifically, the multiphysics collaborative data fusion module is deployed on a shore-based data center or cloud server, serving as the core backend processing component of the entire system. This module receives cleaned observation data from the server-side data noise removal module, including noise-suppressed hydrostatic pressure time-series data, wave characteristic parameter data, and layered velocity profile data. Since these three types of data originate from different sensors and observation platforms, they exhibit device clock deviations in time reference, discrete point-based spatial distribution, and different sampling frequencies and data formats. Upon startup, the fusion module first performs spatiotemporal reference unification processing on the three types of data: mapping all data timestamps to UTC standard time to eliminate clock deviations; unifying the spatial location of all data to the WGS-84 coordinate system, and setting a three-dimensional grid (horizontal latitude and longitude grid superimposed with a vertical depth layer) based on the observation area. On this basis, the module introduces a pre-defined physical constraint system as guiding rules for the fusion process. This system may include physical equations and correlations describing the inherent laws of the marine dynamic environment, such as the positive correlation between pressure and depth, the dynamic correlation between wave height and wave velocity, and the continuity characteristics of the vertical distribution of current velocity.

[0039] Guided by the physical constraints, the module performs spatiotemporal registration, filling the grid nodes with discretely distributed observation data using interpolation methods that conform to physical laws, forming a preliminary continuous background field. Subsequently, a dynamic assimilation and fusion operation is performed, coupling the background field data with the prediction results of the physical model. By continuously adjusting the values ​​of each grid node, the module ensures that both the characteristics of the observation data and the constraints of physical laws are met. After multiple iterations and optimizations, the module ultimately generates an optimal ocean data field that covers the entire observation area, includes all grid nodes, and evolves continuously over time. This data field is typically stored in a three-dimensional or four-dimensional grid format, with each grid node containing fused values ​​of core physical quantities such as hydrostatic pressure, wave characteristic parameters, and stratified current velocity components.

[0040] By deploying a multi-physics collaborative data fusion module, this application integrates discrete, heterogeneous, cleaned observation data into a spatiotemporally continuous and physically self-consistent optimal ocean data field, completely solving the "island effect" problem caused by inconsistent spatiotemporal benchmarks in traditional methods. The pre-set physical constraint system introduced by the module serves as the underlying guiding rule for the fusion process, ensuring that the generated data field strictly conforms to the inherent laws of ocean dynamics while satisfying the characteristics of the observation data. For example, pressure increases with depth, wave height and wave speed satisfy energy conservation, and the current velocity profile maintains continuity, thus effectively eliminating the physical logic contradictions often caused by pure mathematical interpolation methods. The final optimal ocean data field not only fills the numerical gaps in the observation blank areas, but also significantly improves the inversion accuracy of each parameter through the collaborative constraints of multi-physics fields. It provides the first high-value data product that fully characterizes the coupled evolution of pressure field, wave field, and current velocity field for deep-sea scientific research, engineering applications, and disaster early warning, realizing a qualitative leap from "data acquisition" to "information generation".

[0041] In this embodiment, the deep-sea buoy multiphysics field data compression, cleaning, and collaborative fusion system constructs an "end-cloud collaborative" system architecture. This decouples the hardware acquisition and edge compression at the deep-sea buoy end from the data cleaning and multiphysics field fusion at the shore-based or cloud-based end into collaborative functional modules, forming a closed-loop chain from data perception to high-value information generation. At the hardware level, the system ensures continuous and stable acquisition of raw multiphysics quantity data in harsh marine environments through a highly reliable and low-power perception gateway. At the edge end, structured compression based on physical constraints significantly reduces the data transmission pressure for narrowband communications such as BeiDou satellites, improving channel utilization and reducing transmission power consumption. At the server end, an improved Kalman filter algorithm intelligently identifies and suppresses noise in the decompressed raw observation data, effectively solving the data contamination problem under complex sea conditions and improving data quality. Finally, by introducing a multiphysics field collaborative fusion system with a preset physical constraint framework, the cleaned multi-source heterogeneous data is integrated into a spatiotemporally continuous and logically self-consistent optimal marine data field, overcoming the limitations of single-parameter observation and providing high-precision and high-reliability core data support for deep-sea scientific research, engineering applications, and disaster early warning.

[0042] like Figure 2 As shown, optionally, the edge data compression module is specifically applied to: The original data of the multiple physical quantities are identified and labeled to obtain different types of multiple physical quantity data; Based on the different types of multi-physical quantity data, the corresponding multi-level prediction models are used to process them to obtain the corresponding final prediction values. A joint residual sequence is constructed based on the final predicted value and the corresponding observation value in the original data of the multi-physical quantities, and the joint residual sequence is divided and recombined at multiple scales. Based on the statistical distribution characteristics of the recombined joint residual sequence, a lossless encoding method is adaptively selected locally at the edge node of the buoy for encoding; and according to the preset fixed frame length constraint, the encoded data stream is organized and encapsulated at the frame level to generate the compressed data packet.

[0043] Optionally, the multi-level prediction model is constructed based on physical constraints and the correlation of multiple physical quantities. The step of processing according to the corresponding multi-level prediction model to obtain the corresponding final predicted value includes: Based on the first prediction level in the multi-level prediction model, a basic prediction value is generated according to the multi-physical quantity data, and the basic prediction value is corrected based on the second prediction level in the multi-level prediction model to generate the final prediction value of the current sampling point.

[0044] In some embodiments, after the gateway hardware platform (deep-sea wireless sensing gateway hardware platform) completes a round of multi-physical quantity raw data acquisition, the edge data compression module is triggered to start and execute the following processing flow: Step 1: Data identification and labeling.

[0045] The module first acquires the raw data set (i.e., raw data of multiple physical quantities), which contains time-series data of various physical quantity types (data of different types of multiple physical quantities). Based on the configuration information of each channel sensor and the identifier in the data frame header, the module automatically identifies and classifies the raw data: the numerical sequence collected by the pressure sensor is labeled as "hydrostatic pressure data", the effective wave height, wave velocity and other parameters output by the wave sensor are labeled as "wave characteristic parameter data", and the velocity components of each depth layer returned by the acoustic Doppler current profiler are labeled as "layered velocity data".

[0046] Taking a typical data collection as an example, suppose a buoy collects a set of data at time t: pressure value of 2.35 MPa, significant wave height of 1.28 m, wave velocity of 3.42 m / s, and u and v velocity components at 5 depth layers. The module will automatically add a corresponding type identifier to each value to form a structured data set, laying the foundation for subsequent differential processing.

[0047] Step 2: Multi-level prediction model processing.

[0048] For different types of labeled data, the module calls the corresponding multi-level prediction model for prediction processing. Taking hydrostatic pressure data as an example, the processing of the multi-level prediction model is as follows: The first prediction level is based on the changing trend of historical pressure data (such as the periodic gradual change caused by tides), and uses a sliding window statistical model to generate the basic prediction value at the current sampling time, assumed to be 2.34 MPa; the second prediction level is based on the local fluctuation characteristics of the most recent sampling points (such as instantaneous pressure pulsations caused by waves), and uses a differential prediction model to correct the basic prediction value, capturing the small pressure changes caused by the superposition of sea surface waves, and generating the final prediction value of 2.36 MPa.

[0049] Similarly, for wave characteristic parameter data, the first prediction level generates basic predicted values ​​for wave height and wave velocity based on the overall stability of the sea state, while the second prediction level introduces a historical period similarity correction mechanism to reflect wave characteristics within a short timescale. For stratified current velocity data, the first prediction level constructs a vertical structure prediction by combining the current velocity distribution characteristics of different depth layers, while the second prediction level uses observation data from adjacent depth layers to collaboratively correct the prediction results. Through the collaborative processing of these two levels, the module generates a final predicted value with clear physical meaning and high accuracy for each sampling point of each type of data. In some specific embodiments, prediction models conforming to the physical characteristics of different types of multi-physical quantity data are constructed respectively. The prediction model includes a first prediction level model and a second prediction level model. The first prediction level model is used to characterize the long-term change trend of the data and generates basic prediction values ​​using a low-complexity model that satisfies physical continuity constraints (physical constraints). The second prediction level model is constructed based on the data change characteristics within a short time window and is used to capture local fluctuation characteristics and correct the prediction values ​​of the first prediction level. At the same time, correlation constraints between multiple physical quantities are introduced in the prediction process, including hydrostatic pressure continuity constraints, wave energy stability constraints, or adjacent depth layer flow velocity coordination constraints.

[0050] In other embodiments, the multi-level prediction model includes a first prediction level (long-term trend prediction) and a second prediction level (short-term fluctuation prediction). Both levels embed a correction mechanism based on physical constraints, and cross-physical quantity prediction correction is achieved through multi-physical quantity correlation constraints. It does not simply rely on linear / nonlinear fitting of the mathematical model. Physical constraint correction permeates the entire process of "basic prediction value generation - correction value calculation - final prediction value output." Specific level descriptions are provided, along with a clear explanation of the quantification embedding method and specific correction logic of physical constraints. (I) Core Quantitative Embedding Method of Physical Law Constraints: All ocean dynamic physical laws are not abstract constraints, but are embedded in the prediction model in a mathematical and parameterized way, becoming hard constraints on the predicted values. Specifically, they include three core forms: Constraint term embedding: Transform physical relationships into functional relationships / difference constraints between predicted values, which participate in the prediction calculation as additional constraint terms. If a predicted value violates the constraint term, it is directly corrected. Error correction factor embedding: Physical conservation / correlation relationships are transformed into prediction error correction factors to perform secondary correction on the initial prediction results and reduce prediction biases that have no physical meaning; Cross-variable feature input: The coupling relationship of multiple physical quantities is mapped to cross-variable feature input to the second prediction level, which enhances the model's ability to express the physical correlation structure and reduces the generation of non-physical prediction values ​​from the source.

[0051] (II) Correction of physical law constraints for the first prediction level (long-term trend prediction): The first prediction level employs a low-complexity physical continuous model (linear trend / sliding window statistical model) to generate basic predicted values ​​that reflect the long-term trend of data changes. The core of the physical law constraint correction in this process targets the inherent physical characteristics of a single physical quantity, ensuring that the basic predicted values ​​conform to the basic evolution law of that physical quantity and have no deviations that violate common sense physics. Specifically, the correction logic for the three types of core physical quantities is as follows: Hydrostatic pressure data: Embedded constraints on the continuity and monotonicity of hydrostatic pressure in water bodies; Physical laws: hydrostatic pressure increases with depth and changes continuously; the pressure changes caused by tides have a stable temporal continuity and no sudden fluctuations. Correction logic: The difference in pressure change between adjacent time steps is set as a constraint threshold. If the difference between the pressure baseline prediction value generated in the first level and the historical trend exceeds this threshold (indicating a non-physical abrupt change), the baseline prediction value is linearly corrected through a smoothing correction factor to bring it back to the range of physically continuous trends. At the same time, the prediction value is mapped to an equivalent water depth space for consistency verification, based on the functional relationship between static pressure and water depth. If the predicted pressure value corresponds to an unreasonable water depth (such as pressure decreasing with depth), the constraint correction mechanism is directly triggered, and the baseline prediction value is recalculated.

[0052] Wave characteristic data (wave height / wave speed): Embedded wave energy stability and spectral consistency constraints Physical laws: The long-term trend of waves (overall sea state) has energy stability, and there is a fixed dynamic relationship between wave height, wave speed, and period, and the dominant frequency of its spectral distribution remains consistent; Correction logic: Construct a wave energy stability range based on historical time windows. If the energy change rate corresponding to the basic predicted values ​​of wave height / wave speed at the first level exceeds this range, the predicted values ​​are proportionally adjusted and corrected. At the same time, the consistency of the main frequency of the wave spectrum is used as a constraint condition, and the prediction results are projected into the spectrum stability range to ensure that the basic predicted values ​​of wave height and wave speed meet the dynamic correlation (such as the wave speed increasing in a reasonable manner when the wave height increases).

[0053] Layered velocity data: Embedding vertical velocity gradient continuity and shear relationship constraints Physical laws: In the long-term trend of deep-sea stratified current velocity, there are stable velocity gradients and shear relationships between different depth layers, without any abnormal amplification of interlayer velocity gradients; Correction logic: Calculate the gradient difference between the predicted velocity values ​​of adjacent depth layers. If the basic predicted velocity value of the first layer causes the interlayer velocity gradient to exceed the physically reasonable range, then introduce a collaborative correction coefficient to jointly correct the basic predicted velocity values ​​of multiple layers to ensure that the long-term trend of vertical velocity conforms to the shear relationship constraint.

[0054] (III) Correction of physical law constraints for the second prediction level (short-term fluctuation prediction): The second prediction level employs a differential / local statistical model to correct the basic prediction values ​​from the first level, generating final prediction values ​​that capture local high-frequency fluctuations. This process involves two layers of physical constraint correction: short-term fluctuation physical constraints for a single physical quantity + cross-quantity constraint correction involving multiple physical quantities. This is the core of the physical constraint correction process, ensuring that the prediction correction for short-term fluctuations still conforms to the coupling laws of ocean dynamics. Physical constraint correction for short-term fluctuations of a single physical quantity: Core logic: For short-term high-frequency fluctuations in waves, pressure, and current velocity (such as pressure pulsations caused by waves and small fluctuations in local flow fields), embed physical reasonable range constraints on the amplitude / frequency of the fluctuations to avoid non-physical fluctuations in the corrected prediction values ​​that exceed the actual sea conditions; Example: The short-term fluctuation amplitude of wave height is limited by the sea state level (e.g., the short-term fluctuation of wave height in light sea conditions does not exceed 0.5m). If the correction range of the second level to the basic prediction value of wave height exceeds this physical range, the correction range will be automatically limited to the threshold to generate a final prediction value that conforms to the actual sea conditions.

[0055] Transverse constraint correction for multi-physical quantity correlation: Core logic: Introducing the ocean dynamic coupling law across physical quantities as a constraint, the prediction corrections between different physical quantities are mutually verified and constrained, avoiding logical contradictions between the correction results of a single physical quantity and other physical quantities. This is the core innovation of the prediction model in this application that distinguishes it from traditional single physical quantity prediction. Specific example of coupling constraint correction: ① Pressure-wave height coupling: Short-term wave pulsation will cause synchronous pulsation of surface pressure. If the short-term correction value of pressure in the second level has no temporal correlation with the correction value of wave height (such as a large increase in wave height but no pressure pulsation), the pressure correction value will be corrected a second time according to the wave height-pressure pulsation correlation formula. ② Wave height-wave speed coupling: According to deep-water wave theory, wave speed is positively correlated with the square root of wave height. If the correction value of wave speed in the second level violates this relationship with the correction value of wave height, the correction value of wave speed is adjusted by the dynamic dispersion relation correction factor. ③ Velocity-Wave Height Coupling: The shearing of velocity will affect the propagation of waves. If there is a logical contradiction between the surface velocity correction value and the wave height correction value of the stratified velocity (such as high surface velocity but no wave height attenuation), the correction values ​​of the two will be adjusted in a coordinated manner according to the wave-current interaction constraint.

[0056] (iv) Overall output logic of multi-level prediction correction The first prediction level generates basic prediction values ​​that conform to physical continuity by constraining the long-term physical laws of a single physical quantity, thus solving the non-physical trend bias that is prone to occur in long-term predictions of traditional models. The second prediction level first corrects the basic prediction value by using the physical constraints of short-term fluctuations of a single physical quantity, and then performs a second correction by using the cross-quantity constraints of multiple physical quantities to finally generate the final prediction value that conforms to both the short-term fluctuation characteristics of a single physical quantity and the coupling law of multiple physical quantities. Throughout the entire multi-level prediction correction process, the laws of ocean dynamic physics are always used as hard constraints, which completely avoids the problem of "non-physical prediction values" that are prone to occur in pure mathematical model predictions, ensuring the high accuracy of the prediction values, and thus minimizing the amplitude and making the distribution of the subsequently constructed joint residual sequence most concentrated.

[0057] Step 3: Combined residual construction and multi-scale reorganization.

[0058] After obtaining the final predicted value, the module compares it point by point with the original collected value, calculates the difference between the two, and obtains the initial residual sequence (i.e., the joint residual sequence). For example, if the original collected value of the pressure data above is 2.35 MPa and the final predicted value is 2.36 MPa, then the residual is -0.01 MPa. Because the multi-level prediction model makes full use of physical constraints and data correlation, the residual values ​​of most sampling points are concentrated in a small range near zero, but there are still a few points with relatively large residuals. To further improve compression efficiency, the module performs multi-scale partitioning and recombination of the initial residual sequence: residuals with smaller absolute values ​​are assigned to the "primary residual set," and residuals with larger absolute values ​​are assigned to the "secondary residual set." The two sets are then processed differently, so that the recombined joint residual sequence exhibits highly concentrated and low-entropy characteristics in statistical distribution, creating favorable conditions for subsequent encoding.

[0059] Step 4: Adaptive lossless coding.

[0060] The module performs rapid statistical analysis on the recombined joint residual sequence, extracting its statistical distribution characteristics, including the maximum amplitude, the proportion of zero values, and dispersion indicators (such as standard deviation or mean square value). Taking a set of actual data as an example, assuming that the statistics show the maximum absolute value of the joint residual sequence is 0.15, the proportion of zero values ​​is as high as 85%, and the standard deviation is 0.03, these indicators show that the residual sequence is highly concentrated and has a narrow distribution range. At the same time, the module monitors the current operating status of the edge nodes in real time, detecting that the processor utilization rate is 30% and the remaining power is 75%, both at a low load level. That is, the selection based on the statistical distribution characteristics of the recombined joint residual sequence is actually a comprehensive decision-making process: using the maximum amplitude, the proportion of zero values, and dispersion as the three core statistical characteristics, by comparing with preset thresholds A, B, and C, combined with the joint judgment of multiple indicators, the resource status of the edge nodes (processor load, remaining power), and the adaptability of the encoding results to the frame length constraints (back-off mechanism), the most suitable lossless encoding method is finally dynamically selected. This mechanism achieves the optimal balance between compression efficiency, computational overhead, and communication adaptability while ensuring that the data is completely lossless.

[0061] Based on the aforementioned statistical characteristics and resource status, the module determines, according to preset decision rules, that the current residual sequence is suitable for a low-complexity fixed-length bit-width encoding method, uniformly encoding each residual value as a 4-bit binary number. In another scenario, if the statistical characteristics show that the residual sequence is dispersed and the maximum amplitude is large, while the current battery power is sufficient, the module may choose entropy encoding or variable-length bit-width encoding to achieve a higher compression ratio; if the battery power is detected to be below 20%, the module will forcibly switch to the encoding method with the lowest computational complexity to prioritize ensuring battery life. Through this adaptive mechanism, the module achieves a dynamic balance between encoding efficiency and resource consumption while ensuring lossless data recovery.

[0062] Step 5: Frame-level organization and encapsulation.

[0063] After encoding, the module obtains a stream of binary encoded data. At this point, the module reads the fixed frame length constraints of the BeiDou satellite short message, for example, the maximum effective payload per frame is 78 bytes. If the current encoded data length is 60 bytes, the module will encapsulate all data within a single frame; if the length is 150 bytes, the module will divide the data into two frames and add encapsulation information such as frame sequence number and total frame count to the header of each frame; if the length is 75 bytes, the module will compress the data to within 78 bytes by fine-tuning the encoding granularity to avoid padding redundancy. After encapsulation, the module generates a standard compressed data packet and stores it in the transmission buffer for the communication module to use.

[0064] Through the coordinated processing of the above five steps, the edge data compression module independently completes the entire conversion process from raw data to compressed data packets locally on the buoy.

[0065] The construction of joint residuals for multi-physical quantity data is not simply a difference calculation between the observed and predicted values ​​of a single physical quantity, but a structured residual construction process that combines the correlations of multiple physical quantities and undergoes multi-scale reorganization. Furthermore, the correction of multi-level prediction models based on physical constraints and the correlations of multiple physical quantities is not merely a simple mathematical model iteration, but a physical correction that embeds ocean dynamic physics laws as hard constraints into the entire prediction process. The core design of the technical solution will be explained in detail below: I. Construction of Joint Residual Sequences: Not simple residuals, but structured residuals involving multi-physical quantity correlation and multi-scale recombination. The construction of the joint residual sequence in this invention is completely different from the traditional simple residual calculation of "observation value - single predicted value". The core is to achieve structured optimization of the residuals through residual fusion under the constraints of multiple physical quantity correlation and multi-scale partitioning and recombination, ultimately achieving the effect of increasing the proportion of zero values ​​and more concentrated statistical distribution. The specific steps are as follows: Basic residual acquisition: First, for three types of physical quantities, namely hydrostatic pressure, wave characteristics, and stratified flow velocity, calculate the basic residual of each physical quantity between the observed value and the final value of the multi-level prediction, and obtain three independent residual sequences: pressure residual, wave height / wave velocity residual, and stratified flow velocity residual.

[0066] Joint residual fusion based on multi-physical quantity correlation: Introducing multi-physical quantity correlation constraints under ocean dynamic laws (such as water body hydrostatic pressure continuity constraints, wave energy stability constraints, adjacent depth layer flow velocity coordination constraints, pressure-wave height vertical correlation constraints, etc.) to fuse three independent basic residuals to construct a joint residual sequence.

[0067] For example, surface pressure pulsations caused by waves can cause a temporal correlation between pressure residuals and wave height residuals. By using correlation constraints, the residuals of the two are coupled and calculated to eliminate residual fluctuations without physical meaning. The residuals of adjacent depth layers of stratified flow velocities need to satisfy cooperative constraints. If the residual of a certain depth flow velocity exceeds the correlation threshold with the adjacent layer, it is coupled and corrected. The final joint residual sequence is no longer a simple splicing of single physical quantity residuals, but a unified residual sequence containing the inherent correlation between physical quantities.

[0068] Multi-scale partitioning and recombination: Based on the magnitude, variation characteristics, and physical meaning of the joint residual sequence, it is divided into at least two subsequences (such as small-amplitude residual subsequence, large-amplitude residual subsequence, or steady-state residual subsequence, fluctuating residual subsequence), and the different subsequences are recombinated and rearranged. Small-amplitude residuals concentrated near zero are aggregated and reorganized to increase the overall proportion of zero values; sparse large-amplitude residuals are separately segmented and reorganized to avoid them interfering with the statistical distribution of small-amplitude residuals. The entropy of the recombined joint residual sequence is significantly reduced, and the statistical distribution is highly concentrated. This solves the problems of discrete residual distribution and low coding efficiency in traditional simple residuals, laying the foundation for subsequent adaptive lossless coding.

[0069] In short, the residual of this invention is not a simple residual of "single physical quantity, no correlation, original difference", but a joint residual that incorporates physical correlation of multiple physical quantities, undergoes multi-scale structured reorganization, and significantly increases the proportion of zero values. Its core design goal is to mine the physical structural redundancy of data, rather than just performing mathematical difference calculations.

[0070] By deploying an edge-end data compression module and executing the above processing flow, efficient and lossless compression and adaptive encapsulation of raw data on multiple physical quantities in the deep sea were achieved under the strong bandwidth constraints of BeiDou satellite communication, significantly improving data transmission efficiency and channel resource utilization. The module achieved differentiated processing of three types of heterogeneous data—static pressure, wave characteristics, and stratified current velocity—through data identification and annotation, fully utilizing the physical characteristics and structural redundancy of each data type. The multi-level prediction model introduced physical constraints and correlations between multiple physical quantities, resulting in highly concentrated prediction residuals and a significant reduction in information entropy, providing a high-quality data foundation for subsequent encoding. Adaptive lossless compression was achieved. The encoding mechanism dynamically selects the encoding method based on the residual statistical characteristics and the real-time resource status of edge nodes, achieving an intelligent balance between computational load and compression performance while ensuring lossless recovery. The frame-level organization and encapsulation design has been specially optimized for the fixed frame length of BeiDou short messages, eliminating padding redundancy and maximizing the effective information carrying capacity of a single communication. The entire compression process is completed independently locally at the buoy end without the need for remote computing resources, which significantly reduces communication frequency and transmission power consumption, effectively extending the autonomous operation cycle of the deep-sea buoy system in energy-constrained environments and providing reliable technical support for long-term unattended monitoring operations.

[0071] In other embodiments, the physical constraints are not presented as independent computational modules, but are embedded into the prediction model in the form of constraint terms, feature maps, or error correction mechanisms. The quantitative embedding of physical constraints can be achieved in the following ways: (1) Transform the physical relationship into a functional relationship or difference constraint between the predicted values, and use it as an additional constraint term in the prediction calculation; (2) Transform the physical conservation relationship into a prediction error correction factor and perform secondary correction on the initial prediction results; (3) Map the coupling relationship of multiple physical quantities into cross-variable feature inputs to enhance the predictive model's ability to express the physical correlation structure. For example, for hydrostatic pressure data, its physical characteristics satisfy the characteristics of continuous and monotonic variation of hydrostatic pressure with depth. In the predictive modeling process, the difference in pressure change between adjacent time steps can be used as a constraint. When the deviation between the predicted value and the historical trend exceeds the preset continuity threshold, a smoothing correction factor is introduced to adjust the prediction result.

[0072] In another implementation, the predicted values ​​can be mapped to an equivalent water depth space for consistency verification based on the functional relationship between static pressure and water depth. If the predicted pressure value does not match the reasonable water depth range, a constraint correction mechanism is triggered. For wave characteristic parameter data, there is a relatively stable functional relationship between wave period, wave height, and energy. During the prediction modeling process, an energy stability interval based on a historical time window can be constructed. When the energy change rate corresponding to the prediction result exceeds this interval, the prediction result is proportionally adjusted or weighted for correction.

[0073] In another implementation, the consistency of the dominant frequency of the wave spectrum distribution can be used as a constraint to project the prediction results into the stable spectral range, ensuring that the prediction parameters meet physical rationality. For stratified velocity data, there are shear relationships and velocity gradient continuity characteristics between different depth layers. During the prediction modeling process, the velocity difference between adjacent depth layers can be calculated. When the prediction results cause an abnormal amplification of the inter-layer velocity gradient, a collaborative correction coefficient is introduced to jointly adjust the multi-layer prediction results.

[0074] Optionally, the deep-sea wireless sensing gateway hardware platform includes: The core control unit uses a low-power ARM Cortex-M4 microcontroller; The dynamic power management and time-sharing power supply unit supports wide voltage input and adopts independent power switch control technology. It only supplies power to the corresponding sensors and communication modules during the preset acquisition and communication periods, and cuts off the power supply to peripherals during non-working periods, achieving milliwatt-level sleep power consumption. The multi-mode heterogeneous communication and intelligent switching unit integrates a 4G / 5G cellular communication module and a BeiDou satellite short message communication module, and embeds link quality detection logic to prioritize the use of the cellular network and automatically and seamlessly switch to the BeiDou short message channel when it is unavailable. The multi-protocol sensor interface unit is equipped with multiple digital bus interfaces and analog input interfaces with opto-isolation and transient suppression protection, for mounting heterogeneous sensors including hydrostatic pressure gauges, wave sensors and acoustic Doppler current profilers. The highly reliable dual storage unit adopts a dual backup architecture with random access memory as a high-speed cache and flash memory as a large-capacity persistent storage, and introduces ferroelectric memory as a system black box to store key configuration parameters and breakpoint resume pointers.

[0075] In some embodiments, the deep-sea wireless sensing gateway hardware platform serves as the physical carrier and edge computing core of the entire system, deployed on deep-sea buoys or underwater moorings. Through a precisely designed hardware architecture and modular integration, this platform achieves long-term stable operation and multi-source data acquisition and processing in extremely harsh marine environments. Its specific structure and operating process are as follows: The core control unit uses a low-power ARM Cortex-M4 microcontroller as the system's main control chip. This chip integrates a floating-point unit and a digital signal processing instruction set, with a maximum clock frequency of 168 MHz, enabling efficient execution of complex mathematical operations in edge data compression algorithms at milliwatt-level power consumption. In actual operation, the MCU is responsible for task scheduling: periodically waking up various functional units, controlling sensor data reading, executing compression algorithms, managing communication processes, and automatically entering deep sleep mode when idle. For example, in a complete acquisition-processing-transmission cycle, the MCU completes all calculation tasks within 50 milliseconds after being woken up by the RTC, and then immediately goes back to sleep, with an average operating current of only microamps.

[0076] The dynamic power management and time-sharing power supply unit is the core design for solving the problem of unstable energy supply in deep-sea areas. This unit supports a wide DC 10V-24V input voltage, adaptable to solar panels and battery packs of different specifications. Internally, the unit employs independent power switch control technology, dividing the system power supply into a "constant power domain" (MCU core, RTC clock, ferroelectric memory) and a "controlled domain" (various sensors, 4G / 5G modules, BeiDou modules). In actual operation, the system operates according to a preset data acquisition strategy: assuming a buoy is set to acquire data once per hour, during non-acquisition periods (i.e., a 55-minute idle period), the power management unit cuts off power to all controlled domains, and the entire system enters a deep sleep mode, with static power consumption as low as below 0.1W. When the RTC wakes the system at a set time, the MCU first supplies power to the sensors that need to work (such as hydrostatic pressure gauges and ADCPs), and immediately cuts off their power after data acquisition. Then, based on data transmission requirements, it determines whether the communication module needs to be activated—if the data volume reaches the transmission threshold or the scheduled reporting time is reached, power is supplied to the corresponding 4G or BeiDou module, and immediately cut off power after data transmission is completed. Taking a typical operation as an example: The system is woken up by the RTC 10 seconds before the hour. The MCU first powers the pressure sensor and wave sensor on the RS485 interface. After 2 seconds of data acquisition and analysis, the sensor power is immediately cut off. Then, based on the amount of data, it is determined that the 4G network needs to be used for transmission. The MCU turns on the 4G module power, completes the link establishment and data upload within 3 seconds, and cuts off the module power immediately after receiving the ACK from the server. The whole process consumes a total of about 15mAh of power. The system is in sleep mode for the remaining 50 minutes. The average daily power consumption is reduced by more than 80% compared with the traditional solution.

[0077] The multi-mode heterogeneous communication and intelligent switching unit integrates a 4G / 5G cellular communication module and a BeiDou-3 short message communication module, constructing a dual-redundant communication architecture. The unit embeds link quality detection logic to monitor the availability and signal quality of the current communication link in real time. Taking a deep-sea buoy as an example: when the buoy is in the nearshore area, the 4G signal strength is -75dBm, and the unit prioritizes using the cellular network for high-speed data transmission, capable of transmitting hundreds of megabytes of detailed data daily. When the buoy drifts to an area with no cellular coverage in the open ocean, after three consecutive failed 4G network handshakes, the unit automatically detects the link interruption and seamlessly switches to the BeiDou short message channel. Simultaneously, it triggers the edge compression module to compress the data to a format compatible with the 78-byte BeiDou single frame for transmission. When the buoy returns to the cellular coverage area due to ocean currents, the unit detects the recovery of the 4G signal, automatically switches back to the cellular network, and initiates the breakpoint resume function to retransmit the data accumulated during BeiDou transmission to the server. The entire switching process is transparent to upper-layer applications, ensuring that critical monitoring data is never lost.

[0078] The multi-protocol sensor interface unit is equipped with multiple digital bus interfaces and analog input interfaces for mounting various heterogeneous sensors required for deep-sea monitoring. Specifically, it includes: 2 RS485 interfaces (supporting Modbus-RTU protocol) for connecting hydrostatic pressure gauges and ADCPs; 1 SDI-12 interface for connecting meteorological or hydrological sensors; and 4 analog input interfaces (supporting 4-20mA current signals and 0-5V voltage signals, equipped with a 16-bit high-precision ADC) for connecting traditional analog output sensors. All interfaces employ a three-stage protection circuit design: the first stage uses a gas discharge tube to discharge large currents generated by lightning strikes; the second stage uses a thermistor to limit surge currents; and the third stage uses a TVS transient suppression diode to clamp peak voltages. Simultaneously, the digital interfaces use high-speed optocouplers for electrical isolation to prevent surge currents caused by seawater corrosion and leakage from external sensors from burning out the main control circuitry. Taking actual deployment as an example: A certain buoy is equipped with three RS485 interface ADCPs, one SDI-12 interface wave sensor and two 4-20mA output pressure sensors. The interface unit achieves the coordinated operation of all sensors through time-sharing power supply and multi-protocol compatibility design. Even during the summer thunderstorms, the three-level protection circuit successfully discharged lightning surges multiple times, ensuring the safety of the platform's internal circuitry.

[0079] The highly reliable dual storage unit employs a hierarchical storage architecture to ensure data security. An onboard 512MB NAND Flash serves as a large-capacity persistent memory, capable of cyclically storing nearly a year's worth of raw observation data and supporting breakpoint resume functionality—data continues to be written to the Flash when communication is interrupted; when communication resumes, the system automatically retransmits missing data based on the storage pointer. Simultaneously, the unit incorporates a 256KB ferroelectric memory as the system's "black box," utilizing its unlimited read / write cycles, nanosecond-level write speed, and non-volatility under power loss characteristics to record critical system operating states in real time, including the current acquisition pointer position, compressed dictionary version, breakpoint resume progress, and recent fault error codes. For example, in an unexpected power outage scenario: a buoy shuts down due to battery depletion caused by continuous rainy weather. The breakpoint resume pointer stored in the ferroelectric memory recorded that the 1024th data packet had been sent. A week later, when solar power was restored, the MCU first read the recovery pointer from the ferroelectric memory and resumed transmission from the 1025th data packet, ensuring zero data loss for a week.

[0080] Through the collaborative work of the above six units, the Deep-Sea Wireless Sensing Gateway hardware platform constitutes a highly reliable embedded system integrating data acquisition, edge processing, multi-mode communication, and intelligent power management, providing a stable hardware foundation for upper-layer software modules.

[0081] By deploying a deep-sea wireless sensing gateway hardware platform, highly reliable, low-power, all-weather acquisition and edge processing of raw data of multiple physical quantities were achieved in extremely harsh marine environments, laying a solid hardware foundation for the entire system. The low-power ARM Cortex-M4 core control unit adopted by the platform provides ample computing power at milliwatt-level power consumption, supporting the efficient operation of edge compression algorithms. The dynamic power management and time-sharing power supply unit reduces the static sleep power consumption of the entire unit through independent power switch control technology and refined time-sharing strategy, resulting in lower daily energy consumption than traditional solutions and effectively solving the problem of insufficient solar power supply during long-term cloudy and rainy days. The multi-mode heterogeneous communication and intelligent switching unit ensures that the buoy can reliably transmit data in deep-sea areas without cellular coverage through 4G / 5G and Beidou dual-mode redundancy design and automatic seamless switching logic, achieving "never losing contact" of key monitoring information. The multi-protocol sensor interface unit integrates opto-isolation and three-level transient suppression protection. It can directly mount various heterogeneous sensors such as hydrostatic pressure gauges, wave sensors, and ADCPs, and effectively resist damage from lightning strikes and surges, significantly improving the equipment's survivability in storm-prone sea areas. The highly reliable dual storage unit, through the dual backup design of large-capacity Flash storage and a ferroelectric memory "black box," enables rapid self-healing and breakpoint resumption after system abnormal reset, ensuring the integrity and continuity of data during long-term unattended operations. The platform adopts a corrosion-resistant aluminum alloy shell and IP67 protection design, which can adapt to the deep-sea environment of high salt spray, high humidity and heat, and severe vibration for a long time, providing a highly reliable, easy-to-maintain, and long-life hardware solution for the large-scale deployment of deep-sea monitoring networks.

[0082] like Figure 3 As shown, optionally, the server-side data noise cleaning module is specifically applied to: The compressed data packet is decompressed to obtain the original observation data, and the original observation data is preprocessed to construct a standardized observation vector. A state-space model is constructed based on the physical characteristics of the marine dynamic environment. The observation vector and the optimal state estimate of the previous time step are used as inputs. Kalman filtering is performed to update the time and generate the state prediction value and prediction error covariance of the current sampling point. Based on the deviation between the current state prediction value and the corresponding observation value in the observation vector, the information of the current sampling point is obtained; and based on the corresponding prediction error covariance and the preset observation noise covariance, the theoretical standard deviation of the information is obtained. A strong radial constraint threshold is constructed based on the current sampling point's information and the theoretical standard deviation. A preset multiple of the theoretical standard deviation is used as a threshold. The current sampling point's information amplitude is compared with the threshold to obtain a comparison result. Kalman filtering measurement updates are performed based on the comparison result to obtain the optimal state estimate at the current time, which is used as the cleaned observation data.

[0083] Optionally, the step of updating the measurement using Kalman filtering based on the comparison result to obtain the optimal state estimate at the current time, as the cleaned observation data, includes: If the innovation amplitude of the current sampling point does not exceed the threshold, it is determined that the observation value corresponding to the current sampling point belongs to the background noise. Based on the Sage-Husa adaptive filtering algorithm, the system noise covariance and the observation noise covariance are iteratively corrected in real time with the innovation of the current sampling point as input, so as to obtain the updated current system noise covariance and the updated current observation noise covariance. A forgetting factor is introduced to weight the historical data. If the innovation amplitude of the current sampling point exceeds the threshold, it is determined that the observation value corresponding to the current sampling point belongs to a non-physical impulse field value. The update of the system noise covariance and the observation noise covariance is frozen, the system noise covariance and the observation noise covariance at the previous moment are kept unchanged, and the Kalman gain is suppressed. Based on the updated or frozen current system noise covariance, current observation noise covariance, information of the current sampling point, and the state prediction value, Kalman filtering measurement updates are performed to obtain the optimal state estimate at the current moment, which is used as the cleaned observation data.

[0084] Specifically, the server-side data noise cleaning module is deployed in a shore-based data center or cloud server, serving as the core processing link connecting data transmission and data fusion. This module receives compressed data packets from the edge and performs the following processing steps to intelligently clean the raw observation data: Step 1: Data decompression and preprocessing.

[0085] The module first decompresses the received compressed data packet to fully reconstruct the original observation data collected by the buoy, including hydrostatic pressure time-series data, wave characteristic parameter data, and stratified velocity profile data. Taking a set of actually received ADCP velocity data as an example, the original data stream contains surface velocity values ​​at 100 consecutive sampling times, each value stored in ASCII encoding. The module first parses the format of the original data, removing invalid identifiers (such as the "999" missing value code filled in when a sensor malfunctions), and aligns and calibrates the timestamps of each sensor data to ensure that the pressure, wave, and velocity data at the same sampling time are strictly corresponding in time. After the above preprocessing, the module organizes the standardized data into observation vectors, for example, the observation vector at time k is Zk = [velocity value 0.85 m / s, pressure value 2.35 MPa, wave height value 1.28 m], which serves as the input for subsequent filtering processing.

[0086] Step 2: State-Space Modeling and Kalman Filtering Time Update. After data preprocessing, a state-space model conforming to the dynamic characteristics of physical quantities such as ocean current velocity is constructed. This model is based on the assumption that ocean environmental parameters have short-term physical continuity. Its structure typically consists of state transition equations, used to characterize the dynamic evolution of the ocean current dynamic environment. For example, for ocean current velocity data, the velocity can be set as a state variable, and the theoretical expectation of the current moment can be derived from the dynamic state of the previous moment using a linear model. The input data are the parameters of the previous moment, including the "optimal state estimate" and "error covariance matrix" generated at the previous moment, and the original observation values, i.e., the time series of physical quantities such as ocean current velocity collected by the sensor. The output data is the prior state estimate, i.e., the state prediction value of the current sampling point, reflecting the system's theoretical expectation of the physical quantity and the prior error covariance before the introduction of the current measurement information. It is used to quantify the prediction accuracy and is the core basis for subsequent calculation of Kalman gain and innovation statistical characteristics. In this embodiment, the Kalman filtering time update step (prediction step) is performed using the optimal state estimate and error covariance matrix of the previous moment. The system extrapolates the current state estimate and the priori error covariance based on a physical model. This process is based on the assumption that ocean environmental parameters have physical continuity over a short period of time, and the generated predictions reflect the system's theoretical expectations of the ocean current state before the current observations are incorporated.

[0087] For example, flow velocity is set as a state variable, and its evolution is described by a state transition equation. Taking time k-1 as an example, assume the optimal state estimate of the previous time step is 0.83 m / s, and the corresponding error covariance matrix is ​​P(k-1). The module takes these two values ​​as input and performs a Kalman filter time update step (prediction step), deriving the current state prediction value X(k|k-1) and prediction error covariance P(k|k-1) at time k based on the physical model. For example, based on the inertial characteristics of flow velocity, the model predicts that the current flow velocity should be 0.84 m / s, and gives the prediction uncertainty range of ±0.05 m / s (quantized by P(k|k-1)). This prediction reflects the system's theoretical expectation of the ocean current state based on physical laws before the current observation value is introduced.

[0088] For the physical quantity of ocean current velocity, the module constructs a state-space model that conforms to its dynamic characteristics. This model is based on the assumption that ocean environmental parameters have physical continuity over a short period, setting the velocity as a state variable and describing its evolution through state transition equations. Taking time k-1 as an example, the optimal state estimate of the previous time step is assumed to be 0.83 m / s, with a corresponding error covariance matrix P(k-1). The module takes these two values ​​as input and performs a Kalman filter time update step (prediction step), deriving the predicted state value X(k|k-1) and prediction error covariance P(k|k-1) for the current time k based on the physical model. For example, based on the inertial characteristics of the velocity, the model predicts the current velocity to be 0.84 m / s and provides a prediction uncertainty range of ±0.05 m / s (quantified by P(k|k-1)). This predicted value reflects the system's theoretical expectation of the ocean current state based on physical laws before incorporating the current observation value. Step 3: Innovation Calculation and Statistical Feature Extraction. After obtaining the state prediction value, the module compares it with the corresponding observation value in the observation vector obtained in Step 1. Taking flow velocity data as an example, the actual observed value at the current moment is 0.85 m / s, while the state prediction value is 0.84 m / s. The difference between the two is the innovation (residual) νk = 0.01 m / s. This innovation reflects the part of the observed data that was not predicted by the physical model, including measurement noise, system process noise, and potential abnormal interference. At the same time, the module uses the prediction error covariance generated in Step 2 and the preset initial observation noise covariance R (e.g., preset to 0.0025 m² / s² based on sensor accuracy) to calculate the theoretical standard deviation σk of the innovation sequence in real time, as a dynamic benchmark for measuring the current noise level. Through this step, the assessment of data quality is transformed into the analysis of the statistical characteristics of the innovation sequence, providing a quantitative basis for subsequently distinguishing between "normal noise" and "outliers". Step 4: Strong Radial Constraint Threshold Discrimination. To address the limitation of traditional algorithms in distinguishing between high-frequency background noise and sudden outliers, the module constructs strong radial constraint logic. A threshold is set as a multiple of the theoretical standard deviation σk obtained in Step 3, for example, 3 times the standard deviation (3σ). The module continuously monitors whether the amplitude of the current innovation νk is within the "safe threshold" defined by this threshold: if |νk| ≤ 3σk, the current deviation is determined to be normal environmental background noise (such as small velocity fluctuations caused by wind and waves); if |νk| > 3σk, the current deviation is determined to be a non-physical impulse outlier (such as sudden changes caused by biological impacts or sensor malfunctions). Taking the current data as an example, |0.01| ≤ 0.15, therefore it is determined to be normal background noise. This discrimination mechanism gives the algorithm "immunity" to extreme outliers.

[0089] Step 5: Differentiated Parameter Adjustment and Outlier Suppression. Based on the discrimination results of Step 4, the module implements a differentiated parameter adjustment strategy to address the problem of the traditional Sage-Husa algorithm being prone to divergence under outlier interference.

[0090] Scenario 1: The observation value corresponding to the current sampling point is determined to be background noise (the innovation does not exceed the threshold). Taking the current current velocity data as an example, since the innovation of 0.01 m / s does not exceed the 3σ threshold (0.15 m / s), the module determines that the observation value corresponding to the current sampling point belongs to background noise and starts the Sage-Husa adaptive filtering algorithm. This algorithm is based on the maximum a posteriori estimation criterion and uses the current innovation νk to iteratively correct the system noise covariance Q and the observation noise covariance R in real time. Among them, a forgetting factor is introduced to weight historical data: the closer its value is to 1, the higher the degree of retention of historical data and the smoother the filtering; the smaller the value, the greater the weight of the algorithm on the current observation value and the more sensitive the response to sea state changes. Through this adaptive mechanism, the filter parameters can closely follow the changes in sea state. For example, when the wind and wave intensity increases and the noise level rises, the R value automatically increases, reducing the trust weight of the observation value; when the sea state returns to calm, the R value automatically decreases, and the observation value is trusted again. After correction, the module obtains the updated current system noise covariance Qk and the current observation noise covariance Rk.

[0091] Scenario 2: The observed value corresponding to the current sampling point is determined to be an outlier (innovation exceeds the threshold). Suppose that at a certain moment, the observed velocity suddenly changes to 1.50 m / s due to a biological impact, while the predicted state value remains 0.84 m / s. The calculated innovation νk = 0.66 m / s, which is much larger than the current 3σ threshold of 0.15 m / s. At this time, the module immediately triggers the outlier suppression mechanism: it forcibly freezes the updates of the system noise covariance and the observation noise covariance, keeping Qk-1 and Rk-1 unchanged from the previous moment; simultaneously, it suppresses the Kalman gain by forcibly setting the gain matrix to zero or significantly reducing it, so that the current observed value hardly participates in the state update, and the predicted value is directly used as the output. This strategy effectively blocks dirty data from contaminating the filter's internal parameters and prevents algorithm oscillations and divergence. For example, even if outliers persist for multiple sampling points, because Q and R are frozen and the gain is suppressed, the filter can still maintain a stable output, automatically resuming normal tracking after the outliers disappear.

[0092] Step 6: Measurement Update and Optimal State Estimation. Using the corrected or frozen current system noise covariance and current observation noise covariance obtained in Step 5, along with the current innovation obtained in Step 3 and the state prediction value generated in Step 2, the module executes the measurement update step of the Kalman filter. First, the Kalman gain Kk = P(k|k-1)·H is calculated. ·(H·P(k|k-1)·H + Rk) - ¹, and then calculate the optimal state estimate for the current moment: Xk = X(k|k-1) + Kk·νk. Taking a normal background noise scene as an example, assuming Kk = 0.8, then Xk = 0.84 + 0.8×0.01 = 0.848m / s. This optimal estimate utilizes the prediction information of the physical model and incorporates the effective components of the observed data, and the noise level has been adaptively adjusted to the best matching state. The module outputs this value as the cleaned high-fidelity physical data, and simultaneously uses Xk and Pk = (I-Kk·H)·P(k|k-1) as the input for the next moment, entering the next loop.

[0093] Through the coordinated processing of the above six steps, the server-side data noise cleaning module completes a full intelligent cleaning cycle for the raw observation data at each moment.

[0094] It should be noted that when performing measurement updates using Kalman filtering to obtain cleaned observation data, in addition to following the innovation threshold judgment logic described above, updating the noise covariance or freezing the noise covariance and suppressing the Kalman gain respectively before completing the complete measurement update process, a simplified implementation method can also be adopted: when the current sampling point is determined to be a non-physical impulse outlier, the state prediction value can be directly used as the optimal state estimate value at the current moment, and used as the cleaned observation data.

[0095] This simplification method is consistent with the core logic of noise covariance freezing and Kalman gain suppression mentioned above. When the observed data is determined to be an outlier, the measurement update process for the outlier is weakened or skipped to avoid interference from the outlier on the filtering results. This ensures the physical rationality and reliability of the cleaned observed data, while also simplifying the calculation process of edge nodes and improving data processing efficiency.

[0096] By deploying a server-side data noise cleaning module and executing the above processing flow, high-precision and robust real-time cleaning of raw observation data transmitted from deep sea was achieved, effectively solving the data pollution problem caused by the coexistence of background noise and sudden outliers under complex sea conditions. The module introduces a strong radial constraint threshold discrimination mechanism, using a multiple of the real-time standard deviation of the innovation as a dynamic threshold to intelligently distinguish normal background fluctuations from non-physical pulse outliers, avoiding filter divergence caused by outlier impacts in traditional algorithms. For background noise, the Sage-Husa adaptive filtering algorithm is used, which uses real-time iterative correction of the system noise covariance and observation noise covariance using the innovation, and introduces a forgetting factor to process historical data. Weighted processing is applied to ensure that the filter parameters keep pace with the dynamic changes in sea state, maintaining optimal filtering performance from calm seas to typhoon passage. Parameter freezing and gain suppression strategies are employed to address non-physical outliers, forcibly preventing contamination of the filter's internal parameters by dirty data and ensuring the algorithm's stability under extreme outlier interference. The final cleaned data output effectively filters out high-frequency clutter while fully preserving the true physical trends such as gradual changes in ocean current velocity and tidal cycles. Verification by measured data shows that the root mean square error is reduced compared to the original noisy sequence, and there is no phase lag problem as in traditional filtering methods, providing a high-fidelity, high-time-accuracy data foundation for subsequent multiphysics collaborative fusion.

[0097] like Figure 4 As shown, optionally, the multiphysics collaborative data fusion module is specifically applied to: Different types of data in the cleaned observation data are preprocessed separately to obtain corresponding standardized ocean data; The standardized marine data is unified in terms of spatiotemporal reference, and a grid background field is generated through a unified spatiotemporal registration and grid construction process. A unified state vector is constructed based on the grid background field. Under the guidance of the preset physical constraint system, a unified dynamic data assimilation loop is executed to obtain the optimized unified state vector. The dynamic data assimilation loop is repeated until the current unified state vector satisfies the convergence condition, thereby obtaining the optimal ocean data field.

[0098] Optionally, the dynamic data assimilation loop includes: Using the current state value in the unified state vector as input, the state evolution is predicted based on the physical model through ensemble Kalman filtering to obtain the predicted state vector; then, the predicted state vector is updated according to the real-time observation data of the corresponding spatiotemporal nodes in the grid background field to obtain the updated state vector; the core parameters in the updated state vector are corrected for prediction errors through a long short-term memory network model to obtain the optimized unified state vector.

[0099] Optionally, the preset physical constraint system includes at least one of the following: hydrostatic equilibrium equation, simplified shallow water wave equation, geostrophic equilibrium equation, or mass conservation continuity equation.

[0100] Optionally, the core parameters include wave height, wave velocity, and pressure.

[0101] Specifically, the multi-physics collaborative data fusion module is deployed in a shore-based data center or cloud server, serving as the core backend processing component of the entire system. This module receives cleaned observation data from the server-side data noise removal module, including noise-suppressed hydrostatic pressure time-series data, wave characteristic parameter data, and stratified current velocity profile data. Because these three types of data originate from different sensors and observation platforms, they exhibit device clock skew in terms of time reference, discrete point-based spatial distribution, and different sampling frequencies, data formats, and physical units. After the fusion module is activated, it executes the following processing flow to integrate the multi-source heterogeneous data into a spatiotemporally continuous and physically consistent optimal ocean data field: Phase 1: Targeted Preprocessing - Achieving Data Standardization. Based on the physical characteristics and interference factors of the three types of data, the module employs differentiated preprocessing methods: Hydrostatic pressure data preprocessing: Raw pressure observations are affected by changes in seawater temperature and salinity and cannot be directly used for fusion. The module uses a temperature and salinity compensation correction formula to correct the pressure data. ;in, : Corrected pressure value; Original pressure observations; Seawater density (calculated using the TEOS-10 equation of state); : Gravitational acceleration (usually taken as ); Sensor deployment depth deviation; Temperature and salinity compensation coefficients; Standard temperature (e.g., 20°C), standard salinity (e.g., 35‰). Salinity For temperature.

[0102] Taking a deep-sea observation point as an example, the original pressure value is 15.32 MPa, the water temperature is 25℃, and the salinity is 35‰. After compensation and correction, the actual pressure value is 15.28 MPa. Subsequently, the module converts the corrected pressure into a dynamic height field and density profile parameters based on the TEOS-10 equation: Density: ;in, The density of seawater, The density of standard seawater; These are the fitting coefficients.

[0103] Power height: ;in, For power height; For reference depth; The depth of observation.

[0104] Through this transformation, a single pressure observation is sublimated into a standardized dynamic parameter that can be directly coupled with flow velocity and wave data.

[0105] Wave feature data preprocessing: For the two core parameters of significant wave height and wave velocity, the module uses an improved Kalman filter for smoothing and noise reduction.

[0106] Taking wave height data as an example, an initial state vector X and an initial covariance are set. Sensor noise and environmental interference are filtered out through iterative calculations based on state prediction and observation updates. For example, if the original observed wave height is 1.32 m at a certain moment, the filtered output is 1.28 m. Simultaneously, a Kalman gain threshold (K≤0.8) is constrained to ensure that the true evolution trend of the wave height is fully preserved. Wave velocity data is processed in the same way to obtain a smoothed wave velocity sequence that is consistent with the physical relationship between wave height and wave height.

[0107] Layered velocity data preprocessing: For the raw velocity data acquired by ADCP, the module first uses a linear coordinate transformation method to unify the depth layers of different observation points to a standard depth layer: ;in, Standard depth layer; Original observation depth; These are coordinate transformation coefficients (derived from regional terrain fitting).

[0108] Then, the 3σ criterion was used to remove outliers, and the mean flow velocity at each depth layer was calculated. and standard deviation , ; ; Outlier detection: If If so, then remove it.

[0109] Finally, a composite index model is adopted: Vertical interpolation smoothing is performed to fill in missing data and ensure a continuous and complete velocity profile. Among these steps, Let z be the flow velocity at depth z. : Initial value of surface flow velocity; k is the attenuation coefficient (an empirical value can be taken as : ).

[0110] After the first stage of processing, all three types of data were transformed into standardized ocean data with unified units (pressure MPa, wave height m, wave velocity m / s, current velocity m / s), unified format (NetCDF format), and preliminary alignment with spatiotemporal references.

[0111] Phase Two: Spatiotemporal Registration and Gridding – Constructing a Continuous Background Field. The module first maps the timestamps of standardized ocean data to UTC standard time, eliminating clock deviations between different observation devices. Taking a specific data processing example, the original data showed the pressure sensor timestamp as 2024-03-15 08:23:15 (local device time) and the wave sensor timestamp as 2024-03-15 08:23:22. After calibration, these timestamps were unified to UTC time 2024-03-15 00:23:15 ± 1 second. For sporadic missing values ​​in the time series, linear interpolation was used to fill in the gaps (missing ≤ 30 minutes) or historical data fitting was used (missing > 30 minutes).

[0112] Spatially, all location information was unified to the WGS-84 coordinate system, and a three-dimensional grid was set according to the observation area: horizontal resolution of 0.01°×0.01°, and vertically layered at 10m / layer (from 0m at the surface to 5000m at the deep sea). Differential interpolation methods were used to fill the discrete observation points based on the spatial distribution characteristics of different physical quantities. The pressure data exhibits a significant vertical distribution characteristic. Using the Kriging interpolation method, interpolation weights are assigned based on the pressure values ​​of multiple discrete observation points and spatial correlation to calculate the pressure value of each grid node.

[0113] The wave data exhibits a significant horizontal distribution characteristic. The inverse distance weighted interpolation method is adopted, and the weights are assigned according to the horizontal distance between the grid node and each observation point (the closer the distance, the greater the weight, and the weight coefficient is 2) to calculate the wave height and wave velocity values ​​of each grid node.

[0114] Velocity data were interpolated using a combination of a composite exponential model and kriging: first, kriging was used to fill in the missing horizontal values, and then the composite exponential model was used to optimize vertical continuity. Kriging interpolation was used to fill in spatially missing values ​​between different observation points; the interpolation formula is as follows: ,in The standard depth flow velocity value, interpolation weights (satisfying) i=1), is the original depth layer velocity observation value, and m is the original number of samples used for interpolation.

[0115] Taking a 1°×1° sea area as an example, the original five discrete observation points were interpolated to generate a complete three-dimensional grid background field with 100×100 horizontal grid nodes and 50 vertical layers. Each grid node contains parameters such as pressure, wave height, wave velocity, and current velocity u / v / w components. Finally, through physical rationality verification, such as the law of pressure increasing with depth and the dynamic correlation between wave height and wave velocity, local anomalous data were corrected to ensure the physical logic of the background field is consistent.

[0116] Phase Three: Dynamic Assimilation and Fusion under Physical Constraints to Generate the Optimal Data Field. Based on the grid background field generated in Phase Two, the module constructs a unified state vector, with wave height, wave velocity, and pressure as core parameters. Guided by a pre-defined physical constraint system, a dynamic data assimilation loop is executed: The pre-defined physical constraint system includes physical equations describing the intrinsic laws of ocean dynamics, mainly including: the hydrostatic equilibrium equation, which constrains the vertical distribution of pressure with depth; the simplified shallow water wave equation, which constrains the horizontal coupling relationship between wave height and current velocity; the geostrophic equilibrium equation, which constrains the balance relationship between horizontal current velocity and pressure field; and the mass conservation continuity equation, which constrains the overall coordination of the three-dimensional velocity field.

[0117] The dynamic data assimilation loop executes the following steps: Step 1: Ensemble Kalman Filter Prediction. Using the current unified state vector X(k-1) as input, the state evolution is predicted based on a marine dynamic physics model (such as a simplified form of the shallow water equation), resulting in the predicted state vector X(k|k-1). For example, based on the wave height and current velocity fields from the previous moment, the predicted wave height at the current moment should evolve to 1.30 m, and the wave velocity to 3.45 m / s. During this process, a physical constraint system is used to regulate the prediction direction, ensuring that the prediction results do not violate fundamental laws such as hydrostatic equilibrium and energy conservation; for example, the predicted pressure value must monotonically increase with depth.

[0118] Step 2: Observation Update. The real-time observation data of the corresponding spatiotemporal nodes in the grid background field (e.g., a measured wave height of 1.28 m at a certain grid point) is compared with the predicted state vector. The predicted state is corrected using a gain matrix to obtain the updated state vector. During this process, a physical constraint system is used to verify the rationality of the observation data—if an observation value deviates significantly from physical laws (e.g., a wave height of 5 m but a wave speed of only 1 m / s, violating the deep-water wave dispersion relation), the confidence weight of that observation value is reduced to avoid unreasonable data contaminating the fusion results.

[0119] Step 3: LightGBM Error Correction. The updated state vector is input into a Long Short-Term Memory (LSTM) network model (such as the LightGBM machine learning model) to correct prediction errors for the core parameters (wave height, wave velocity, and pressure). This model is pre-trained using historical data to learn the nonlinear pattern of the residuals after EnKF (Kalman Filter) assimilation. Taking wave height as an example, assuming the wave height after EnKF update is 1.29 m, the LightGBM model, based on the current flow and pressure field characteristics, determines that there is a systematic underestimation of 0.02 m, outputting a correction of +0.02 m, resulting in a final optimized value of 1.31 m. The corrected state vector is used as the input for the next loop.

[0120] Iterative convergence. Repeat the above loop, performing one round of "prediction-update-correction" process for each new batch of grid background field data. Continue iterating until the change in the core parameters in the state vector is less than the preset threshold (e.g., wave height change < 0.01 m, pressure change < 0.01 MPa), and the overall vector tends to stabilize. At this point, the optimal ocean data field is output.

[0121] Taking a specific fusion as an example, the initial background field at a certain grid point had a pressure of 15.28 MPa, a wave height of 1.28 m, a wave velocity of 3.42 m / s, and a current velocity of 0.85 m / s. After five rounds of dynamic assimilation, the pressure was optimized to 15.31 MPa, a wave height of 1.31 m, a wave velocity of 3.45 m / s, and a current velocity of 0.87 m / s. All parameters satisfied the hydrostatic equilibrium equation (pressure and depth are strictly correlated), the shallow water wave equation (wave height and wave velocity satisfy energy conservation), and the continuity equation (current velocity divergence is zero), forming a physically self-consistent and spatiotemporally continuous optimal ocean data field. This data field is stored in a four-dimensional grid format (time + longitude + latitude + depth), and each grid node contains complete parameters such as wave height, wave velocity, pressure, three-dimensional current velocity, density, and sea surface height.

[0122] In some embodiments, the specific steps of the improved Kalman filter algorithm are as follows: 1. State initialization: Set wave height ( The initial state vector of wave speed (c) Initial covariance matrix ,in, and The initial noise variances for wave height and wave velocity are respectively taken as 0.1m and 0.2m / s, respectively. 2. State prediction: X(k|k-1)=A X(k-1)+B wk, P(k|k-1)=A P(k-1) A T +Q, Where A is the state transition matrix (taken as identity matrix I), B is the input matrix, and wk is the process noise (following a normal distribution N(0,Q), where Q is the process noise covariance matrix). 3. Observation Update: Kk = P(k|k-1)·H ·(H·P(k|k-1)·H + Rk) - ¹; Xk=X(k|k-1)+Kk (Zk-H X(k|k-1)), Pk=(I-Kk) H) P(k|k-1); Where H is the observation matrix (taken as the identity matrix I), Zk=[ ,craw] T R is the original observation vector, R is the observation noise covariance matrix (empirical value diag(0.01,0.04)), and Kk is the Kalman gain. Through the above algorithm, on the one hand, it effectively filters out sensor noise and interference from extreme deep-sea environments (such as strong currents and undercurrents) causing fluctuations in wave height and wave velocity data, and eliminates abnormal fluctuation values; To further enhance the smoothing effect, the algorithm employs a sliding window averaging method for auxiliary smoothing after the Kalman filter output. The window size is set to 5 observation points, with a step size of 1, meaning the arithmetic mean is taken for the current point and the two points before and after it (a total of 5 points). The formula is: ; (yi represents the smoothed data, (This refers to the original observation data).

[0123] By deploying a multi-physics collaborative data fusion module and executing the above processing flow, this application achieves the integration of discrete, heterogeneous, cleaned observation data into a spatiotemporally continuous and physically consistent optimal ocean data field, completely solving the "island effect" problem caused by inconsistent spatiotemporal references in traditional methods. The module, through targeted preprocessing, eliminates temperature and salinity interference in pressure data, high-frequency noise in wave data, and outliers and vertical discontinuities in current velocity data, providing a high-quality, standardized data foundation for subsequent fusion. The spatiotemporal registration and gridding construction process employs a differentiated interpolation strategy, enabling pressure, wave, and current velocity data to generate a continuous and complete three-dimensional / four-dimensional grid background field under a unified UTC time and WGS-84 coordinate system, filling numerical gaps in observational areas. The pre-set physical constraint system explicitly embeds fundamental ocean dynamic laws such as hydrostatic equilibrium equations, simplified shallow water wave equations, geostrophic equilibrium equations, and mass conservation continuity equations into the fusion process, ensuring that the generated data field strictly conforms to physical constraints while satisfying the characteristics of the observation data. The logical reasoning—for example, pressure strictly increases with depth, wave height and wave speed satisfy energy conservation, and the velocity field satisfies continuity—effectively eliminates the physical contradictions often caused by pure mathematical interpolation methods. The dynamic data assimilation loop innovatively combines the physical modeling capabilities of ensemble Kalman filtering with the data-driven advantages of LightGBM machine learning—EnKF is responsible for evolution prediction and observation updates based on physical laws, while LightGBM specifically corrects the nonlinear prediction errors of core parameters. The two complement each other, significantly improving the fusion accuracy. The final generated optimal ocean data field uses wave height, wave speed, and pressure as core parameters, fully covering the three major physical fields of pressure, wave, and velocity. It has a complete four-dimensional structure of "time + horizontal + vertical" and can further derive high-value parameters such as dynamic height, density profile, and sea surface roughness. It provides the first high-precision data product that fully characterizes the multi-physics coupled evolution characteristics for deep-sea scientific research, engineering applications, and disaster early warning, realizing a qualitative leap from "data acquisition" to "information generation."

[0124] Optionally, the adaptive selection of lossless coding method includes selecting an appropriate coding method from fixed-length bit-width coding, variable-length bit-width coding, or entropy coding based on at least one of the maximum amplitude, zero-value ratio, and dispersion index of the joint residual sequence, as well as the current computing resource or power consumption status of the edge node.

[0125] In some embodiments, the adaptive lossless coding strategy of the edge data compression module is based on statistical analysis of the joint residual sequence and dynamic decision-making on the real-time resource status of the edge nodes, so as to balance compression efficiency and computational overhead while ensuring lossless data recovery.

[0126] Specifically, the module first performs rapid statistical analysis on the recombined joint residual sequence, extracting its maximum absolute value, the proportion of zero values, and dispersion indicators (such as standard deviation or mean square value). Based on a preset threshold system, the module initially determines the suitable encoding method according to the following rules: If the maximum absolute value of the residual sequence is less than the preset amplitude threshold A, it indicates that the residual distribution is concentrated within a finite range, and fixed-length bit-width encoding or small-bit-width encoding can be used efficiently; if it exceeds threshold A, variable-length bit-width encoding or entropy encoding is triggered to avoid bit waste caused by fixed bit width. The amplitude threshold A can be preset based on the statistical interval of historical data or the calibration results at the initial stage of buoy operation. If the proportion of zero values ​​in the residual sequence is higher than the preset proportion threshold B, it indicates that the sequence has sparse distribution characteristics, and encoding methods supporting run-length encoding or zero-value optimization are preferentially selected to further improve the compression ratio; conversely, if the proportion of zero values ​​is low, zero-value enhancement encoding is not enabled to avoid increasing unnecessary control overhead. The proportion threshold B can be set based on the mean fluctuation range of the historical residual distribution. If the dispersion index (such as standard deviation) of the residual sequence is less than the preset dispersion threshold C, it indicates that the residuals are highly concentrated and a low-complexity encoding method can be used; if it exceeds the threshold C, then switch to a method with a higher compression ratio, such as entropy encoding or segmented bit-width encoding.

[0127] To further enhance the robustness of decision-making, one implementation method employs a multi-indicator joint judgment mechanism: when at least two of the maximum amplitude, the proportion of zero values, and the dispersion index simultaneously meet the concentrated distribution condition (i.e., below threshold A, above threshold B, and below threshold C, respectively), the residual sequence is determined to be suitable for low-complexity coding as a whole; otherwise, it switches to a conservative lossless coding method to ensure reliability.

[0128] In addition, adaptive decision-making also considers the real-time resource status of edge nodes. When the processor utilization rate is detected to exceed the preset load threshold D, or the power management module reports that the remaining power is lower than the preset power threshold E, the system forcibly triggers a low-complexity coding strategy to prioritize reducing computational overhead and power consumption, ensuring the continuous operation of the system under resource-constrained conditions.

[0129] If, during the encoding process, the length of the encoded data exceeds the maximum effective payload limit of a single frame of BeiDou short message, an automatic rollback mechanism is triggered: the system immediately terminates the current encoding and switches to a preset high compression ratio lossless encoding method to re-encode the original residual sequence, ensuring that the generated compressed data packet can adapt to the fixed frame length requirement of BeiDou short message and avoid transmission failure due to excessive length.

[0130] Through the aforementioned multi-dimensional and multi-level adaptive decision-making mechanism, the edge device can flexibly select the optimal encoding strategy under different data characteristics and resource conditions, achieving a dynamic balance between compression efficiency, computational overhead, and communication adaptability.

[0131] In some specific embodiments, to verify the effectiveness of the ocean buoy data noise removal process based on improved Kalman filtering in the server-side data noise removal module, the following comparative experiment was designed: Experimental Data and Noise Simulation: The experiment used ocean current velocity data measured at the XX Marine Data Center buoy station as the baseline. This dataset represents the continuous variation of ocean current velocity in a real ocean environment, including real physical characteristics such as tidal cycles and gradual changes in current eddies.

[0132] To simulate data contamination scenarios under extreme sea conditions in the deep ocean, artificial noise was superimposed on the original clean data: Background Gaussian noise: Superimposed Gaussian white noise with a mean of 0 and a standard deviation of 2.5 cm / s to simulate continuous random fluctuations caused by wind and wave disturbances, sensor electronic thermal noise, etc. Pulse-type outliers: Five instantaneous spike signals with amplitudes between 20-40 cm / s are randomly injected to simulate non-physical anomalies caused by biological impact, entanglement of floating objects, and instantaneous electrical failure of sensors.

[0133] The above processing generates an original noisy sequence containing mixed noise, which serves as the input for each comparison algorithm.

[0134] Comparison of algorithms and evaluation metrics: The experiment selected four typical data processing algorithms for comparison and verification: (1) moving average filter (MA), (2) median filter (MF), (3) standard Kalman filter (SKF), and (4) the improved Kalman filter proposed in this application (including strong radial constraint threshold and Sage-Husa adaptive correction).

[0135] The root mean square error (RMSE) is used as the core evaluation metric. This metric comprehensively reflects the degree of deviation between the algorithm output and the benchmark true value. The lower the RMSE value, the better the data cleaning effect. At the same time, the percentage improvement in accuracy of each algorithm compared to the original noisy sequence is calculated to quantify the degree of improvement.

[0136] The experimental data processing results are as follows: Original noisy sequence: RMSE is 4.0893 cm / s.

[0137] Improved Kalman filtering: RMSE is significantly reduced to 2.8605 cm / s, and the accuracy is improved by 30.05% compared with noisy sequences.

[0138] Comparison of algorithm performance: Moving average (MA) accuracy improved by 14.78%, median filter (MF) by 20.82%, and standard Kalman filter (SKF) by 24.14%.

[0139] Data analysis shows that traditional filtering methods can suppress noise to some extent, but the effect is limited. Moving average filtering improves noise by 14.78%, median filtering by 20.82%, and standard Kalman filtering by 24.14%.

[0140] The improved Kalman filter algorithm proposed in this application reduces the RMSE from 4.0893 cm / s to 2.8605 cm / s, improving the accuracy by up to 30.05%, which is significantly better than all the comparison algorithms.

[0141] Compared to the standard Kalman filter, the improved algorithm in this application achieves an additional accuracy gain of approximately 6% (24.14% → 30.05%).

[0142] Experimental data fully demonstrate that the improved Kalman filter algorithm proposed in this application can intelligently distinguish between background Gaussian noise and non-physical impulse outliers by introducing a strong radial constraint threshold mechanism: for normal fluctuations in the innovation amplitude within the 3σ threshold, the Sage-Husa adaptive algorithm is used to track noise changes in real time; for abnormal outliers exceeding the 3σ threshold, the covariance update is frozen and the Kalman gain is suppressed, effectively preventing dirty data from polluting the filter parameters.

[0143] It is this "tiered processing" strategy that enables the algorithm in this application to accurately identify and suppress sudden pulse interference while filtering out high-frequency background noise, thereby achieving an additional accuracy gain of approximately 6% on top of standard Kalman filtering. Experimental results verify the superiority, robustness, and engineering applicability of the technical solution in this application for data cleaning in complex marine environments.

[0144] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys, characterized in that, include: Deep-sea wireless sensing gateway hardware platform, edge data compression module, server-side data noise cleaning module, and multi-physics field collaborative data fusion module; The deep-sea wireless sensing gateway hardware platform is deployed on deep-sea buoys or underwater moorings to collect raw data of multiple physical quantities in the deep-sea environment. The edge data compression module is used to perform structured compression and frame-level encapsulation of the original data of the multiple physical quantities based on physical constraints to generate compressed data packets; The server-side data noise cleaning module is used to receive and decompress the compressed data packet to restore the original observation data, and to identify and suppress background noise and non-physical outliers in the original observation data in real time based on the improved Kalman filter algorithm, and output the cleaned observation data. The multiphysics collaborative data fusion module is used to perform spatiotemporal registration and dynamic assimilation fusion on the cleaned observation data based on a preset physical constraint system, so as to generate a spatiotemporally continuous optimal ocean data field.

2. The system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys according to claim 1, characterized in that, The edge data compression module is specifically used for: The original data of the multiple physical quantities are identified and labeled to obtain different types of multiple physical quantity data; Based on the different types of multi-physical quantity data, the corresponding multi-level prediction models are used to process them to obtain the corresponding final prediction values. A joint residual sequence is constructed based on the final predicted value and the corresponding observation value in the original data of the multi-physical quantities, and the joint residual sequence is divided and recombined at multiple scales. Based on the statistical distribution characteristics of the recombined joint residual sequence, a lossless encoding method is adaptively selected locally at the edge node of the buoy for encoding; and according to the preset fixed frame length constraint, the encoded data stream is organized and encapsulated at the frame level to generate the compressed data packet.

3. The system for compressing, cleaning, and collaboratively fusing multiphysics field data of deep-sea buoys according to claim 2, characterized in that, The multi-level prediction model is constructed based on physical constraints and the correlation of multiple physical quantities. The processing according to the corresponding multi-level prediction model to obtain the corresponding final prediction value includes: Based on the first prediction level in the multi-level prediction model, a basic prediction value is generated according to the multi-physical quantity data, and the basic prediction value is corrected based on the second prediction level in the multi-level prediction model to generate the final prediction value of the current sampling point.

4. The system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys according to claim 1, characterized in that, The server-side data noise cleaning module is specifically used for: The compressed data packet is decompressed to obtain the original observation data, and the original observation data is preprocessed to construct a standardized observation vector. A state-space model is constructed based on the physical characteristics of the marine dynamic environment. The observation vector and the optimal state estimate of the previous time step are used as inputs. Kalman filtering is performed to update the time and generate the state prediction value and prediction error covariance of the current sampling point. Based on the deviation between the predicted state value of the current sampling point and the corresponding observed value in the observation vector, the information of the current sampling point is obtained; The theoretical standard deviation of the innovation is obtained based on the corresponding prediction error covariance and the preset observation noise covariance. A strong radial constraint threshold is constructed based on the current sampling point's information and the theoretical standard deviation. A preset multiple of the theoretical standard deviation is used as a threshold. The current sampling point's information amplitude is compared with the threshold to obtain a comparison result. Kalman filtering measurement updates are performed based on the comparison result to obtain the optimal state estimate at the current time, which is used as the cleaned observation data.

5. The system for compressing, cleaning, and collaboratively fusing multiphysics field data of deep-sea buoys according to claim 4, characterized in that, The step of updating the measurements using Kalman filtering based on the comparison results to obtain the optimal state estimate at the current moment, which is then used as the cleaned observation data, includes: If the innovation amplitude of the current sampling point does not exceed the threshold, it is determined that the observation value corresponding to the current sampling point belongs to the background noise. Based on the Sage-Husa adaptive filtering algorithm, the system noise covariance and the observation noise covariance are iteratively corrected in real time with the innovation of the current sampling point as input, so as to obtain the updated current system noise covariance and the updated current observation noise covariance. A forgetting factor is introduced to weight the historical data. If the innovation amplitude of the current sampling point exceeds the threshold, it is determined that the observation value corresponding to the current sampling point belongs to a non-physical impulse field value. The update of the system noise covariance and the observation noise covariance is frozen, the system noise covariance and the observation noise covariance at the previous moment are kept unchanged, and the Kalman gain is suppressed. Based on the updated or frozen current system noise covariance, current observation noise covariance, information of the current sampling point, and the state prediction value, Kalman filtering measurement updates are performed to obtain the optimal state estimate at the current moment, which is used as the cleaned observation data.

6. The system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys according to claim 1, characterized in that, The multiphysics collaborative data fusion module is specifically applied to: Different types of data in the cleaned observation data are preprocessed separately to obtain corresponding standardized ocean data; The standardized marine data is unified in terms of spatiotemporal reference, and a grid background field is generated through a unified spatiotemporal registration and grid construction process. A unified state vector is constructed based on the grid background field. Under the guidance of the preset physical constraint system, a unified dynamic data assimilation loop is executed to obtain the optimized unified state vector. The dynamic data assimilation loop is repeated until the current unified state vector satisfies the convergence condition, thereby obtaining the optimal ocean data field.

7. The system for compressing, cleaning, and collaboratively fusing multiphysics field data of deep-sea buoys according to claim 6, characterized in that, The dynamic data assimilation loop includes: Using the current state value in the unified state vector as input, the state evolution is predicted based on the physical model through ensemble Kalman filtering to obtain the predicted state vector; then, the predicted state vector is updated according to the real-time observation data of the corresponding spatiotemporal nodes in the grid background field to obtain the updated state vector; the core parameters in the updated state vector are corrected for prediction errors through a long short-term memory network model to obtain the optimized unified state vector.

8. The system for compressing, cleaning, and collaboratively fusing multiphysics data from deep-sea buoys according to claim 1, characterized in that, The preset physical constraint system includes at least one of the following: hydrostatic equilibrium equation, simplified shallow water wave equation, geostrophic equilibrium equation, and mass conservation continuity equation.

9. The system for compressing, cleaning, and collaboratively fusing multiphysics field data of deep-sea buoys according to claim 2, characterized in that, The adaptive selection of lossless coding method includes selecting an appropriate coding method from fixed-length bit-width coding, variable-length bit-width coding, or entropy coding based on at least one of the maximum amplitude, zero-value ratio, and dispersion index of the joint residual sequence, as well as the current computing resource or power consumption status of the edge node.

10. The system for compressing, cleaning, and collaboratively fusing multiphysics field data of deep-sea buoys according to claim 7, characterized in that, The core parameters include wave height, wave velocity, and pressure.