A Data-Driven Approach to Monitoring Marine Phytoplankton Abundance Using Underwater Acoustic Networks
By constructing a multi-level spatiotemporal feature integration model and adaptive modulation coding, the location of underwater acoustic network nodes and communication strategies were optimized, solving the problem of unstable operation of underwater acoustic networks in complex environments and realizing efficient and real-time monitoring of marine phytoplankton abundance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies are insufficient for high-resolution, full-coverage real-time monitoring of marine phytoplankton abundance. Furthermore, underwater acoustic networks are unstable in complex environments, have low energy efficiency, and suffer from unreasonable node deployment.
A multi-level spatiotemporal feature integration model is constructed to optimize the location and time window of underwater acoustic network nodes, match suitable underwater acoustic communication models and routing strategies, perform adaptive modulation and coding, and combine spatiotemporal knowledge distillation technology to reduce energy consumption and achieve efficient data transmission.
It has achieved efficient and stable operation of the underwater acoustic network, improved the spatiotemporal representativeness and coverage efficiency of monitoring data, and provided real-time and adaptive phytoplankton abundance monitoring capabilities.
Smart Images

Figure CN121542644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of marine environmental monitoring and underwater acoustic communication, and in particular to a method for monitoring the abundance of marine phytoplankton based on data-driven and underwater acoustic networks. Background Technology
[0002] For decades, monitoring of marine phytoplankton abundance has primarily relied on satellite remote sensing, field sampling, and automated monitoring systems. While commonly used field sampling and fixed monitoring station methods can provide high-precision, long-term data on phytoplankton abundance and its corresponding environmental conditions, these methods are limited by high cost, long cycles, and limited spatial coverage, making it difficult to achieve large-scale, high-density, long-term real-time monitoring, thus missing many key marine phenomena. Satellite remote sensing has played a crucial role in global-scale research, providing long-term, large-scale data support and reflecting ocean surface characteristics well. However, its observation accuracy is susceptible to interference from clouds and aerosols, and it suffers from drawbacks such as large equipment size, high power consumption, and high cost. Therefore, achieving high-resolution, full-coverage real-time monitoring has become a pressing scientific and technological bottleneck that needs to be overcome.
[0003] Against this backdrop, underwater acoustic networks and automated monitoring systems composed of cableless and cabled nodes offer new possibilities for marine ecological monitoring. This system can stably transmit phytoplankton abundance and corresponding marine environmental sensor data over long periods, offering advantages such as miniaturization, low cost, and ease of deployment. Underwater acoustic communication, currently the only feasible long-distance underwater communication method, has been applied in various scenarios, including extreme environment communication. However, most underwater nodes rely on battery power, and replacing batteries underwater is both time-consuming and costly. Furthermore, complex environmental factors such as wave impacts can easily lead to network instability. Digital empowerment is a key means to improve the real-time, accurate, and intelligent monitoring of the marine environment. Traditional node deployment, which relies heavily on historical data and expert experience (such as proximity to land-based pollution sources), is ill-suited to the complex and ever-changing marine environment, limiting the improvement of monitoring efficiency. Therefore, improving the energy efficiency of underwater acoustic networks and achieving scientific node deployment have become core issues for ensuring their long-term stable operation and efficient transmission of monitoring data. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention proposes a data-driven method for monitoring the abundance of marine phytoplankton using underwater acoustic networks. The specific technical solution is as follows:
[0005] A data-driven method for monitoring marine phytoplankton abundance using underwater acoustic networks includes:
[0006] S1: Based on historical marine phytoplankton abundance and its corresponding marine environment data, a multi-level spatiotemporal feature integration model is constructed and trained to solve for the optimal spatial location and monitoring time window of the underwater acoustic network nodes; the multi-level spatiotemporal feature integration model includes three parallel feature extraction branches: a three-dimensional convolutional neural network, a convolutional long short-term memory network, and a Transformer.
[0007] S2: Construct a teacher model for channel prediction, including a multi-layer cascaded structure and residual modules, and use spatiotemporal knowledge distillation to compress the teacher model into a student model consisting of a two-layer one-dimensional convolutional neural network and a long short-term memory network; train the student model to achieve real-time channel state prediction at underwater acoustic network nodes.
[0008] S3: Based on the optimal spatial location of the underwater acoustic network nodes and the real-time predicted channel status of the underwater acoustic network nodes, match suitable underwater acoustic communication devices and seasonal optimal routing strategies for each underwater acoustic network node, and perform adaptive modulation and coding at the link level. At the same time, rely on the underwater acoustic network to complete the monitoring of phytoplankton abundance.
[0009] Further, step S1 includes the following steps:
[0010] S1.1: Map the latitude and longitude of all samples in the historical marine phytoplankton abundance dataset to a grid with the first spatial resolution.
[0011] S1.2: Construct the test dataset and training dataset respectively;
[0012] S1.3: Match the corresponding marine environmental data based on the latitude and longitude information of the marine phytoplankton abundance data, and generate marine environmental data images at multiple different second spatial resolution scales with the latitude and longitude of the marine phytoplankton abundance sample as the center.
[0013] S1.4: Construct a multi-level spatiotemporal feature ensemble model, and input marine environmental data images generated at multiple different second spatial resolution scales into the multi-level spatiotemporal feature ensemble model respectively. Select the second spatial resolution corresponding to the multi-level spatiotemporal feature ensemble model with the best prediction performance as the optimal second spatial resolution. The multi-level spatiotemporal feature ensemble model includes three parallel feature extraction branches: a three-dimensional convolutional neural network, a convolutional long short-term memory network, and a Transformer.
[0014] S1.5: Arrange the underwater acoustic network nodes in the optimal spatial position according to the principle that the spatial distance between adjacent underwater acoustic network nodes is equal to the optimal second spatial resolution;
[0015] S1.6: Introducing the time window method, by adjusting the number of data years involved in the training, the multi-level spatiotemporal feature integration model trained in S1.4 is retrained to obtain the monitoring time window of the underwater acoustic network node at the optimal spatial location.
[0016] Furthermore, the multi-layered cascaded structure is constructed by stacking multiple basic modules, which consist of a multi-scale one-dimensional convolutional neural network, a long short-term memory network, batch normalization, and a linear rectified function.
[0017] Furthermore, the kernel sizes of the multi-scale one-dimensional convolutional neural network are 1, 3, 5, 7, and 11.
[0018] Furthermore, the loss function for training the student model is a weighted sum of soft-target distillation loss and hard-target loss; the soft-target distillation loss is used to quantify the prediction difference between the student model output and the teacher model output, and the hard-target loss is used to quantify the difference between the student model output and the real channel.
[0019] Furthermore, in step S3, matching and adapting underwater acoustic communication devices for each node specifically includes:
[0020] S3.1: Establish a set of candidate underwater acoustic communication devices that include bandwidth, transmit power, modulation and coding parameters, power amplifier efficiency, and circuit power consumption;
[0021] S3.2: Obtain the link channel gain, ambient noise power spectral density, and interference power spectral density while the candidate underwater acoustic communication device is running;
[0022] S3.3: Based on the parameters in S3.2, calculate the link capacity and total power consumption of each candidate underwater acoustic communication device, and evaluate each candidate underwater acoustic communication device using the energy efficiency index of capacity / total power consumption; select the underwater acoustic communication device with the highest energy efficiency that meets the preset constraints for the link.
[0023] Furthermore, in S3, matching an appropriate seasonal optimal routing strategy for each underwater acoustic network node specifically includes:
[0024] S3.4: At time t, obtain the environmental state and estimate the link utility and unit transmission energy consumption for each candidate link and its modulation and coding scheme;
[0025] S3.5: Using throughput-energy consumption weighted index as path score, the optimal path from source to destination and the modulation scheme of each hop are obtained through weighted shortest path, and updated online when environmental or residual energy changes exceed the threshold.
[0026] Furthermore, in S3, adaptive modulation and coding are performed at the link level, specifically including:
[0027] S3.6: Receive real-time channel prediction information;
[0028] S3.7: Select the optimal modulation and coding combination from a predefined set of modulation and coding schemes based on a neural network model;
[0029] S3.8: Control the underwater acoustic communication device to perform a switching operation.
[0030] Furthermore, it also includes:
[0031] S4: Periodically update and optimize all model parameters of S1~S3 online based on the marine phytoplankton detection data returned by the underwater acoustic network nodes.
[0032] A marine phytoplankton abundance monitoring device based on data-driven and underwater acoustic networks is characterized by comprising one or more processors for implementing a marine phytoplankton abundance monitoring method based on data-driven and underwater acoustic networks.
[0033] The beneficial effects of this invention are as follows:
[0034] (1) This invention introduces a spatiotemporal knowledge distillation mechanism to construct a teacher-student network structure, which significantly reduces the computational complexity and energy consumption of the model while ensuring prediction accuracy. It effectively overcomes the limitations of underwater acoustic communication devices in terms of energy and hardware resources, and provides an efficient and reliable model foundation for realizing real-time and adaptive modulation and coding.
[0035] (2) By fully exploring the spatiotemporal characteristics of monitoring data and comprehensively considering the dual influence of natural climate and human activities on phytoplankton abundance, this invention achieves coordinated dynamic optimization of the spatial distribution location of nodes and the frequency of time monitoring, thereby improving network coverage efficiency and spatiotemporal representativeness of data collection under limited resources.
[0036] (3) Based on the optimization results of node spatiotemporal distribution, this invention further matches the optimal communication model and routing protocol to form an integrated design from node deployment to communication transmission, providing data support and system-level solutions for the high-efficiency, low-latency and stable operation of underwater observation networks. Attached Figure Description
[0037] Figure 1 This is a flowchart of the data-driven and underwater acoustic network-based method for monitoring the abundance of marine phytoplankton in this embodiment.
[0038] Figure 2 This is a schematic diagram of a three-dimensional convolutional neural network and a convolutional long short-term memory network.
[0039] Figure 3 This is a schematic diagram of a Transformer.
[0040] Figure 4This is a schematic diagram of images corresponding to three different second spatial resolutions.
[0041] Figure 5 This is a schematic diagram of a multi-layered cascaded structure. Detailed Implementation
[0042] The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. The purpose and effects of the present invention will become clearer. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0043] Explanation of technical terms:
[0044] 3D Convolution: Three-dimensional convolution;
[0045] 2D Convolution: Two-dimensional convolution;
[0046] ConvLSTM (Convolutional Long Short-Term Memory) is a convolutional long short-term memory network.
[0047] Flatten, a flattened layer;
[0048] Fully Connected;
[0049] MaxPooling 3D: Max pooling operation for 3D data;
[0050] MLP, or Multilayer Perceptron, is the multilayer perceptron part that transforms the feature maps processed by the Transformer layer into the final classification result.
[0051] Norm, Normalization;
[0052] Multi-Head Attention;
[0053] Patch+Position Embedding divides the image into small patches and configures their positions for embedding.
[0054] Transformer Encoder;
[0055] Liner Projection of Flattened Patches: The image is divided into small patches, then these patches are flattened, and finally the flattened vectors are mapped to a lower-dimensional space through linear projection.
[0056] Filter.
[0057] like Figure 1 As shown, the marine phytoplankton abundance monitoring method based on data-driven and underwater acoustic networks of the present invention includes the following steps one through four.
[0058] Step 1: Based on historical marine phytoplankton abundance and corresponding marine environmental data, construct and train a multi-level spatiotemporal feature ensemble model to solve for the optimal spatial location and monitoring time window of the underwater acoustic network nodes.
[0059] Step one includes the following sub-steps:
[0060] S1.1: Map the latitude and longitude of all samples in the historical marine phytoplankton abundance dataset to a grid with the first spatial resolution.
[0061] The preferred spatial resolution here is 0.25°×0.25° or 1°×1°.
[0062] S1.2: Construct test datasets and training datasets respectively. The test dataset consists of historical samples from the latest year of the mapped marine phytoplankton abundance dataset, and is uniformly labeled as high-risk areas (labeled as 1) and low-risk areas (labeled as 0). The other years' mapped marine phytoplankton abundance data constitute the training set, which is also labeled in the same way.
[0063] S1.3: Match the corresponding marine environmental data based on the latitude and longitude information of the marine phytoplankton abundance data, and generate marine environmental data images at multiple different second spatial resolution scales with the latitude and longitude of the marine phytoplankton abundance sample as the center.
[0064] The second spatial resolution here is an integer multiple of the first spatial resolution. For example, if the first spatial resolution is 0.25°×0.25°, the second spatial resolution here is 0.25°×0.25°, 1°×1°, and 5°×5°; if the first spatial resolution is 1°×1°, the second spatial resolution here can be 1°×1°, 2°×2°, and 5°×5°.
[0065] The marine environmental data here includes silicates, phosphates, nitrates, salinity, sea surface temperature, and dissolved oxygen.
[0066] S1.4: Construct a multi-level spatiotemporal feature integration model, and input marine environmental data images generated at multiple different second spatial resolution scales into the multi-level spatiotemporal feature integration model respectively. Select the second spatial resolution corresponding to the multi-level spatiotemporal feature integration model with the best prediction performance as the optimal second spatial resolution. The multi-level spatiotemporal feature integration model includes three parallel feature extraction branches: a three-dimensional convolutional neural network, a convolutional long short-term memory network, and a Transformer.
[0067] Among them, the three-dimensional convolutional neural network, the convolutional long short-term memory network, and the Transformer are used to perform hierarchical modeling and ensemble learning for local spatial structure features, long temporal dependencies, and long sequence global correlations, respectively. The input of the three networks is the same set of marine environmental data images, and the output is a binary classification result. Figure 4 Marine environmental data images generated at three different second spatial resolution scales are presented.
[0068] The output of the trained multi-level spatiotemporal feature ensemble model corresponds to the optimal second spatial resolution, i.e. the second spatial resolution with the highest F1 score.
[0069] The network structures of 3D convolutional neural networks and convolutional long short-term memory networks are as follows: Figure 2 As shown, the network architecture of Transformer is as follows: Figure 3 As shown. The predictive performance of the multi-level spatiotemporal feature ensemble model is presented for three different second spatial resolutions (5°×5°, 1°×1°, 0.25°×0.25°). To obtain the prediction performance at the second spatial resolution of the optimal marine environment, i.e., the highest F1 score, the specific formula is:
[0070]
[0071] S1.5: After obtaining the optimal second spatial resolution, arrange the optimal spatial positions of the underwater acoustic network nodes according to the principle that the spatial distance between adjacent underwater acoustic network nodes is equal to the optimal second spatial resolution.
[0072] S1.6: Introducing the time window method, by adjusting the number of data years involved in the training, the multi-level spatiotemporal feature integration model trained in S1.4 is retrained so that the selected time window can capture the seasonal change characteristics of phytoplankton in the area as completely and clearly as possible, that is, the highest F1 score, and the monitoring time window of the underwater acoustic network node at the optimal spatial location is obtained.
[0073] Step 2: Construct a teacher model for channel prediction, including a multi-layer cascaded structure and residual modules, and a student model consisting of a two-layer one-dimensional convolutional neural network (1D CNN) and a long short-term memory network (LSTM). Use spatiotemporal knowledge distillation to compress the teacher model into a student model. Real-time channel state prediction is achieved at underwater acoustic network nodes, reducing energy consumption for transmitting phytoplankton and other data based on underwater acoustic communication devices.
[0074] Among them, the multi-level cascaded structure of the teacher model enables multi-scale representation of channel dynamics and fully explores historical channel information; the residual module extracts deep features.
[0075] When constructing the teacher model for channel prediction, the channel is uniformly defined in the time domain as a coherent multipath channel, as shown in the formula:
[0076]
[0077] In the formula, P is the number of channel taps, t is the time to observe the channel, and τ is the delay variable; h p (t) and τ p (t) represents the p-th tap at time t and the corresponding delay, respectively.
[0078] Discretizing h(τ, t) over time, the channel h(n) at time n can be expressed as: .
[0079] When predicting the channel, it is necessary to use the estimated channels from the past M times, for example, predicting the channel at time n+N. When, it is necessary to use { The channel prediction model is trained using the formula:
[0080]
[0081] When predicting the p-th tap at time n+N At that time, the input of the predictor is The output of the predictor is This refers to the predicted value of the p-th tap at time n+N. The prediction process of the predictor can be expressed by the formula:
[0082]
[0083]
[0084] like Figure 5 As shown, the teacher model for channel prediction constructs a multi-layered cascaded structure by stacking multiple basic modules and uses residual modules for deep feature extraction, thereby accelerating the model's convergence speed while effectively avoiding gradient vanishing / exploding and network degradation problems. Figure 5As shown, the basic modules consist of a multi-scale one-dimensional convolutional neural network (kernel sizes 1, 3, 5, 7, 11), a long short-term memory network (LSTM), batch normalization (BN), and a linear rectified function (ReLU). The multi-scale one-dimensional convolutional neural network is used to increase the model width, and the combination of the one-dimensional convolutional neural network and the long short-term memory network is used to enhance the spatiotemporal representation and generalization capabilities. The relationships between the various basic modules are formed by stacking them to create a deep network structure, with input data... The input layer processes data sequentially through multiple basic modules, each responsible for extracting specific features, ultimately forming the teacher model's prediction result X. st_tch .
[0085] In the basic module, the channel state data at time t is input. The data is first fed into multiple 1D convolutional layers (CNN), each using different filters and kernel sizes (k=3, filters=3; k=5, filters=5; k=7, filters=7; k=11, filters=11). These convolutional layers process the input data in parallel and extract features. Next, the outputs of the convolutional layers are fused through addition and passed to subsequent 1D convolutional layers (k=1, filters=1) for further processing and extraction of finer-grained features. The output is then standardized using batch normalization (BN) to reduce internal covariate bias. Subsequently, the data undergoes a non-linear transformation using the ReLU activation function to enhance the model's expressive power. Next, an LSTM layer processes the temporal data, enhancing the global memory capability of multi-scale features. Finally, the output undergoes another batch normalization and ReLU activation to generate the final predicted output. .
[0086] The student model consists of two layers of one-dimensional convolutional neural networks (1D CNN) and a long short-term memory network (LSTM). First, the input data... After processing by the first 1D CNN layer, features are extracted using convolutional kernels of size 3. Next, the output of the first layer is passed as input to the second 1D CNN layer, which uses convolutional kernels of size 1. Finally, the extracted features are passed to an LSTM layer, which enhances global memory and outputs the student model's prediction.
[0087] Spatiotemporal knowledge distillation is used to compress the teacher model into a student model, enabling the student model to effectively inherit the key information extracted from the teacher model, while maintaining high prediction performance while reducing computational complexity and energy consumption.
[0088] Spatiotemporal knowledge distillation requires comprehensive control of the proportions of soft target loss and hard target loss in the total loss. The specific formula is:
[0089]
[0090]
[0091]
[0092] In the formula, The soft-target distillation loss is used to quantify the prediction difference between the student model output and the teacher model output. This represents the predicted output of the student model at time t+k. This represents the teacher model's predicted output at time t+k. This is a hard target loss used to quantify the difference between the student model output and the real channel; is the total loss function, which combines the loss functions of soft targets and hard targets; K is the prediction step size, representing the number of future prediction time points; T is the distillation temperature coefficient, used to control the smoothness of the soft target distribution; The actual channel value is used as a hard target; This is a weighting factor that controls the proportion of soft and hard targets in the total loss.
[0093] Step 3: Based on the optimal spatial location of the underwater acoustic network nodes obtained in Step 1 and the teacher and student models of channel prediction obtained in Step 2, match suitable underwater acoustic communication devices and seasonal optimal routing strategies for each underwater acoustic network node, and perform adaptive modulation and coding at the link level. At the same time, rely on the underwater acoustic network to complete the monitoring of phytoplankton abundance.
[0094] Step three involves matching suitable underwater acoustic communication devices to each underwater acoustic network node, including the following sub-steps:
[0095] S3.1: Establish a set of candidate underwater acoustic communication devices that include bandwidth, transmit power, modulation and coding parameters, power amplifier efficiency, and circuit power consumption.
[0096] S3.2: Obtain the link channel gain, ambient noise power spectral density, and interference power spectral density while the candidate underwater acoustic communication device is running.
[0097] S3.3: Based on the parameters in S3.2, calculate the link capacity and total power consumption (the sum of transmit power and circuit power consumption converted according to power amplifier efficiency) of each candidate underwater acoustic communication device, and evaluate each candidate underwater acoustic communication device using the energy efficiency index of capacity / total power consumption; select the underwater acoustic communication device with the highest energy efficiency that meets the preset constraints for the link.
[0098] The above process can be expressed by the following formula:
[0099]
[0100] In the formula, The optimal underwater acoustic communication unit model for the link from node i to j. For bandwidth, For transmission power, For power amplifier efficiency. For circuit power consumption, For channel gain, For noise power spectral density, This represents the interference power spectral density.
[0101] Step three involves matching each underwater acoustic network node with an appropriate seasonal optimal routing strategy, which specifically includes the following sub-steps:
[0102] S3.4: At time t, acquire the environmental conditions (temperature, salinity, ocean current, etc.), and estimate the link utility η and unit transmission energy consumption E for each candidate link and its modulation and coding scheme.
[0103] S3.5: The path score is based on the "throughput-energy consumption" weighted index. The optimal path from source to destination and the modulation scheme of each hop are obtained through the weighted shortest path, and the path is updated online when the environmental or residual energy changes exceed the threshold.
[0104] The expression for the optimal path is:
[0105]
[0106] In the formula, t represents the moment / time window for routing decisions;
[0107] This represents the optimal routing strategy selected at time t;
[0108] (i,j) represents a directed link in the path (from node i to node j);
[0109] : The specific transmission method selected for link (i,j) at time t.
[0110] This indicates the "utility" of the link under mode m and environment t (e.g., achievable throughput / number of successfully transmitted bits / effective rate). This represents the "energy cost" of the link under mode m and environment t;
[0111] λ represents the throughput-energy consumption tradeoff coefficient, which can be fixed or adaptively adjusted with the remaining energy to achieve the optimal balance between throughput and energy consumption.
[0112] Step three, performing adaptive modulation and coding at the link level, specifically includes:
[0113] S3.6: Receive real-time channel prediction information;
[0114] S3.7: Select the optimal modulation and coding combination from a predefined set of modulation and coding schemes based on a neural network model;
[0115] S3.8: Control the underwater acoustic communication device to perform a switching operation.
[0116] Step 4: Based on the marine phytoplankton detection data transmitted back from the underwater acoustic network nodes, all model parameters from Step 1 to Step 3 are periodically updated and optimized online to form a closed loop of "deployment-perception-decision-redeployment".
[0117] The periodic online updates and optimizations here take place every 1 week to 3 months.
[0118] Another embodiment of the present invention provides a marine phytoplankton abundance monitoring device based on data-driven and underwater acoustic networks, including one or more processors for implementing the marine phytoplankton abundance monitoring method based on data-driven and underwater acoustic networks in the above embodiments.
[0119] The marine phytoplankton abundance monitoring device based on data-driven and underwater acoustic networks in this embodiment can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data-processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. Besides the processor, memory, network interface, and non-volatile memory, the data-processing device in this embodiment may also include other hardware depending on its actual functions, which will not be elaborated further.
[0120] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0121] It will be understood by those skilled in the art that the above descriptions are merely preferred examples of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A method for monitoring the abundance of marine phytoplankton based on data-driven methods and underwater acoustic networks, characterized in that, include: S1: Based on historical marine phytoplankton abundance and its corresponding marine environment data, a multi-level spatiotemporal feature integration model is constructed and trained to solve for the optimal spatial location and monitoring time window of the underwater acoustic network nodes; the multi-level spatiotemporal feature integration model includes three parallel feature extraction branches: a three-dimensional convolutional neural network, a convolutional long short-term memory network, and a Transformer. S2: Construct a teacher model for channel prediction, including a multi-layer cascaded structure and residual modules, and use spatiotemporal knowledge distillation to compress the teacher model into a student model consisting of a two-layer one-dimensional convolutional neural network and a long short-term memory network; train the student model to achieve real-time channel state prediction at underwater acoustic network nodes. S3: Based on the optimal spatial location of the underwater acoustic network nodes and the real-time predicted channel status of the underwater acoustic network nodes, match suitable underwater acoustic communication devices and seasonal optimal routing strategies for each underwater acoustic network node, and perform adaptive modulation and coding at the link level. At the same time, rely on the underwater acoustic network to complete the monitoring of phytoplankton abundance. S1 includes the following steps: S1.1: Map the latitude and longitude of all samples in the historical marine phytoplankton abundance dataset to a grid with the first spatial resolution. S1.2: Construct the test dataset and training dataset respectively; S1.3: Match the corresponding marine environmental data based on the latitude and longitude information of the marine phytoplankton abundance data, and generate marine environmental data images at multiple different second spatial resolution scales with the latitude and longitude of the marine phytoplankton abundance sample as the center. S1.4: Construct a multi-level spatiotemporal feature ensemble model, and input marine environmental data images generated at multiple different second spatial resolution scales into the multi-level spatiotemporal feature ensemble model respectively. Select the second spatial resolution corresponding to the multi-level spatiotemporal feature ensemble model with the best prediction performance as the optimal second spatial resolution. The multi-level spatiotemporal feature ensemble model includes three parallel feature extraction branches: a three-dimensional convolutional neural network, a convolutional long short-term memory network, and a Transformer. S1.5: Arrange the underwater acoustic network nodes in the optimal spatial position according to the principle that the spatial distance between adjacent underwater acoustic network nodes is equal to the optimal second spatial resolution; S1.6: Introducing the time window method, by adjusting the number of data years involved in the training, the multi-level spatiotemporal feature integration model trained in S1.4 is retrained to obtain the monitoring time window of the underwater acoustic network node at the optimal spatial location.
2. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 1, characterized in that, The multi-layered cascaded structure is constructed by stacking multiple basic modules, which consist of a multi-scale one-dimensional convolutional neural network, a long short-term memory network, batch normalization, and a linear rectified function.
3. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 2, characterized in that, The kernel sizes of the multi-scale one-dimensional convolutional neural network are 1, 3, 5, 7, and 11.
4. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 1, characterized in that, The loss function used to train the student model is a weighted sum of soft-target distillation loss and hard-target loss; the soft-target distillation loss is used to quantify the prediction difference between the student model output and the teacher model output, and the hard-target loss is used to quantify the difference between the student model output and the real channel.
5. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 1, characterized in that, In step S3, the underwater acoustic communication device models matched and adapted for each node specifically include: S3.1: Establish a set of candidate underwater acoustic communication devices that include bandwidth, transmit power, modulation and coding parameters, power amplifier efficiency, and circuit power consumption; S3.2: Obtain the link channel gain, ambient noise power spectral density, and interference power spectral density while the candidate underwater acoustic communication device is running; S3.3: Based on the parameters in S3.2, calculate the link capacity and total power consumption of each candidate underwater acoustic communication device, and evaluate each candidate underwater acoustic communication device using the energy efficiency index of capacity / total power consumption; select the underwater acoustic communication device with the highest energy efficiency that meets the preset constraints for the link.
6. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 1, characterized in that, In step S3, a seasonal optimal routing strategy is matched and adapted for each underwater acoustic network node, specifically including: S3.4: In t The system continuously acquires environmental status and estimates link utility and unit transmission energy consumption for each candidate link and its modulation and coding scheme. S3.5: Using throughput-energy consumption weighted index as path score, the optimal path from source to destination and the modulation scheme of each hop are obtained through weighted shortest path, and updated online when environmental or residual energy changes exceed the threshold.
7. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 6, characterized in that, In step S3, adaptive modulation and coding are performed at the link level, specifically including: S3.6: Receive real-time channel prediction information; S3.7: Select the optimal modulation and coding combination from a predefined set of modulation and coding schemes based on a neural network model; S3.8: Control the underwater acoustic communication device to perform a switching operation.
8. The method for monitoring marine phytoplankton abundance based on data-driven and underwater acoustic networks according to claim 1, characterized in that, Also includes: S4: Periodically update and optimize all model parameters of S1~S3 online based on the marine phytoplankton detection data returned by the underwater acoustic network nodes.
9. A marine phytoplankton abundance monitoring device based on data-driven and underwater acoustic networks, characterized in that, It includes one or more processors for implementing the data-driven and underwater acoustic network-based method for monitoring the abundance of marine phytoplankton as described in any one of claims 1 to 8.