An ocean ecological monitoring method based on ecological representation and target perception matching

By employing a multi-module collaborative optimization method for marine ecological monitoring, the problems of low data synchronization accuracy and difficulty in target identification in marine ecological monitoring have been solved, achieving efficient and accurate marine ecological monitoring and supporting red tide early warning and biodiversity conservation.

CN121834243BActive Publication Date: 2026-05-12SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)
Filing Date
2026-03-13
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing marine ecological monitoring technologies suffer from low synchronization accuracy and weak anti-interference ability in the multimodal data preprocessing stage. In the ecological feature extraction, the distinction between target and background is blurred. In the target matching and model iteration, multi-scale adaptation is poor and generalization ability is weak, making it difficult to achieve efficient and accurate marine ecological monitoring.

Method used

A multi-module collaborative optimization approach is adopted, including multimodal data preprocessing, distinguishable ecological feature representation, and target perception matching. Through spatiotemporal synchronous calibration, improved polarization weighted filtering, task-adaptive feature decoupling network, dynamic anchor box and hierarchical diffusion matching, combined with lightweight design and knowledge distillation technology, efficient and accurate ecological monitoring is achieved.

Benefits of technology

It significantly improves data synchronization accuracy and signal-to-noise ratio, enhances target identification capabilities and monitoring accuracy, reduces energy consumption and deployment costs, and enables routine, wide-coverage monitoring of marine ecosystems, supporting red tide early warning and biodiversity conservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834243B_ABST
    Figure CN121834243B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of marine ecological monitoring, and more particularly to a marine ecological monitoring method based on ecological representation and target perception matching. The method comprises the following steps: acquiring multi-source marine monitoring data; performing multi-modal data preprocessing based on the acquired multi-source marine monitoring data; generating distinguishable ecological feature vectors according to the preprocessed data; generating a fusion confidence using the distinguishable ecological feature vectors; generating a monitoring report based on the fusion confidence result and feeding back. The present application realizes multiple breakthroughs in the accuracy, efficiency and practicability of marine ecological monitoring through multi-module collaborative optimization, and is significantly superior to the traditional technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine ecological monitoring technology, and in particular to a marine ecological monitoring method based on matching ecological representation with target perception. Background Technology

[0002] Marine ecosystems are the core carriers of fishery resource supply and climate regulation. Traditional monitoring relies mainly on manual sampling and single-device systems, which are insufficient in both efficiency and accuracy. Manual sampling depends on research vessels operating at fixed points, with a single voyage covering only 60 to 90 square kilometers of sea area. The response to sudden red tides is delayed by more than 48 hours, and the false negative rate for micron-sized plankton reaches 58%. Single-sensor monitoring also has significant limitations: hydrological buoys can only collect basic parameters such as temperature and salinity, and cannot directly identify biological targets; underwater cameras are limited by seawater transparency, with an effective monitoring distance generally less than 5 meters, and are completely ineffective in turbid waters; optical sensors have a false positive rate of up to 37% for red tides, often mistaking common algae aggregations for harmful red tides. Traditional technologies are also costly, with a monitoring cost of approximately 800 yuan per square kilometer, nine times that of intelligent monitoring solutions, making it difficult to achieve large-scale, routine coverage.

[0003] However, existing marine ecological monitoring technologies suffer from low synchronization accuracy and weak anti-interference capabilities in the multimodal data preprocessing stage. Traditional methods for integrating satellite remote sensing, underwater acoustic, and buoy sensing data rely solely on simple timestamp alignment, often resulting in spatiotemporal synchronization errors exceeding 5 seconds. Furthermore, the spatial coordinates are not unified to the WGS-84 coordinate system, leading to data correlation failure. To address interference from cloud cover and ship noise in the marine environment, the commonly used filtering algorithms can only improve the signal-to-noise ratio by about 15%, which is insufficient for subsequent feature extraction requirements.

[0004] Existing technologies face core problems in ecological feature extraction, namely, blurred distinction between target and background and low utilization of true features. Conventional feature extraction models do not have separation mechanisms designed for marine scenes, mixing ecological target features with background features such as ocean current textures and temperature gradients, resulting in a target-background feature distinction of less than 50%. Even with the introduction of masking techniques, the lack of adaptive capabilities leads to the loss of key features such as red tide spectra and fish acoustic pulses, resulting in a true feature utilization rate of only about 60%.

[0005] Existing technologies suffer from shortcomings in target matching and model iteration, including poor multi-scale adaptation, weak generalization ability, and low update efficiency. Traditional matching algorithms use a fixed anchor frame design, which cannot cover multi-scale targets ranging from 16×16 pixel micrometer-level algae to 256×256 pixel meter-level red tides, resulting in a false negative rate of over 45% for small targets. When applied across sea areas, the lack of a dynamic update mechanism leads to a decrease in recognition accuracy of over 50%, and model parameter updates require full retraining, taking more than 15 days. Summary of the Invention

[0006] To address the aforementioned problems, this invention provides a marine ecological monitoring method based on matching ecological representation with target perception.

[0007] In a first aspect, the present invention provides a marine ecological monitoring method based on ecological representation and target perception matching, which adopts the following technical solution:

[0008] A marine ecological monitoring method based on matching ecological representation with target perception includes:

[0009] Acquire multi-source ocean monitoring data;

[0010] Multimodal data preprocessing was performed based on the acquired multi-source marine monitoring data;

[0011] Generate distinguishable ecological feature vectors based on the preprocessed data;

[0012] Target perception is generated and fusion confidence is achieved using distinguishable ecological feature vectors;

[0013] A monitoring report is generated and fed back based on the fusion confidence results.

[0014] Secondly, a marine ecological monitoring system based on matching ecological representation with target perception includes:

[0015] The data acquisition module is configured to acquire multi-source ocean monitoring data;

[0016] The preprocessing module is configured to perform multimodal data preprocessing based on the acquired multi-source ocean monitoring data;

[0017] The feature module is configured to generate distinguishable ecological feature vectors based on the preprocessed data.

[0018] The confidence module is configured to generate fusion confidence scores by using distinguishable ecological feature vectors for target perception.

[0019] The monitoring module is configured to generate and provide feedback monitoring reports based on the fusion confidence results.

[0020] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the marine ecological monitoring method based on ecological representation and target perception matching.

[0021] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a marine ecological monitoring method based on ecological representation and target perception matching.

[0022] In summary, the present invention has the following beneficial technical effects:

[0023] This invention achieves multiple breakthroughs in the accuracy, efficiency, and practicality of marine ecological monitoring through multi-module collaborative optimization, significantly outperforming traditional technologies. At the multimodal data processing level, the spatiotemporal stamp dynamic calibration technology compresses data synchronization errors from over 5 seconds to within ±0.1 seconds. Combined with WGS-84 coordinate system processing, it completely solves the problem of multi-source data correlation failure. The improved polarization weight enhancement algorithm specifically filters interference from clouds, ships, etc., improving the signal-to-noise ratio by over 42%, a qualitative leap compared to the 15% improvement of ordinary filtering algorithms. This provides a highly reliable data foundation for subsequent feature extraction, increasing the effective data utilization rate from less than 70% to over 95%.

[0024] Optimization in the feature extraction stage has led to a core improvement in target recognition capabilities. The task-adaptive feature decoupling network achieves precise separation of target and background features through dual-channel convolution. Combined with a self-supervised masking mechanism adapted to marine scenes, the target-background feature discrimination has increased from less than 50% to over 65%, and the utilization rate of true features is significantly improved compared to the 60% of traditional techniques. The EPG directivity score reaches 0.89, indicating a significant enhancement in the accuracy of capturing key features such as red tide spectra and fish acoustic pulses. This high-purity feature output reduces the false recognition rate of red tides and common algae from 35% to below 8%, providing high-quality input for subsequent target matching and improving the reliability of monitoring results from the source.

[0025] The optimization of target matching and model deployment has completely solved the practicality bottleneck of traditional technologies. The application of dynamic anchor frames and hierarchical diffusion matching system has reduced the false negative rate of small targets smaller than 10 cm from over 45% to 12%, controlled the decrease in cross-sea monitoring accuracy to within 15%, and stabilized the matching error to within 2 pixels. The automatic database entry and incremental update mechanism has improved the model's generalization ability by 40%, eliminating the need for manual intervention when facing unknown invasive species, and shortening the update cycle from 15 days to 2 days. At the same time, the combination of lightweight design and knowledge distillation technology has kept the number of model parameters to within 15M while ensuring a single frame processing rate of over 17 FPS and power consumption of ≤5W. It is perfectly compatible with embedded devices such as buoys, reducing the deployment cost by 60% compared to traditional models. It has achieved a balance of "high precision, high real-time performance, and low power consumption", providing a feasible solution for the normalized and wide-coverage monitoring of marine ecology and helping to carry out red tide early warning, biodiversity conservation, and other work efficiently. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of a marine ecological monitoring method based on ecological representation and target perception matching according to Embodiment 1 of the present invention;

[0027] Figure 2 This is a schematic diagram of step S1 in Embodiment 1 of the present invention;

[0028] Figure 3 This is a schematic diagram of step S2 in Embodiment 1 of the present invention;

[0029] Figure 4 This is a schematic diagram of step S3 in Embodiment 1 of the present invention.

[0030] Figure 5 This is a comparison chart of the experimental verification and predicted distribution of Embodiment 1 of the present invention. Detailed Implementation

[0031] The present invention will be further described in detail below with reference to the accompanying drawings.

[0032] Example 1

[0033] Reference Figure 1 This embodiment of a marine ecological monitoring method based on ecological representation and target perception matching includes:

[0034] S1. Multimodal data preprocessing

[0035] The multimodal data preprocessing layer forms the foundation of the intelligent marine ecological monitoring system. Its core task is to transform raw, heterogeneous, and noisy multi-source marine monitoring data into a standardized, high-quality, and directly usable unified data representation for deep feature learning. This processing flow follows a strict sequence of "spatiotemporal synchronization → noise suppression → data augmentation," aiming to systematically address the three major challenges of data alignment, signal-to-noise ratio improvement, and insufficient sample diversity. This section will elaborate on the process from raw multimodal datasets... To the final preprocessed dataset The complete and coherent mathematical transformation process and data flow.

[0036] Original multimodal dataset It contains data tuples from four main sensors: optical remote sensing images Infrared remote sensing images Underwater acoustic spectrum diagram and buoy sensing timing data Each data tuple is accompanied by a timestamp recorded by the sensor at the time of acquisition. and spatial coordinates The unprocessed data exhibits significant inconsistencies: asynchronous clocks from different sensors cause deviations in timestamp t; differing coordinate systems lead to inconsistencies in spatial coordinates. Direct fusion is not possible; the data contains interference from clouds, ship noise, etc.; and the number of samples targeting rare ecological targets is extremely limited. The goal of preprocessing is to construct a deterministic transformation that... Mapped to a clean, aligned, and rich standard dataset .

[0037] The entire preprocessing process can be viewed as a composite of three sequentially executed sub-transformations. Let... For spatiotemporal synchronization transformation operators, For noise suppression transform operator, For data augmentation transform operators, the complete preprocessing transform... Defined as:

[0038] ,

[0039] in This indicates function composition. Therefore, the final output data is:

[0040] ,

[0041] Next, this patent will define and expand upon these three core operators one by one. Spatiotemporal synchronization transformation The aim is to establish a unified spatiotemporal reference framework to eliminate spatiotemporal misalignments caused by sensor differences. Its processing objects are... Each data tuple in ,in This represents data content such as image matrices, spectrograms, and time-series vectors.

[0042] First, time synchronization is performed. This patent uses a high-precision Global Positioning System (GPS) time signal. As an absolute time reference, a time calibration function is defined for each data tuple. This function models the fixed offset and linear drift of the sensor clock by fitting historical synchronization point data using the least squares method. (Calibrated timestamp) The calculation is as follows:

[0043] ,

[0044] function The output ensures that for all Satisfying constraints This allows for time alignment of all modal data with sub-second precision, achieving time alignment of all modal data.

[0045] Secondly, spatial synchronization is performed. This involves synchronizing the original coordinates attached to all data. Transform to the universal WGS-84 geodetic coordinate system. This transformation is performed through a parameterized mapping function. The parameters are determined by the extrinsic calibration data of each sensor.

[0046] ,

[0047] Finally, based on unified time and space coordinates Regarding the original data content Spatial resampling or temporal interpolation is performed to fill or align to a standard spatiotemporal grid. This process is accomplished through resampling operators. Finish:

[0048] ,

[0049] in, For image data, bilinear interpolation may be used, while for time-series data, linear interpolation is used. Thus, this patent yields a spatiotemporally synchronized dataset:

[0050] ,

[0051] Dataset All samples in the dataset have consistent temporal and spatial labels, forming an aligned data cube that can be used for cross-modal association analysis.

[0052] Noise suppression transformation The goal is to improve The signal-to-noise ratio of the data content is improved, focusing on interferences such as cloud cover in optical / infrared images, specular reflection from the sea surface, and broadband ship noise in acoustic data. The core of this transformation is an adaptive filtering algorithm based on feature statistics.

[0053] To process synchronized infrared image blocks For example, the algorithm first uses a pre-designed convolutional kernel. Extract a noise-sensitive feature map from the image. The convolution kernel was optimized to enhance the local contrast difference between noise and signal regions:

[0054] ,

[0055] here, This represents a two-dimensional convolution operation. It is The polarization difference or Gaussian Laplace edge detection kernel.

[0056] Based on feature maps The algorithm at each pixel position Calculate its local neighborhood ( The statistical characteristics within the window are analyzed. A soft noise probability mask is generated by comparing the deviation of the center pixel feature value from the local mean. :

[0057] ,

[0058] in, and Neighborhood Mean and standard deviation of the internal eigenvalues; It is a very small positive number (e.g.) ), used to prevent division by zero errors; It is the Sigmoid function, which maps the standardized score to the (0,1) interval. The closer the value is to 1, the higher the probability that the pixel is noise.

[0059] Then, noise masking was used. Guided filtering process. In high-noise-probability regions indicated by the mask, Gaussian-smoothed image values ​​are used to replace the original values ​​to suppress noise; in low-noise-probability regions, original image details are preserved to protect the signal. This guided filtering operation can be described as follows:

[0060] ,

[0061] In the above formula, This represents element-wise multiplication (Hadamard product). Is with A matrix of all ones of the same size This represents the convolution operation. The standard deviation is A two-dimensional Gaussian smoothing kernel. For After performing this operation on all the images and spectral data in the dataset, this patent yields the denoised dataset:

[0062] ,

[0063] go through The transformation improved the overall signal-to-noise ratio of the data by an average of over 42%, providing a cleaner input for subsequent feature extraction.

[0064] Data augmentation transformation It is the final step in the preprocessing process, and its purpose is to improve the quality of preprocessing. The dataset is artificially expanded in size and diversity by applying a series of random transformations that conform to the physical laws of the marine environment to the samples. This is particularly helpful in mitigating the overfitting problem caused by insufficient samples of rare ecological targets during subsequent model training.

[0065] For the denoised acoustic spectrum This patent defines a time-domain stretching transformation. frequency domain translation transform This simulates the effects of changes in sound wave propagation speed, the Doppler effect of the sound source, or absorption characteristics at different depths. The enhanced acoustic spectrum... This is produced by the sequential action of these two transformations:

[0066] ,

[0067] Among them, time-domain stretching parameters Uniform random sampling within the interval [0.8, 1.2] enables compression or stretching of the spectral time axis; frequency domain shift parameters... Uniform random sampling within the interval [-2.0, +2.0] kHz is used to achieve vertical shifting of the entire spectrum. These two transformations work together to generate new and reasonable acoustic data variants without changing the essential category of the acoustic event.

[0068] For the denoised image patch This patent applies a series of spatial and photometric parameters. Transformation. First, perform random rotation. The angle is Internal sampling was performed to simulate different observation perspectives. Random cropping and scaling were then applied. Crop the original image The area was then resampled to a standard size (e.g.) (pixels) to simulate different viewing distances. Finally, random color dithering is performed. In the HSV color space, small, independent random perturbations (typically ranging from 0.5 to 100%) are applied to the three channels of hue (H), saturation (S), and lightness (V). This simulates changes in water turbidity under different lighting conditions. Enhanced image patch. It is produced by the composition of these transformations:

[0069] ,

[0070] Enhancement Transformation Applied to For each sample (including acoustic and image data), this patent yields the final, standardized preprocessed dataset. :

[0071] ,

[0072] In summary, through , and By sequentially executing and combining three operators, this patent successfully transforms raw, messy multimodal ocean data. It was systematically transformed into a standardized dataset that is spatiotemporally aligned, has a high signal-to-noise ratio, and is rich in diversity. This dataset serves as a reliable input for the entire technical framework, laying a solid data foundation for the subsequent efficient and accurate feature learning of the "distinguished ecological representation layer".

[0073] S2. Distinguishing ecological representation layer

[0074] The distinguishable ecological representation layer is the core feature learning module of this technical framework, which receives a standardized multimodal dataset from the preprocessing layer. This layer is responsible for learning and extracting robust feature representations that can clearly distinguish different ecological targets from complex marine backgrounds. The core design of this layer addresses the feature obfuscation problem prevalent in traditional methods, where the essential features of ecological targets are easily masked or interfered with by their surrounding variable environmental background (such as seawater textures under different lighting conditions, ocean current disturbances, and symbiotic biomes). To address this, this layer constructs a deep network consisting of three innovative cascaded sub-modules, aiming to perform a progressive processing of "feature decoupling - feature purification - feature fusion," ultimately outputting a high-dimensional, compact, and highly discriminative ecological feature vector. This provides semantically clear query vectors for subsequent target-aware matching.

[0075] The entire processing flow of the representation layer can be formally defined as a composition of three parameterized transformation operators. Let... This represents the Task Adaptive Feature Decoupling Network (TA-FDN). This represents the self-supervised masked real feature enhancement module (AIM-Marine). The hierarchical ecological feature fusion representation module represents the following: , , These are the learnable parameters for each module. The complete mapping from input data to final features is then:

[0076] ,

[0077] Next, this patent will elaborate on these three core transformations and clearly demonstrate the data transformation process. Initially, how does the flow gradually pass through each module and evolve into... .

[0078] The TA-FDN module aims to perform preliminary source separation on the input multimodal mixed features. Its core objective is to learn two independent feature subspaces: one specifically encodes the intrinsic properties of the ecological object itself (such as the morphology, texture, spectral or acoustic spectral patterns of a particular species), while the other focuses on encoding information about its surrounding environmental context. The input to this module is a preprocessed, standardized data sample. The output is the decoupled target feature tensor. and background feature tensor .

[0079] The network employs a parallel dual-branch encoder architecture. Each branch consists of a convolutional layer followed by instance normalization and a ReLU activation function, but the convolutional kernel weights of the two branches are initialized and optimized independently to guide them to focus on different patterns. Target branch and background branches For input Perform parallel processing:

[0080] ,

[0081] in, and These are the convolution weight parameters for the two branches. To dynamically adapt to different monitoring tasks, this patent introduces a lightweight task adaptive unit (TAU). This unit generates a pair of dynamic modulation vectors based on the metadata of the input data, such as the main modality type, acquisition time period, and geolocation coding. (C represents the number of feature channels), and channel-level weighting is applied to the initially extracted features:

[0082] ,

[0083] in, This represents a broadcast multiplication indicating the channel direction. Through this mechanism, the network can flexibly adjust the importance of each feature channel based on contextual information.

[0084] In order to and To truly achieve statistical decoupling, this patent defines a decoupling loss function based on minimizing mutual information. Mutual information The degree of dependence between two random variables was measured. The optimization objective of this patent is:

[0085] ,

[0086] Directly calculating mutual information is difficult. In practice, this patent employs a variational estimation method based on adversarial learning. An auxiliary discriminator network D is introduced, whose goal is to distinguish feature pairs. Is it a genuine decoupled pair from the same sample, or a fake pair randomly combined from different samples? By alternately optimizing the Feature Encoder-FDN (TA-FDN) to "deceive" the discriminator (making genuine pairs appear fake), this patent can effectively reduce... and The mutual information between them. After training, this module improved the distinguishability between target and background features by 65%, providing a preliminarily separated and purer feature stream for subsequent processing.

[0087] Although TA-FDN has decoupled features, each feature flow (especially) The internal structure may still contain "false" or redundant feature fragments that are irrelevant to the target category. The purpose of the AIM-Marine module is to further refine these features. Its core idea is to dynamically generate a binary mask for each training sample. This mask can act like a "spotlight," retaining only the feature regions that are crucial to the ecological target identification of the current sample, while obscuring other irrelevant parts.

[0088] First, the decoupled two-stream features are concatenated and input into a lightweight multi-scale feature encoder E (based on ResNet-50 compression) to construct a four-level feature pyramid:

[0089] ,

[0090] in, This indicates splicing along the channel dimension. This is the l-th stage of the encoder, outputting a feature map. The spatial dimensions decrease step by step ( , , , This allows you to capture information ranging from fine-grained details to the global context.

[0091] Next, for each level of the pyramid A lightweight mask generation network G_l generates a spatial binary mask for it. To enable gradient backpropagation in discrete mask decisions, this patent employs the Gumbel-Softmax reparameterization technique. Specifically, For each spatial location Output a two-dimensional logic value Then, the hard mask value (0 or 1) at that position is obtained through Gumbel-Softmax sampling:

[0092] ,

[0093] In the forward propagation of actual training, the above sampling is used to obtain the hard mask; in the backward propagation, the temperature parameter is used. The gradient is calculated using a continuous approximation of the controlled Gumbel-Softmax distribution.

[0094] The mask generation network G is not trained using additional annotations, but rather through a self-supervised "feature reconstruction" task. Its training objective is to perform element-wise multiplication of the feature pyramid using the generated mask, i.e. Afterwards, the remaining features after being obscured It should be able to predict the predefined pseudo-label of a sample (e.g., the ecological scene category or principal component obtained through clustering) with high accuracy using a simple multilayer perceptron (MLP) classification head. This forces G to learn to automatically identify and retain the feature regions that are most informative for sample discrimination. The resulting multi-scale feature set is purified through this step. Its Ecological Objectives (EPG) score reached 0.89.

[0095] This module is responsible for deep fusing and high-order encoding of refined features from different scales to generate a final unified ecological feature representation. This process simulates a cognitive hierarchy from local perception to global synthesis. First, feature maps from different scales are... Aggregated features are obtained by upsampling to the same medium size using bilinear interpolation and then concatenating them along the channel dimension. ,in,

[0096] To achieve deep fusion of features refined at different scales, the first step is to achieve unified alignment of spatial dimensions. This is necessary for multi-scale feature sets. Each feature map in Where l represents the scale level, with a value ranging from 1 to 4. We use bilinear interpolation for upsampling to unify its spatial dimensions to a pre-defined medium size. The upsampled feature map can be simply represented as:

[0097] ,

[0098] in Represents the bilinear interpolation upsampling function. This step, used to specify the target spatial size, ensures that feature maps of different scales are perfectly aligned in spatial dimensions.

[0099] The feature map after upsampling is denoted as Its shape is , Let be the number of channels in the l-th feature map. Based on this, we perform a channel-dimensional concatenation operation on all upsampled feature maps to generate the initial aggregated features. The calculation formula is as follows:

[0100] ,

[0101] in This represents the concatenation function along the channel dimension, which merges feature maps of different scales along the channel direction. After concatenation... The shape is This operation integrates discriminative information from different scales into a single feature tensor, fully preserving the original feature representations at each scale. To further enhance the task relevance of the aggregated features, we will... The input channel attention module performs dynamic weighting. This module first performs global average pooling on the features of each channel through a "squeezing" operation, compressing the spatial dimension information into a single-value vector.

[0102] Subsequently, a channel attention module (SE-Block) is applied to compute weights for each channel to highlight those feature channels that are more important to the current task:

[0103] ,

[0104] Wherein, GAP represents global average pooling. It is the ReLU activation function. , These are the weights of the fully connected layer. It is the Sigmoid function. It is the learned channel attention vector.

[0105] Next, the enhanced features Given a Transformer encoder consisting of L layers (usually L=6), let the self-attention mechanism of the l-th layer in the Transformer encoder be:

[0106] ,

[0107] in, , , They are obtained from the input features through linear projection. This refers to the dimension of the key vector. After L layers of encoding, this patent obtains the deeply encoded features. .

[0108] Finally, for Global average pooling is applied to compress it into a one-dimensional feature vector, which is then projected onto a predefined 512-dimensional space through a fully connected layer, and then processed... Normalization, outputting the final ecological feature vector:

[0109] ,

[0110] At this point, the distinguishable ecological representation layer has completed its entire task. It takes the raw, mixed multimodal data... Through deep transformation involving feature decoupling, purification, and fusion, it is transformed into a feature vector that possesses high discriminativeness, high robustness, and rich semantic information. This vector, as the core representation of the entire monitoring framework, lays a solid foundation for achieving accurate and dynamic target perception and matching in the next stage.

[0111] S3. Target-aware matching layer

[0112] The target-aware matching layer is the core of this framework's decision-making process, responsible for matching distinguishable ecological feature vectors. Accurately associate and locate the target with known ecological templates. The input to this layer is a feature vector. and their corresponding geographical locations The output is a series of high-confidence ecological target identification results. The entire processing flow transforms abstract features into concrete monitoring conclusions through three core steps: dynamic anchor box generation, hierarchical diffusion matching, and confidence fusion.

[0113] To accommodate the varying scales of marine targets, ranging from centimeters to kilometers, this module dynamically generates initial spatial hypotheses based on input features. (Preset) Basic anchor frame dimensions Input features Adjustment parameters for each base anchor box, including center point offset, are predicted using a lightweight regression network. and scale logarithmic shift :

[0114] ,

[0115] Combined with input position Dynamic anchor frame Calculated as:

[0116] ,

[0117] in:

[0118] ,

[0119] Finally, a dynamic anchor frame set is obtained. This serves as the initial spatial assumption for subsequent matching. The SD-Match network uses a dynamic set of anchor boxes. and characteristics For input, matching is performed in two stages. First, a coarse matching is performed for each dynamic anchor box. Extract its corresponding local feature vector And calculate its relationship with the template library. cosine similarity Simultaneously, to assess geometric fit, the normalized Wasserstein distance (NWD) is calculated. The anchor frame is then... and templates Standard frame They are modeled as two-dimensional Gaussian distributions. and Its mean Centered on the box, the covariance matrix The diagonal is proportional to the width and height of the frame. The square of the second-order Wasserstein distance between the two distributions. It can be analytically calculated, and NWD is defined as:

[0120] ,

[0121] in This is a normalization constant. The overall score for coarse matching is obtained by weighting the two:

[0122] ,

[0123] For each anchor frame The template with the highest score is selected as its coarse matching category. And record the score. Only when The anchor frame has entered the fine matching stage.

[0124] In the fine-matching stage, a hierarchical diffusion model is used to iteratively optimize the anchor frames that have passed the coarse screening. (The anchor frames are used as an example.) and its coarse matching category As initial conditions, the anchor box parameters are encoded as vectors. The diffusion model generates the optimized parameter vector through a reverse process. Specifically, at time step... Noise prediction network Based on current parameters Steps Local features and categories Predicted noise:

[0125] ,

[0126] By utilizing the predicted noise, a better parameter estimate is calculated through the inverse update rule of the diffusion model. :

[0127] ,

[0128] in , and These are hyperparameters of the diffusion model. It is standard Gaussian noise. After passing through... arrive Through iteration, the optimized parameter vector is obtained. The vector is decoded into a fine anchor box. ,Right now:

[0129] ,

[0130] The decoding operation converts the vector into the coordinates and size of the bounding box. Therefore, intermediate variables in the diffusion model... Eventually converges to And directly output as a fine anchor frame. .

[0131] After obtaining the fine anchor frame Then, features are re-extracted based on their location. And calculate its relationship with the template. cosine similarity as well as With standard frame of value The perfect match score is:

[0132] ,

[0133] To improve the reliability of the results, the confidence scores are integrated with those of data-driven and knowledge-driven approaches. The feature matching confidence score is obtained by normalizing the fine-match score:

[0134] ,

[0135] Ecological logic confidence is determined by querying the rule base. The library encodes the constraints between target types and environmental parameters (such as water temperature and chlorophyll). Environmental data vectors are extracted based on the position of the precise anchor frame. ,calculate:

[0136] ,

[0137] The final confidence level is the weighted sum of the two:

[0138] ,

[0139] Only when The result was then adopted. The target perception matching layer ultimately outputs a set of structured monitoring results:

[0140] ,

[0141] Each result includes the target type and the optimized geographic bounding box, i.e., the fine anchor box. And fusion confidence. This output This provides a direct and reliable input for subsequent monitoring, feedback, and application.

[0142] S4. Monitoring Result Output and Feedback Layer

[0143] This layer serves as the terminal of the entire monitoring process and undertakes two core tasks: first, to convert the high-confidence results generated by the target perception and matching layer into standardized monitoring reports; and second, to build a closed-loop feedback mechanism to use the low-quality matching samples generated in this monitoring to perform lightweight fine-tuning of the core model, thereby achieving online evolution of system performance.

[0144] This module receives the final result set from the target perception and matching layer. For each valid result, the system executes a standardized output process. First, the precise anchor box... The coordinates are mapped from the image space to The system uses a geographic coordinate system to generate a standard geographic location description; the positioning error in this process has been verified to be no greater than 5 meters through actual measurements. Subsequently, the system will classify the target type... The geolocation result and the corresponding final confidence level. The data is packaged into a structured monitoring record. All records are integrated in chronological order and pushed to the monitoring center in real time through a standard data interface. They are then automatically visualized and rendered on electronic nautical charts or remote sensing image base maps to form a real-time ecological monitoring situation map that includes target type, precise location, and confidence score.

[0145] To achieve adaptive optimization of the system, this module designs a feedback process based on incremental learning. Its core is to use low-confidence matching samples generated in this monitoring task to fine-tune key parameters in the distinguishable ecological representation layer and the target perception matching layer, without initiating a computationally expensive full model retraining.

[0146] Specifically, after each monitoring task is completed, the system automatically filters out all cases with a confidence level below a set threshold. The matched samples constitute an incremental dataset. This dataset contains the original multimodal data of these samples. and its corresponding initial feature vector extracted from the distinguishable ecological representation layer. The system caches these "hard samples" and their features. The goal of model optimization is to fine-tune the network parameters. It can better handle such samples. This is achieved by minimizing a composite loss function that combines representation learning and matching tasks:

[0147] ,

[0148] in, This refers to loss functions designed for feature representation, such as contrastive loss, which aim to bring positive sample features of the same target closer together and push negative sample features of different targets further apart, thereby improving feature quality. Distinctiveness; It is a loss designed for matching tasks, such as focus loss based on optimized matching results, which aims to make the model pay more attention to the correct classification and localization of these difficult samples; These are the weighting coefficients that balance the two losses. This is achieved through small-scale incremental datasets. By performing a few rounds of gradient descent optimization, model parameters can be updated efficiently. This process is executed asynchronously in the background, and the corresponding parameters of the online model are replaced with hot updates upon completion. This enables the entire system to continuously learn and improve from actual error monitoring, significantly enhancing its generalization and robustness in the face of new environments or new disturbance patterns.

[0149] Experimental verification

[0150] To verify the effectiveness, robustness, and real-time performance of the marine ecological monitoring method based on ecological representation and target perception matching proposed in this invention in practical applications, this embodiment constructs a comprehensive experimental environment containing multi-source heterogeneous data for detailed testing. The dataset used in the experiment is named MarineEco-Fusion, which consists of 15,000 high-resolution optical images collected by an underwater robot in different sea areas such as deep sea, shallow waters, and coral reefs, as well as 15,000 frames of multibeam sonar and side-scan sonar images acquired simultaneously. To comprehensively evaluate the model performance, the dataset covers five main ecological targets: corals, fish schools, seagrass, marine debris, and seabed sediments. Based on water visibility, the data is divided into a high-resolution group with visibility greater than 5 meters and a turbid, low-light group with visibility less than 2 meters, to focus on testing the model's performance under harsh conditions. The experiment was run on a high-performance computing platform equipped with two NVIDIA A100 graphics cards and an Intel Xeon Gold processor. It was implemented based on the PyTorch 2.0 deep learning framework. The input images were normalized in size and trained for 200 rounds using the AdamW optimizer.

[0151] In the quantitative analysis phase, the experiment selected mean precision (mAP@0.5), recall, and F1 score as the main evaluation indicators, and introduced frames per second (FPS) to measure the real-time performance of the monitoring. The method of this invention was rigorously compared with traditional image processing methods, mainstream single-modal deep learning models such as YOLOv8, purely acoustic detection methods, and existing multimodal early fusion methods. The comprehensive performance comparison data shown in Table 1 below demonstrates that, under a unified test set, the method proposed in this invention exhibits significant advantages. Specifically, as... Figure 5As shown, the proposed method achieves an mAP of 89.5%, which is 10.9 percentage points higher than the 78.6% achieved by the YOLOv8 model using only optical data, and 8.3 percentage points higher than the 81.2% achieved by existing simple fusion methods. Particularly noteworthy is the recall rate, which reaches 87.8%, indicating a significant reduction in missed detections. Although the inference speed of 72 FPS is slightly lower than the 85 FPS of the pure vision model, it still far exceeds the 30 FPS standard typically required for real-time monitoring, demonstrating that the proposed method successfully achieves real-time performance suitable for engineering applications while maintaining high accuracy.

[0152] Table 1. Performance comparison of different methods on the MarineEco-Fusion dataset.

[0153] method Input mode mAP@0.5% Recall percentage F1 score Inference speed (FPS) Baseline Method A, i.e., traditional image processing Optical only 54.2 48.5 0.51 120 Benchmark Method B, i.e., YOLOv8 Optical only 78.6 75.2 0.77 85 Baseline Method C, i.e., sonar detection Acoustics only 65.4 60.1 0.62 90 Existing fusion methods, i.e., early fusion Optics plus acoustics 81.2 79.5 0.8 45 Method of the present invention Optics plus acoustics 89.5 87.8 0.88 72

[0154] Example 2

[0155] This embodiment provides a marine ecological monitoring system based on ecological representation and target perception matching, including:

[0156] The data acquisition module is configured to acquire multi-source ocean monitoring data;

[0157] The preprocessing module is configured to perform multimodal data preprocessing based on the acquired multi-source ocean monitoring data;

[0158] The feature module is configured to generate distinguishable ecological feature vectors based on the preprocessed data.

[0159] The confidence module is configured to generate fusion confidence scores by using distinguishable ecological feature vectors for target perception.

[0160] The monitoring module is configured to generate and provide feedback monitoring reports based on the fusion confidence results.

[0161] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the marine ecological monitoring method based on ecological representation and target perception matching.

[0162] A terminal device includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a marine ecological monitoring method based on ecological representation and target perception matching.

[0163] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A marine ecological monitoring method based on ecological representation and target perception matching, characterized in that, include: Acquire multi-source ocean monitoring data; Multimodal data preprocessing was performed based on the acquired multi-source marine monitoring data; Generate distinguishable ecological feature vectors based on the preprocessed data; Target perception is generated and fusion confidence is achieved using distinguishable ecological feature vectors; A monitoring report is generated and fed back based on the fusion confidence results; The process of generating distinguishable ecological feature vectors based on preprocessed data includes receiving a standardized multimodal dataset from the preprocessing layer. The Task Adaptive Feature Decoupling Network (TA-FDN) is used to generate the decoupled target feature tensor. and background feature tensor The TA-FDN network employs a parallel dual-branch encoder architecture, where each branch consists of a convolutional layer followed by normalization and ReLU activation functions. The target branch is utilized... and background branches For input Perform parallel processing: in, and These are the convolution weight parameters for the two branches; to dynamically adapt to different monitoring tasks, a lightweight task adaptive unit (TAU) is introduced, which generates a pair of dynamic modulation vectors based on the metadata of the input data. C represents the number of feature channels, and the initially extracted features are weighted at the channel level. ,in, Broadcast multiplication indicating channel direction; to make and To achieve decoupling, a decoupling loss function based on minimizing mutual information is defined. Utilizing mutual information The optimization objective is to measure the degree of dependence between two random variables. Finally, a variational estimation method based on adversarial learning is adopted, and an auxiliary discriminator network D is introduced to distinguish feature pairs. ; The process of generating distinguishable ecological feature vectors based on preprocessed data also includes using the self-supervised masked real feature enhancement network AIM-Marine to purify the decoupled features. First, the decoupled dual-stream features are concatenated and input into a lightweight multi-scale feature encoder E to construct a four-level feature pyramid. in, This indicates splicing along the channel dimension. This is the l-th stage of the encoder, outputting a feature map. The spatial dimensions decrease progressively, thereby capturing information from fine-grained details to global context; then for each level of the pyramid... A lightweight mask generation network G_l generates a spatial binary mask for it. To enable gradient backpropagation in discrete mask decisions, the Gumbel-Softmax reparameterization technique is employed. Specifically, For each spatial location Output a two-dimensional logic value Then, the hard mask value at that location is obtained through Gumbel-Softmax sampling: The mask generation network G is trained through a self-supervised feature reconstruction task. The training objective is to perform element-wise multiplication of the feature pyramid using the generated mask, i.e. Afterwards, the remaining features after being obscured It should be able to predict predefined pseudo-labels for samples using a multilayer perceptron (MLP) classification head, forcing G to automatically identify and retain the feature regions with the most information for sample discrimination, resulting in a refined multi-scale feature set. Its ecological target orientation score (EPG) reached 0.89; The process of generating distinguishable ecological feature vectors based on preprocessed data also includes using a hierarchical ecological feature fusion representation network to deeply fuse and encode purified features from different scales to generate a final unified ecological feature representation. First, feature maps from different scales are... Aggregated features are obtained by upsampling to the same medium size using bilinear interpolation and then concatenating them along the channel dimension. Subsequently, the channel attention module SE-Block is applied to calculate weights for each channel to highlight those feature channels that are more important to the current task: Wherein, GAP represents global average pooling. It is the ReLU activation function. , These are the weights of the fully connected layer. It is the Sigmoid function. These are the learned channel attention vectors; then the enhanced features are... The input is a Transformer encoder consisting of L layers to model long-range dependencies within features and to fuse multimodal information. The self-attention mechanism of the l-th layer in the Transformer encoder is denoted as: in, , , They are obtained from the input features through linear projection. The dimension of the key vector, after being encoded through L layers, yields the deeply encoded features. Finally, regarding Global average pooling is applied to compress it into a one-dimensional feature vector, and then it is projected onto a predefined 512-dimensional space through a fully connected layer, and then processed... Normalization, outputting the final ecological feature vector: This enables the processing of raw, mixed multimodal data. Through deep transformation involving feature decoupling, purification, and fusion, it is transformed into feature vectors that possess high discriminativeness, high robustness, and rich semantic information. .

2. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 1, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data includes utilizing... The spatiotemporal synchronization transformation operator performs time synchronization on the original multimodal dataset. It contains four types of sensor data tuples: optical remote sensing images Infrared remote sensing images Underwater acoustic spectrum diagram and buoy sensing timing data First, for each data tuple, define a time calibration function. By fitting historical synchronization point data using the least squares method, the fixed offset and linear drift of the sensor clock are modeled, and the calibrated timestamps are obtained. The calculation is as follows: ,function The output ensures that for all Satisfying constraints First, time alignment of all modal data is achieved with sub-second precision; second, spatial synchronization is performed, using the original coordinates attached to all data. Transform to the universal WGS-84 geodetic coordinate system using a parameterized mapping function. The parameters are determined by the extrinsic calibration data of each sensor. Finally, based on unified time and space coordinates Regarding the original data content Spatial resampling and temporal interpolation are performed to fill or align to a standard spatiotemporal grid, using resampling operators. Represented as: in, Bilinear interpolation was used for image data, and linear interpolation was used for time-series data, resulting in a spatiotemporally synchronized dataset. .

3. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 2, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data also includes using convolution kernels. Extract noise-sensitive feature maps from images. , convolution kernel Optimization is performed to enhance the local contrast difference between noise and signal regions: ,in This represents a two-dimensional convolution operation. It is Polarization difference edge detection kernel; based on feature map At each pixel position Calculate its local neighborhood The statistical properties within the area are used to generate a soft noise probability mask by comparing the deviation of the central pixel feature value from the local mean. : in, and Neighborhood Mean and standard deviation of the internal eigenvalues; It is a very small positive number used to prevent division by zero errors; It is the Sigmoid function, which maps the standardized scores to the (0,1) interval; then it uses a noise mask. The guided filtering process, in the high-noise-probability region indicated by the mask, replaces the original value with a Gaussian-smoothed image value to suppress noise; in the low-noise-probability region, it preserves the original image details to protect the signal, as described below: in, This represents element-wise multiplication. Is with A matrix of all ones of the same size This represents the convolution operation. The standard deviation is Two-dimensional Gaussian smoothing kernel; for After performing operations on all images and spectral data, the denoised dataset is obtained: 。 4. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 3, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data also includes data augmentation transformation, including the denoised acoustic spectrogram. Define time-domain stretching transformation frequency domain translation transform To simulate the effects of changes in sound wave propagation speed and the Doppler effect of the sound source, the enhanced acoustic spectrum was obtained. Produced by the action of changing the order: Among them, time-domain stretching parameters Uniform random sampling is performed within the interval [0.8, 1.2]; frequency domain shift parameter Uniform random sampling is performed within the interval [-2.0, +2.0] kHz; for the denoised image patch... Apply space and light intensity Transformation, through random rotation The angle is Internal sampling was performed to simulate different observation perspectives, followed by random cropping and scaling. Crop the original image The area was then resampled to a standard size, and finally random color dithering was performed. Enhanced image patches The expression generated by transformation composition is as follows: This will enhance the transformation. Applied to The samples in the dataset are used to obtain the final standardized preprocessed dataset. : 。 5. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 4, characterized in that, The method of generating fusion confidence scores using distinguishable ecological feature vectors for target perception includes using a target perception matching layer to combine the distinguishable ecological feature vectors... Accurately associate and locate the target with known ecological templates, with the input being a feature vector. and their corresponding geographical locations The output is a series of high-confidence ecological target identification results, achieved through three steps: dynamic anchor box generation, hierarchical diffusion matching, and confidence fusion. Dynamic anchor box generation includes dynamically generating initial spatial hypotheses based on input features and pre-setting... Basic anchor frame dimensions Input features Adjustment parameters for each base anchor frame, including center point offset, are predicted using a lightweight regression network. and scale logarithmic shift : Combined with input position Dynamic anchor frame Calculated as: , in: Finally, a dynamic anchor frame set is obtained. .

6. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 5, characterized in that, The hierarchical diffusion matching includes utilizing the SD-Match network with a dynamic set of anchor boxes. and characteristics For input, matching is performed in two stages. First, a coarse matching is performed for each dynamic anchor box. Extract its corresponding local feature vector And calculate its relationship with the template library. cosine similarity Meanwhile, to assess geometric fit, the normalized Wasserstein distance (NWD) is calculated, and the anchor frame is... and templates Standard frame They are modeled as two-dimensional Gaussian distributions. and Its mean Centered on the box, the covariance matrix The diagonal is proportional to the width and height of the frame, and the square of the second-order Wasserstein distance between the two distributions is given. Analytical calculation, NWD is defined as: in As a normalization constant, the coarse matching comprehensive score is obtained by weighting the two: Finally, for each anchor frame The template with the highest score is selected as its coarse matching category. And record the score. Only when The anchor frames then enter the fine matching stage; in the fine matching stage, a hierarchical diffusion model is used to iteratively optimize the anchor frames that have passed the coarse screening, with the anchor frames... and its coarse matching category As initial conditions, the anchor box parameters are encoded as vectors. The diffusion model generates an optimized parameter vector through a reverse process. Specifically, in time step Noise prediction network Based on current parameters Steps Local features and categories Predicted noise: Using the predicted noise, a better parameter estimate is calculated through the inverse update rule of the diffusion model. : in , and These are hyperparameters of the diffusion model. It is standard Gaussian noise, after being processed from... arrive Through iteration, the optimized parameter vector is obtained. The vector is decoded into a fine anchor box. ,Right now: The decoding operation converts the vector into the coordinates and size of the bounding box; therefore, the intermediate variables in the diffusion model... Eventually converges to And directly output as a fine anchor frame. ; Obtain the fine anchor frame Then, features are re-extracted based on their location. And calculate its relationship with the template. cosine similarity as well as With standard frame of value The perfect match score is: .

7. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 6, characterized in that, The confidence fusion includes fusing data-driven and knowledge-driven confidence scores, where the feature matching confidence score is obtained by normalizing the fine-matching score. Ecological logic confidence is determined by querying the rule base. Obtained, based on the fine anchor frame Location Extraction of Environmental Data Vectors ,calculate: The final confidence level is the weighted sum of the two: Only when When this result is adopted, the target perception matching layer finally outputs a set of structured monitoring results: Each result includes the target type and the optimized geographic bounding box, i.e., the fine anchor box. And fusion confidence.