Marine ecological monitoring method based on ecological representation and target perception matching
The marine ecological monitoring method, which utilizes multi-module collaborative optimization, solves the problems of data synchronization, feature extraction, and target matching in marine ecological monitoring, achieving high-precision, high-efficiency, and low-energy-consumption marine ecological monitoring. It is applicable to marine ecological monitoring systems and equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-13
- Publication Date
- 2026-04-10
AI Technical Summary
Existing marine ecological monitoring technologies suffer from problems such as low synchronization accuracy, weak anti-interference ability, blurred distinction between target and background in ecological feature extraction, poor multi-scale adaptation in target matching and model iteration, weak generalization ability, and low update efficiency in the multimodal data preprocessing stage.
A multi-module collaborative optimization approach is adopted, including multimodal data preprocessing, distinguishable ecological feature representation, and target perception matching. Through spatiotemporal synchronous transformation, an improved polarization weight enhancement algorithm, a task-adaptive feature decoupling network, a dynamic anchor frame and hierarchical diffusion matching system, combined with lightweight design and knowledge distillation technology, high-precision and high-efficiency marine ecological monitoring is achieved.
It significantly improved the accuracy and efficiency of marine ecological monitoring, reducing the data synchronization error from more than 5 seconds to within ±0.1 seconds, increasing the signal-to-noise ratio by 42%, improving the target recognition capability by 65%, reducing the small target false detection rate to 12%, stabilizing the accuracy of cross-sea monitoring, improving the model generalization capability by 40%, and reducing energy consumption by 60%, achieving a balance between high accuracy, high real-time performance, and low energy consumption.
Smart Images

Figure CN121834243A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of marine ecological monitoring, and in particular to a marine ecological monitoring method based on ecological representation and target perception matching. BACKGROUND
[0002] The marine ecosystem is the core carrier of fishery resource supply and climate regulation. Traditional monitoring mainly relies on manual sampling and single equipment, which is inefficient and inaccurate. Manual sampling relies on fixed-point operation of research vessels, and only covers 60 to 90 square kilometers of sea area per voyage. The response to sudden red tides lags more than 48 hours, and the missing detection rate of micron-level plankton is as high as 58%. Single sensor monitoring is also obviously limited: hydrological buoys can only collect basic parameters such as temperature and salinity, and cannot directly identify biological targets; underwater cameras are limited by water transparency, with an effective monitoring distance generally less than 5 meters, and are completely ineffective in turbid waters; optical sensors have a false positive rate of 37% for red tides, often misjudging ordinary algae aggregation as harmful red tides. Traditional technology is costly, with a monitoring cost of about 800 yuan per square kilometer, which is 9 times that of intelligent monitoring solutions, making it difficult to achieve large-scale and normalized coverage.
[0003] However, the existing marine ecological monitoring technology has the defects of low synchronization accuracy and weak anti-interference ability in the preprocessing of multi-modal data. The traditional method integrates satellite remote sensing, underwater acoustic and buoy sensor data, only relying on simple time stamp alignment, with a time and space synchronization error often exceeding 5 seconds, and the spatial coordinates are not unified to the WGS-84 coordinate system, resulting in invalid data correlation; for cloud cover, ship noise and other interference in the marine environment, the ordinary filtering algorithm can only improve the signal-to-noise ratio by about 15%, which cannot meet the needs of subsequent feature extraction.
[0004] The existing technology faces the core problems of blurred target and background differentiation and low utilization rate of real features in ecological feature extraction. The conventional feature extraction model does not design a separation mechanism for marine scenes, mixing ecological target features with background features such as sea current texture and temperature gradient, with a target-background feature differentiation degree of less than 50%; even if the mask technology is introduced, due to the lack of adaptive ability, key features such as red tide spectrum and fish acoustic pulse are lost, and the utilization rate of real features is only about 60%.
[0005] The existing technology has the shortcomings of poor multi-scale adaptation, weak generalization ability and low update efficiency in target matching and model iteration. The traditional matching algorithm uses a fixed anchor frame design, which cannot cover multi-scale targets from 16x16 pixel micron-level algae to 256x256 pixel meter-level red tides, with a small target missing detection rate of more than 45%; when applied across sea areas, the recognition accuracy decreases by more than 50% due to the lack of dynamic update mechanism, and the model parameter update requires full retraining, taking more than 15 days. SUMMARY
[0006] In order to solve the above-mentioned problems, the present application provides a marine ecological monitoring method based on ecological representation and target perception matching.
[0007] In a first aspect, the present application provides a marine ecological monitoring method based on ecological representation and target perception matching, which adopts the following technical solution: A marine ecological monitoring method based on ecological representation and target perception matching comprises: Obtaining multi-source marine monitoring data; Performing multi-modal data preprocessing based on the obtained multi-source marine monitoring data; Generating distinguishable ecological feature vectors according to the preprocessed data; Generating fusion confidence using distinguishable ecological feature vectors for target perception; Generating a monitoring report based on the fusion confidence result and feeding back.
[0008] In a second aspect, a marine ecological monitoring system based on ecological representation and target perception matching comprises: A data acquisition module configured to obtain multi-source marine monitoring data; A preprocessing module configured to perform multi-modal data preprocessing based on the obtained multi-source marine monitoring data; A feature module configured to generate distinguishable ecological feature vectors according to the preprocessed data; A confidence module configured to generate fusion confidence using distinguishable ecological feature vectors for target perception; A monitoring module configured to generate a monitoring report based on the fusion confidence result and feed back.
[0009] In a third aspect, the present application provides a computer-readable storage medium having a plurality of instructions stored therein, the instructions being adapted to be loaded and executed by a processor of a terminal device to implement the marine ecological monitoring method based on ecological representation and target perception matching.
[0010] In a fourth aspect, the present application provides a terminal device comprising a processor and a computer-readable storage medium, the processor being configured to implement instructions, and the computer-readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor to implement the marine ecological monitoring method based on ecological representation and target perception matching.
[0011] In summary, the present application has the following beneficial technical effects: The present application realizes multiple breakthroughs in the precision, efficiency and practicality of marine ecological monitoring through multi-module collaborative optimization, which is significantly superior to traditional technologies. In the aspect of multi-modal data processing, the spatio-temporal stamp dynamic calibration technology compresses the data synchronization error from more than 5 seconds to within ±0.1 seconds, combined with WGS-84 coordinate system processing, completely solves the problem of multi-source data correlation failure; the improved polarization weight enhancement algorithm filters out interference such as clouds and ships, which makes the signal-to-noise ratio increase by more than 42%, which is a qualitative leap compared with the 15% improvement rate of ordinary filtering algorithm, and provides a high-reliability data foundation for subsequent feature extraction, and the effective utilization rate of data is increased from less than 70% to more than 95%.
[0012] The optimization of the feature extraction link brings a core improvement of the target recognition ability. The task-adaptive feature decoupling network realizes the accurate separation of target and background features through double-channel convolution, and cooperates with the self-supervised mask mechanism adapted to the marine scene, so that the target-background feature discrimination degree is increased from less than 50% to more than 65%, the real feature utilization rate is greatly improved compared with the 60% of traditional technologies, and the EPG directivity score reaches 0.89, which means that the capture accuracy of key features such as red tide spectrum and fish acoustic pulse is significantly enhanced. This high-purity feature output reduces the misrecognition rate of red tide and ordinary algae from 35% to less than 8%, providing high-quality input for subsequent target matching and improving the credibility of monitoring results from the source.
[0013] The optimization of target matching and model deployment completely solves the practicality bottleneck of traditional technologies. The application of dynamic anchor frame and hierarchical diffusion matching system reduces the small target detection rate below 10 cm from more than 45% to 12%, the cross-sea monitoring accuracy is reduced by less than 15%, and the matching error is stable within 2 pixels; the automatic warehousing and incremental updating mechanism improves the model generalization ability by 40%, and when facing unknown invasive species, it does not need manual intervention, and the updating cycle is shortened from 15 days to 2 days. At the same time, the combination of lightweight design and knowledge distillation technology controls the model parameter quantity within 15M while ensuring the single-frame processing rate to be more than 17 FPS, the energy consumption is less than or equal to 5W, which perfectly adapts to embedded devices such as buoys, reduces the deployment cost of traditional models by 60%, realizes the balance of "high precision-high real-time-low energy consumption", and provides a feasible solution for marine ecological normalization and wide coverage monitoring, and helps the efficient development of red tide early warning and biodiversity protection. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is a schematic diagram of a marine ecological monitoring method based on ecological representation and target perception matching according to an embodiment of the present application; Figure 2 is a schematic diagram of step S1 according to an embodiment of the present application; Figure 3 is a schematic diagram of step S2 according to an embodiment of the present application; Figure 4 is a schematic diagram of step S3 of embodiment 1 of the invention.
[0015] Figure 5 is a comparison chart of the predicted distribution verified by experiments of embodiment 1 of the invention. DETAILED DESCRIPTION
[0016] The invention is further described in detail below with reference to the accompanying drawings.
[0017] Embodiment 1 With reference to Figure 1 , the marine ecological monitoring method based on ecological representation and target perception matching of the embodiment comprises: S1. Multi-modal data preprocessing The multi-modal data preprocessing layer constitutes the foundation of the intelligent marine ecological monitoring system, and its core task is to transform the original, heterogeneous, noisy multi-source marine monitoring data into a standardized, high-quality, unified data representation that can be directly used for deep feature learning. This processing procedure follows a strict order of "spatiotemporal synchronization → noise suppression → data augmentation", aiming to systematically address the three major challenges of data alignment, signal-to-noise ratio improvement, and insufficient sample diversity. This section will elaborate on the complete, coherent mathematical transformation process and data flow from the original multi-modal data set to the final preprocessed data set .
[0018] Original multi-modal data set contains data tuples from four main sensors: optical remote sensing images , infrared remote sensing images , underwater acoustic spectrograms , and buoy sensor time series data . Each data tuple is accompanied by a timestamp and spatial coordinates recorded by the sensor at the time of collection. The unprocessed data exhibits significant inconsistencies: the different clocks of each sensor result in a deviation in the timestamp t; the different coordinate systems make it impossible to directly fuse the spatial coordinates ; the data is contaminated with clouds, ship noise, and other interference; and the number of samples for rare ecological targets is extremely limited. The goal of preprocessing is to construct a deterministic transformation that maps to a clean, aligned, and rich standard data set .
[0019] The entire preprocessing procedure can be considered as a composition of three sequentially executed sub-transformations. Let be the spatiotemporal synchronization transformation operator, be the noise suppression transformation operator, For data augmentation transform operators, the complete preprocessing transform... Defined as: , in This indicates function composition. Therefore, the final output data is: , Next, this patent will define and expand upon these three core operators one by one. Spatiotemporal synchronization transformation The aim is to establish a unified spatiotemporal reference framework to eliminate spatiotemporal misalignments caused by sensor differences. Its processing objects are... Each data tuple in ,in This represents data content such as image matrices, spectrograms, and time-series vectors.
[0020] First, time synchronization is performed. This patent uses a high-precision Global Positioning System (GPS) time signal. As an absolute time reference, a time calibration function is defined for each data tuple. This function models the fixed offset and linear drift of the sensor clock by fitting historical synchronization point data using the least squares method. (Calibrated timestamp) The calculation is as follows: , function The output ensures that for all Satisfying constraints This allows for time alignment of all modal data with sub-second precision, achieving time alignment of all modal data.
[0021] Secondly, spatial synchronization is performed. This involves synchronizing the original coordinates attached to all data. Transform to the universal WGS-84 geodetic coordinate system. This transformation is performed through a parameterized mapping function. The parameters are determined by the extrinsic calibration data of each sensor. , Finally, based on unified time and space coordinates Regarding the original data content Spatial resampling or temporal interpolation is performed to fill or align to a standard spatiotemporal grid. This process is accomplished through resampling operators. Finish: , in, For image data, bilinear interpolation may be used, while for time-series data, linear interpolation is used. Thus, this patent yields a spatiotemporally synchronized dataset: , Dataset All samples in the dataset have consistent time and spatial labels, which form an aligned data cube for cross-modal correlation analysis.
[0022] Noise suppression transform The goal of the transform is to improve the signal-to-noise ratio of data content in , focusing on cloud cover, sea surface specular reflection in optical / infrared images, and broadband ship noise in acoustic data. The core of the transform is an adaptive filtering algorithm based on feature statistics.
[0023] Take the processed infrared image block as an example. First, the algorithm uses a pre-designed convolution kernel to extract a noise-sensitive feature map from the image. The convolution kernel is optimized to enhance the local contrast difference between noise and signal regions: , Here, denotes a two-dimensional convolution operation, is a polarization difference or Gaussian Laplacian edge detection kernel of .
[0024] Based on the feature map , the algorithm calculates the statistical properties of the local neighborhood ( window) of each pixel position . By comparing the deviation of the central pixel feature value from the local mean, a soft noise probability mask is generated: , where and are the mean and standard deviation of the feature values in the neighborhood ; is a very small positive number (e.g. ) to prevent division by zero errors; is a Sigmoid function that maps the normalized score to the interval (0,1), with a value closer to 1 indicating a higher probability of the pixel being noise.
[0025] Subsequently, the noise mask is used to guide the filtering process. In the high-noise probability regions indicated by the mask, the original values are replaced with Gaussian smoothed image values to suppress noise; in the low-noise probability regions, the original image details are preserved to protect the signal. This guided filtering operation can be expressed as: , In the above formula, represents element-wise multiplication (Hadamard product), is an all-one matrix of the same size as , represents convolution operation, is a two-dimensional Gaussian smoothing kernel with a standard deviation of . After performing this operation on all image and spectrum data in , the patent obtains a denoised data set: , After transformation, the overall signal-to-noise ratio of the data is improved by an average of more than 42%, providing purer input for subsequent feature extraction.
[0026] Data augmentation transformation is the last step of the preprocessing process, which aims to artificially expand the size and diversity of the data set by applying a series of random transformations consistent with the physical laws of the marine environment to the samples in . This is especially helpful in alleviating the overfitting problem caused by the lack of rare ecological target samples in subsequent model training.
[0027] For the denoised acoustic spectrum , the patent defines a time-domain stretching transformation and a frequency-domain translation transformation to simulate the effects of changes in sound propagation speed, Doppler effect of sound sources, or different depth absorption characteristics. The enhanced acoustic spectrum is generated by the sequential action of these two transformations: , where the time-domain stretching parameter is uniformly randomly sampled in the interval [0.8, 1.2] to compress or stretch the time axis of the spectrum; the frequency-domain translation parameter is uniformly randomly sampled in the interval [-2.0, +2.0] kHz to translate the entire spectrum up and down. These two transformations work together to generate new, reasonable acoustic data variants without changing the essential categories of acoustic events.
[0028] For the denoised image block , the patent applies a series of spatial and photometric transformations. First, random rotation is performed with an angle sampled in to simulate different observation angles. Then, random cropping and scaling is performed to crop a region of the original image and resample it to a standard size (such as (pixels) to simulate different viewing distances. Finally, random color dithering is performed. In the HSV color space, small, independent random perturbations (typically ranging from 0.5 to 100%) are applied to the three channels of hue (H), saturation (S), and lightness (V). This simulates changes in water turbidity under different lighting conditions. Enhanced image patch. It is produced by the composition of these transformations: , Enhancement Transformation Applied to For each sample (including acoustic and image data), this patent yields the final, standardized preprocessed dataset. : , In summary, through , and By sequentially executing and combining three operators, this patent successfully transforms raw, messy multimodal ocean data. It was systematically transformed into a standardized dataset that is spatiotemporally aligned, has a high signal-to-noise ratio, and is rich in diversity. This dataset serves as a reliable input for the entire technical framework, laying a solid data foundation for the subsequent efficient and accurate feature learning of the "distinguished ecological representation layer".
[0029] S2. Distinguishing ecological representation layer The distinguishable ecological representation layer is the core feature learning module of this technical framework, which receives a standardized multimodal dataset from the preprocessing layer. This layer is responsible for learning and extracting robust feature representations that can clearly distinguish different ecological targets from complex marine backgrounds. The core design of this layer addresses the feature obfuscation problem prevalent in traditional methods, where the essential features of ecological targets are easily masked or interfered with by their surrounding variable environmental background (such as seawater textures under different lighting conditions, ocean current disturbances, and symbiotic biomes). To address this, this layer constructs a deep network consisting of three innovative cascaded sub-modules, aiming to perform a progressive processing of "feature decoupling - feature purification - feature fusion," ultimately outputting a high-dimensional, compact, and highly discriminative ecological feature vector. This provides semantically clear query vectors for subsequent target-aware matching.
[0030] The entire processing flow of the representation layer can be formally defined as a composition of three parameterized transformation operators. Let... This represents the Task Adaptive Feature Decoupling Network (TA-FDN). This represents the self-supervised masked real feature enhancement module (AIM-Marine). representation layer, where , , are the learnable parameters of each module, respectively. Then the complete mapping from input data to final features is: , Next, this patent will elaborate on these three core transformations in detail and clearly show how data flows through each module and evolves from to .
[0031] The TA-FDN module aims to perform a preliminary source separation on the input multi-modal mixed features. Its core objective is to learn two independent feature subspaces: one that encodes the intrinsic attributes of the ecological object itself (e.g., the morphological, textural, spectral, or acoustic spectral patterns of a specific species), and another that focuses on encoding the surrounding environmental context information. The input to this module is the pre-processed standardized data samples , and the output is the decoupled object feature tensor and context feature tensor .
[0032] The network adopts a parallel dual-branch encoder architecture. Each branch consists of a convolutional layer followed by instance normalization and ReLU activation functions, but the convolution kernel weights of the two branches are independently initialized and optimized to guide them to focus on different patterns. The object branch and context branch process the input in parallel: , where and are the convolutional weight parameters of the two branches, respectively. To dynamically adapt to different monitoring tasks, this patent introduces a lightweight task-adaptive unit (TAU). This unit takes the meta-information of the input data, such as the main modality type, acquisition period, geographical location encoding, etc., as conditions to generate a pair of dynamic modulation vectors (C is the number of feature channels), which performs channel-level weighting on the preliminary extracted features: , where denotes the broadcast multiplication in the channel direction. Through this mechanism, the network can flexibly adjust the importance of each feature channel according to the context information.
[0033] To truly achieve statistical decoupling of and , this patent defines a decoupling loss function based on mutual information minimization Mutual information Measures the degree of dependence between two random variables. The optimization goal of this patent is: , Directly calculating mutual information is difficult. In practice, this patent adopts a variational estimation method based on adversarial learning. A auxiliary discriminator network D is introduced, whose goal is to distinguish between feature pairs are real decoupled pairs from the same sample, or false pairs randomly composed from different samples. By alternately optimizing the feature encoder (TA-FDN) to "deceive" the discriminator (make the real pairs look like false pairs), this patent can effectively reduce Mutual information between After training, the distinguishability between target and background features is improved by 65%, providing a preliminary separation and higher purity feature stream for subsequent processing.
[0034] Although TA-FDN has decoupled the features, each feature stream (especially ) may still contain "false" or redundant feature fragments that are irrelevant to target class discrimination. The purpose of the AIM-Marine module is to purify these features again. The core idea is to dynamically generate a binary mask for each training sample, which can act like a "spotlight" to only retain the feature area essential for the ecological target recognition of the current sample, while masking other irrelevant parts.
[0035] First, the decoupled double-stream features are spliced and input into a lightweight multi-scale feature encoder E (based on ResNet-50 compression) to build a four-level feature pyramid: , where represents splicing along the channel dimension, is the l-th level of the encoder, and the spatial size of the output feature map decreases step by step ( , , , ), thus capturing information from fine-grained details to global context.
[0036] Next, for each layer of the pyramid, a lightweight mask generation network G_l generates a spatial binary mask for it. To enable gradient backpropagation in discrete mask decision, this patent uses the Gumbel-Softmax reparameterization trick. Specifically, for each spatial position Output a two-dimensional logical value Then get the hard mask value (0 or 1) of this position by Gumbel-Softmax sampling: , In the actual training of the forward propagation, the above sampling is used to obtain the hard mask; in the back propagation, the temperature parameter The continuous approximation of the Gumbel-Softmax distribution controlled is used to calculate the gradient.
[0037] The training of the mask generation network G does not rely on additional annotations, but through a self-supervised "feature reconstruction" task. Its training target is: when using the generated mask to perform element-wise multiplication on the feature pyramid, that is, After being masked, the remaining features Should be able to predict the pre-defined pseudo-label (for example, the ecological scene category or principal component obtained by clustering) of the sample with high accuracy through a simple multi-layer perception (MLP) classification head. This forces G to learn to automatically identify and retain those feature regions that are most informative for sample discrimination. The multi-scale feature set After this step of purification, the ecological goal-oriented score (EPG) reaches 0.89.
[0038] This module is responsible for deep fusion and high-order coding of the purified features from different scales to generate the final unified ecological feature representation. This process simulates the cognitive level from local perception to global synthesis. First, the feature maps are upsampled to the same medium size through bilinear interpolation and concatenated in the channel dimension to obtain the aggregated features where, To achieve deep fusion of purified features of different scales, first, the spatial dimension needs to be unified and aligned. For each feature map in the multi-scale feature set where l represents the scale level, taking values from 1 to 4, we use bilinear interpolation method for upsampling to unify the spatial size to a pre-set medium size The upsampled feature map can be simply represented as: , where represents the bilinear interpolation upsampling function, is used to specify the target spatial size, which ensures that the feature maps of different scales are completely aligned in the spatial dimension.
[0039] Let the upsampled feature map be with a shape of , is the channel number of the l-th feature map. On this basis, we perform a channel dimension splicing operation on all the up-sampled feature maps to generate the initial aggregated feature , whose calculation formula is: , where represents the channel dimension splicing function, which combines feature maps of different scales in the channel direction. After splicing, the shape of is This operation integrates discriminative information of different scales into the same feature tensor, fully preserving the original feature expression of each scale. To further enhance the task relevance of the aggregated feature, we input into the channel attention module for dynamic weighting. This module first performs global average pooling on the features of each channel through the "squeeze" operation, compressing the spatial dimension information into a single-value vector.
[0040] Subsequently, a channel attention module (SE-Block) is applied to calculate the weight for each channel to highlight those more important feature channels for the current task: , where GAP represents global average pooling, is the ReLU activation function, , is the fully connected layer weight, is the Sigmoid function, and
[0041] Next, the enhanced feature is input into a Transformer encoder composed of L layers (usually L = 6), and the self-attention mechanism of the l-th layer in the Transformer encoder is denoted as: , where , , are obtained by linear projection of the input feature, is the dimension of the key vector. After L layers of such encoding, the deep encoded feature is obtained.
[0042] Finally, global average pooling is applied to to compress it into a one-dimensional feature vector, which is projected into a preset 512-dimensional space through a fully connected layer, and then normalized, outputting the final ecological feature vector: , At this point, the distinguishable ecological representation layer has completed its entire task. It takes the raw, mixed multimodal data... Through deep transformation involving feature decoupling, purification, and fusion, it is transformed into a feature vector that possesses high discriminativeness, high robustness, and rich semantic information. This vector, as the core representation of the entire monitoring framework, lays a solid foundation for achieving accurate and dynamic target perception and matching in the next stage.
[0043] S3. Target-aware matching layer The target-aware matching layer is the core of this framework's decision-making process, responsible for matching distinguishable ecological feature vectors. Accurately associate and locate the target with known ecological templates. The input to this layer is a feature vector. and their corresponding geographical locations The output is a series of high-confidence ecological target identification results. The entire processing flow transforms abstract features into concrete monitoring conclusions through three core steps: dynamic anchor box generation, hierarchical diffusion matching, and confidence fusion.
[0044] To accommodate the varying scales of marine targets, ranging from centimeters to kilometers, this module dynamically generates initial spatial hypotheses based on input features. (Preset) Basic anchor frame dimensions Input features Adjustment parameters for each base anchor box, including center point offset, are predicted using a lightweight regression network. and scale logarithmic shift : , Combined with input position Dynamic anchor frame Calculated as: , in: , Finally, a dynamic anchor frame set is obtained. This serves as the initial spatial assumption for subsequent matching. The SD-Match network uses a dynamic set of anchor boxes. and characteristics For input, matching is performed in two stages. First, a coarse matching is performed for each dynamic anchor box. Extract its corresponding local feature vector And calculate its relationship with the template library. cosine similarity Simultaneously, to assess geometric fit, the normalized Wasserstein distance (NWD) is calculated. The anchor frame is then... and templates Standard frame are modeled as two-dimensional Gaussian distributions and with mean at the center of the bounding box, and covariance matrix whose diagonal entries are proportional to the width and height of the bounding box. The squared 2nd Wasserstein distance between the two distributions can be analytically computed, and the NWD is defined as: , where is a normalization constant. The coarse matching comprehensive score is obtained by weighting the two: , For each anchor box , the template with the highest score is selected as its coarse matching class , and the score is recorded. Only anchor boxes with enter the fine matching stage.
[0045] The fine matching stage employs a hierarchical diffusion model to iteratively optimize anchor boxes that pass the coarse screening. With anchor box and its coarse matching class as initial conditions, the anchor box parameters are encoded into a vector . The diffusion model generates an optimized parameter vector through a reverse process. Specifically, at time step , the noise prediction network predicts the noise based on the current parameters , the step number , the local features , and the class
[0046] , Using the predicted noise, the diffusion model's inverse update rule is used to calculate a better parameter estimate : , where , and are diffusion model hyperparameters, is a standard Gaussian noise. After iterations from to , the optimized parameter vector is obtained. This vector is decoded into the fine anchor box , i.e.: , where the decoding operation converts the vector into the coordinates and dimensions of the bounding box. Therefore, the intermediate variable in the diffusion model converges to And directly output as a fine anchor frame. .
[0047] After obtaining the fine anchor frame Then, features are re-extracted based on their location. And calculate its relationship with the template. cosine similarity as well as With standard frame of value The perfect match score is: , To improve the reliability of the results, the confidence scores are integrated with those of data-driven and knowledge-driven approaches. The feature matching confidence score is obtained by normalizing the fine-match score: , Ecological logic confidence is determined by querying the rule base. The library encodes the constraints between target types and environmental parameters (such as water temperature and chlorophyll). Environmental data vectors are extracted based on the position of the precise anchor frame. ,calculate: , The final confidence level is the weighted sum of the two: , Only when The result was then adopted. The target perception matching layer ultimately outputs a set of structured monitoring results: , Each result includes the target type and the optimized geographic bounding box, i.e., the fine anchor box. And fusion confidence. This output This provides a direct and reliable input for subsequent monitoring, feedback, and application.
[0048] S4. Monitoring Result Output and Feedback Layer This layer serves as the terminal of the entire monitoring process and undertakes two core tasks: first, to convert the high-confidence results generated by the target perception and matching layer into standardized monitoring reports; and second, to build a closed-loop feedback mechanism to use the low-quality matching samples generated in this monitoring to perform lightweight fine-tuning of the core model, thereby achieving online evolution of system performance.
[0049] This module receives the final result set from the target perception and matching layer. For each valid result, the system executes a standardized output process. First, the precise anchor box... The coordinates are mapped from the image space to geographic coordinate system, generating a standard geographic position description, the positioning error of the process is verified by actual measurement not greater than 5 meters. Subsequently, the system encapsulates the target type , the geographic positioning result and the corresponding final confidence into a structured monitoring record. All records are integrated in chronological order, pushed to the monitoring center in real time through a standard data interface, and automatically visualized on an electronic chart or remote sensing image base map to form a real-time ecological monitoring situation map containing target type, accurate location and confidence score.
[0050] To achieve adaptive optimization of the system, this module designs a feedback process based on incremental learning. The core is to use the low-confidence matching samples generated in this monitoring task to fine-tune the key parameters in the distinguishable ecological representation layer and the target perception matching layer, without starting the full model retraining with high computational cost.
[0051] Specifically, after each monitoring task is completed, the system automatically selects all matching samples with confidence scores below a certain threshold to form an incremental data set . This data set contains the original multi-modal data of these samples and their corresponding initial feature vectors extracted by the distinguishable ecological representation layer. The system caches these "difficult samples" and their features. The goal of model optimization is to fine-tune the network parameters so that they can better handle such samples. This is achieved by minimizing a composite loss function that combines representation learning and matching tasks: , where is the loss designed for feature representation, such as contrastive loss, aiming to bring positive sample features of the same target closer and negative sample features of different targets farther apart, thereby improving the distinguishability of the feature ; is the loss designed for the matching task, such as the focal loss based on the optimized matching results, aiming to make the model pay more attention to the correct classification and positioning of these difficult samples; is the weight coefficient that balances the two losses. By performing a small number of gradient descent optimizations on the small incremental data set , the model parameters can be efficiently updated. This process is executed asynchronously in the background, and after completion, the corresponding parameters of the online model are replaced in a hot update manner, so that the entire system has the ability to continuously learn and improve from actual monitoring errors, significantly improving the generalization and robustness when facing new environments or new interference patterns.
[0052] Experimental verification To verify the effectiveness, robustness, and real-time performance of the marine ecological monitoring method based on ecological representation and target perception matching proposed in this invention in practical applications, this embodiment constructs a comprehensive experimental environment containing multi-source heterogeneous data for detailed testing. The dataset used in the experiment is named MarineEco-Fusion, which consists of 15,000 high-resolution optical images collected by an underwater robot in different sea areas such as deep sea, shallow waters, and coral reefs, as well as 15,000 frames of multibeam sonar and side-scan sonar images acquired simultaneously. To comprehensively evaluate the model performance, the dataset covers five main ecological targets: corals, fish schools, seagrass, marine debris, and seabed sediments. Based on water visibility, the data is divided into a high-resolution group with visibility greater than 5 meters and a turbid, low-light group with visibility less than 2 meters, to focus on testing the model's performance under harsh conditions. The experiment was run on a high-performance computing platform equipped with two NVIDIA A100 graphics cards and an Intel Xeon Gold processor. It was implemented based on the PyTorch 2.0 deep learning framework. The input images were size normalized and trained for 200 rounds using the AdamW optimizer.
[0053] In the quantitative analysis phase, the experiment selected mean precision (mAP@0.5), recall, and F1 score as the main evaluation indicators, and introduced frames per second (FPS) to measure the real-time performance of the monitoring. The method of this invention was rigorously compared with traditional image processing methods, mainstream single-modal deep learning models such as YOLOv8, purely acoustic detection methods, and existing multimodal early fusion methods. The comprehensive performance comparison data shown in Table 1 below demonstrates that, under a unified test set, the method proposed in this invention exhibits significant advantages. Specifically, as... Figure 5 As shown, the proposed method achieves an mAP of 89.5%, which is 10.9 percentage points higher than the 78.6% achieved by the YOLOv8 model using only optical data, and 8.3 percentage points higher than the 81.2% achieved by existing simple fusion methods. Particularly noteworthy is the recall rate, which reaches 87.8%, indicating a significant reduction in missed detections. Although the inference speed of 72 FPS is slightly lower than the 85 FPS of the pure vision model, it still far exceeds the 30 FPS standard typically required for real-time monitoring, demonstrating that the proposed method successfully achieves real-time performance suitable for engineering applications while maintaining high accuracy.
[0054] Table 1. Performance comparison of different methods on the MarineEco-Fusion dataset. Method Input modality mAP@0.5 percentage Recall percentage F1 score Inference speed FPS Baseline method A i.e. traditional image processing Optical only 54.2 48.5 0.51 120 Baseline method B i.e. YOLOv8 Optical only 78.6 75.2 0.77 85 Baseline method C i.e. sonar detection Acoustic only 65.4 60.1 0.62 90 Existing fusion method i.e. early fusion Optical plus acoustic 81.2 79.5 0.8 45 Inventive method Optical plus acoustic 89.5 87.8 0.88 72
[0055] Example 2 This embodiment provides a marine ecological monitoring system based on ecological representation and target perception matching, including: a data acquisition module configured to acquire multi-source marine monitoring data; a preprocessing module configured to perform multi-modal data preprocessing based on the acquired multi-source marine monitoring data; a feature module configured to generate distinguishable ecological feature vectors according to the preprocessed data; a confidence module configured to generate a fusion confidence using the distinguishable ecological feature vectors for target perception; a monitoring module configured to generate a monitoring report based on the fusion confidence result and feedback.
[0056] A computer-readable storage medium, wherein a plurality of instructions are stored, the instructions being adapted to be loaded and executed by a processor of a terminal device to implement the marine ecological monitoring method based on ecological representation and target perception matching.
[0057] A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement instructions, and the computer-readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor to implement the marine ecological monitoring method based on ecological representation and target perception matching.
[0058] The above are preferred embodiments of the present application, which do not limit the protection scope of the present application, and therefore: any equivalent changes made in the structure, shape, and principle of the present application should be covered within the protection scope of the present application.
Claims
1. A marine ecological monitoring method based on ecological representation and target perception matching, characterized in that, include: Acquire multi-source ocean monitoring data; Multimodal data preprocessing was performed based on the acquired multi-source marine monitoring data; Generate distinguishable ecological feature vectors based on the preprocessed data; Target perception is generated and fusion confidence is achieved using distinguishable ecological feature vectors; A monitoring report is generated and fed back based on the fusion confidence results.
2. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 1, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data includes utilizing... The spatiotemporal synchronization transformation operator performs time synchronization on the original multimodal dataset. It contains four types of sensor data tuples: optical remote sensing images Infrared remote sensing images Underwater acoustic spectrum diagram and buoy sensing timing data First, for each data tuple, define a time calibration function. By fitting historical synchronization point data using the least squares method, the fixed offset and linear drift of the sensor clock are modeled, and the calibrated timestamps are obtained. The calculation is as follows: ,function The output ensures that for all Satisfying constraints First, time alignment of all modal data is achieved with sub-second precision; second, spatial synchronization is performed, using the original coordinates attached to all data. Transform to the universal WGS-84 geodetic coordinate system using a parameterized mapping function. The parameters are determined by the extrinsic calibration data of each sensor. Finally, based on unified time and space coordinates Regarding the original data content Spatial resampling and temporal interpolation are performed to fill or align to a standard spatiotemporal grid, using resampling operators. Represented as: in, Bilinear interpolation was used for image data, and linear interpolation was used for time-series data, resulting in a spatiotemporally synchronized dataset. .
3. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 2, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data also includes using convolution kernels. Extract noise-sensitive feature maps from images , convolution kernel Optimization is performed to enhance the local contrast difference between noise and signal regions: ,in This represents a two-dimensional convolution operation. It is Polarization difference edge detection kernel; based on feature map At each pixel position Calculate its local neighborhood The statistical properties within the area are used to generate a soft noise probability mask by comparing the deviation of the central pixel feature value from the local mean. : in, and Neighborhood Mean and standard deviation of the internal eigenvalues; It is a very small positive number used to prevent division by zero errors; It is the Sigmoid function, which maps the standardized scores to the (0,1) interval; then it uses a noise mask. The guided filtering process replaces the original values with Gaussian-smoothed image values in high-noise-probability regions indicated by the mask to suppress noise; in low-noise-probability regions, it preserves the original image details to protect the signal. This can be described as follows: in, This represents element-wise multiplication. Is with A matrix of all ones of the same size This represents the convolution operation. The standard deviation is Two-dimensional Gaussian smoothing kernel; for After performing operations on all images and spectral data, the denoised dataset is obtained: 。 4. The marine ecological monitoring method based on ecological representation and target perception matching according to claim 3, characterized in that, The multimodal data preprocessing based on the acquired multi-source ocean monitoring data also includes data augmentation transformation, including the denoised acoustic spectrogram. Define time-domain stretching transformation frequency domain translation transform To simulate the effects of changes in sound wave propagation speed and the Doppler effect of the sound source, the enhanced acoustic spectrum was obtained. Produced by the action of changing the order: Among them, time-domain stretching parameters Uniform random sampling is performed within the interval [0.8, 1.2]; frequency domain shift parameter Uniform random sampling is performed within the interval [-2.0, +2.0] kHz; for the denoised image patch... Apply space and light intensity Transformation, through random rotation The angle is Internal sampling was performed to simulate different observation perspectives, followed by random cropping and scaling. Crop the original image The area was then resampled to a standard size, and finally random color dithering was performed. Enhanced image patches The expression generated by transformation and composition is as follows: This will enhance the transformation. Applied to The samples in the dataset are used to obtain the final standardized preprocessed dataset. : 。 5. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 4, characterized in that, The process of generating distinguishable ecological feature vectors based on preprocessed data includes receiving a standardized multimodal dataset from the preprocessing layer. The Task Adaptive Feature Decoupling Network (TA-FDN) is used to generate the decoupled target feature tensor. and background feature tensor The TA-FDN network employs a parallel dual-branch encoder architecture, where each branch consists of a convolutional layer followed by normalization and ReLU activation functions. The target branch is utilized... and background branches For input Perform parallel processing: in, and These are the convolution weight parameters for the two branches; to dynamically adapt to different monitoring tasks, a lightweight task adaptive unit (TAU) is introduced, which generates a pair of dynamic modulation vectors based on the metadata of the input data. C represents the number of feature channels, and the initially extracted features are weighted at the channel level. ,in, Broadcast multiplication indicating channel direction; to make and To achieve decoupling, a decoupling loss function based on minimizing mutual information is defined. Utilizing mutual information The optimization objective is to measure the degree of dependence between two random variables. Finally, a variational estimation method based on adversarial learning is adopted, and an auxiliary discriminator network D is introduced to distinguish feature pairs. .
6. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 5, characterized in that, The process of generating distinguishable ecological feature vectors based on preprocessed data also includes using the self-supervised masked real feature enhancement network AIM-Marine to purify the decoupled features. First, the decoupled dual-stream features are concatenated and input into a lightweight multi-scale feature encoder E to construct a four-level feature pyramid. in, This indicates splicing along the channel dimension. This is the l-th stage of the encoder, outputting a feature map. The spatial dimensions decrease progressively, thereby capturing information from fine-grained details to global context; then for each level of the pyramid... A lightweight mask generation network G_l generates a spatial binary mask for it. To enable gradient backpropagation in discrete mask decisions, the Gumbel-Softmax reparameterization technique is employed. Specifically, For each spatial location Output a two-dimensional logic value Then, the hard mask value at that location is obtained through Gumbel-Softmax sampling: The mask generation network G is trained through a self-supervised feature reconstruction task. The training objective is to perform element-wise multiplication of the feature pyramid using the generated mask, i.e. Afterwards, the remaining features after being obscured It should be able to predict predefined pseudo-labels for samples using a multilayer perceptron (MLP) classification head, forcing G to automatically identify and retain the feature regions with the most information for sample discrimination, resulting in a refined multi-scale feature set. Its Ecological Target Orientation Score (EPG) reached 0.
89.
7. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 6, characterized in that, The process of generating distinguishable ecological feature vectors based on preprocessed data also includes using a hierarchical ecological feature fusion representation network to deeply fuse and encode purified features from different scales to generate a final unified ecological feature representation. First, feature maps from different scales are... Aggregated features are obtained by upsampling to the same medium size using bilinear interpolation and then concatenating them along the channel dimension. Subsequently, the channel attention module SE-Block is applied to calculate weights for each channel to highlight those feature channels that are more important to the current task: Wherein, GAP represents global average pooling. It is the ReLU activation function. , These are the weights of the fully connected layer. It is the Sigmoid function. These are the learned channel attention vectors; then the enhanced features are... The input is a Transformer encoder consisting of L layers to model long-range dependencies within features and to fuse multimodal information. The self-attention mechanism of the l-th layer in the Transformer encoder is denoted as: in, , , They are obtained from the input features through linear projection. The dimension of the key vector, after being encoded through L layers, yields the deeply encoded features. Finally, regarding Global average pooling is applied to compress it into a one-dimensional feature vector, and then it is projected onto a predefined 512-dimensional space through a fully connected layer, and then processed... Normalization, outputting the final ecological feature vector: This enables the processing of raw, mixed multimodal data. Through deep transformation involving feature decoupling, purification, and fusion, it is transformed into feature vectors that possess high discriminativeness, high robustness, and rich semantic information. .
8. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 7, characterized in that, The method of generating fusion confidence scores using distinguishable ecological feature vectors for target perception includes using a target perception matching layer to combine the distinguishable ecological feature vectors... Accurately associate and locate the target with known ecological templates, with the input being a feature vector. and their corresponding geographical locations The output is a series of high-confidence ecological target identification results, achieved through three steps: dynamic anchor box generation, hierarchical diffusion matching, and confidence fusion. Dynamic anchor box generation includes dynamically generating initial spatial hypotheses based on input features and pre-setting... Basic anchor frame dimensions Input features Adjustment parameters for each base anchor frame, including center point offset, are predicted using a lightweight regression network. and scale logarithmic shift : Combined with input position Dynamic anchor frame Calculated as: , in: Finally, a dynamic anchor frame set is obtained. .
9. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 8, characterized in that, The hierarchical diffusion matching includes utilizing the SD-Match network with a dynamic set of anchor boxes. and characteristics For input, matching is performed in two stages. First, a coarse matching is performed for each dynamic anchor box. Extract its corresponding local feature vector And calculate its relationship with the template library. cosine similarity Meanwhile, to assess geometric fit, the normalized Wasserstein distance (NWD) is calculated, and the anchor frame is... and templates Standard frame They are modeled as two-dimensional Gaussian distributions. and Its mean Centered on the box, the covariance matrix The diagonal is proportional to the width and height of the frame, and the square of the second-order Wasserstein distance between the two distributions is given. Analytical calculation, NWD is defined as: in As a normalization constant, the coarse matching comprehensive score is obtained by weighting the two: Finally, for each anchor frame The template with the highest score is selected as its coarse matching category. And record the score. Only when The anchor frames then enter the fine matching stage; in the fine matching stage, a hierarchical diffusion model is used to iteratively optimize the anchor frames that have passed the coarse screening, with the anchor frames... and its coarse matching category As initial conditions, the anchor box parameters are encoded as vectors. The diffusion model generates an optimized parameter vector through a reverse process. Specifically, in time step Noise prediction network Based on current parameters Steps Local features and categories Predicted noise: Using the predicted noise, a better parameter estimate is calculated through the inverse update rule of the diffusion model. : in , and These are hyperparameters of the diffusion model. It is standard Gaussian noise, after being processed from... arrive Through iteration, the optimized parameter vector is obtained. The vector is decoded into a fine anchor box. ,Right now: The decoding operation converts the vector into the coordinates and size of the bounding box; therefore, the intermediate variables in the diffusion model... Eventually converges to And directly output as a fine anchor frame. ; Obtain the fine anchor frame Then, features are re-extracted based on their location. And calculate its relationship with the template. cosine similarity as well as With standard frame of value The perfect match score is: .
10. A marine ecological monitoring method based on ecological representation and target perception matching according to claim 9, characterized in that, The confidence fusion includes fusing data-driven and knowledge-driven confidence scores, where the feature matching confidence score is obtained by normalizing the fine-matching score. Ecological logic confidence is determined by querying the rule base. Obtained, based on the fine anchor frame Location Extraction of Environmental Data Vectors ,calculate: The final confidence level is the weighted sum of the two: Only when When this result is adopted, the target perception matching layer finally outputs a set of structured monitoring results: Each result includes the target type and the optimized geographic bounding box, i.e., the fine anchor box. And fusion confidence.
Citation Information
Patent Citations
Marine ecology-oriented time-space diagram neural network anomaly detection method and system
CN119312267A
Marine remote sensing coastline segmentation method based on text-guided semi-supervised pseudo-tag
CN120219407A
Ocean red tide anomaly detection method and system fusing multi-source remote sensing and graph neural network
CN120656076A
Fine ocean forecasting method based on frequency domain enhanced neural network
CN121167642A
Marine ecological protection and restoration project comprehensive evaluation method based on multi-dimensional analysis
CN121638951A