Marine organism community monitoring method and system based on memory hybrid prototype and knowledge distillation

Through the method of memory mixed prototype and knowledge distillation, problems such as optical and sonar interference, multimodal data fusion and large calculation volume of edge equipment in marine biome monitoring are solved, and high-precision marine biome monitoring and prediction are achieved, supporting marine ecological protection and fishery management.

CN120354318AActive Publication Date: 2025-07-22SHANDONG MARINE RESOURCE AND ENVIRONMENT RESEARCH INSTITUTE (SHANDONG MARINE ENVIRONMENTAL MONITORING CENTER SHANDONG AQUATIC PRODUCTS QUALITY INSPECTION CENTER)

Patent Information

Application Number
CN202510837471.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing marine biome monitoring technologies face multiple challenges in complex marine environments, including the problems of optical equipment being affected by turbidity and insufficient light in water, sonar signals being disturbed by noise, insufficient multimodal data fusion, large amount of edge equipment computing and lack of expertise in timing dynamic characteristics.

Method used

The method of memory mixed prototype and knowledge distillation is adopted to construct a lightweight model to adapt to edge equipment through multi-source heterogeneous data preprocessing, cross-modal teacher model training, student model strengthening knowledge distillation and time series prediction.

Benefits of technology

It improves the accuracy of biome feature extraction, reduces equipment energy consumption, realizes real-time monitoring and long-term battery life, improves the accuracy of abnormal pattern recognition and trend prediction, and provides a scientific basis for marine ecological protection and fishery management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354318A_ABST
    Figure CN120354318A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of marine organism community monitoring, in particular to a marine organism community monitoring method and system based on a memory mixed prototype and knowledge distillation. The method comprises the following steps: carrying out data preprocessing on obtained multi-source heterogeneous monitoring data; constructing a multi-modal joint feature by using the preprocessed data; performing cross-modal teacher model training by using the multi-modal joint features; training a student model based on the historical marine ecological sequence data; performing enhanced knowledge distillation on the student model by using the cross-modal teacher model; performing marine ecological community time sequence prediction by using the trained student model; and performing anomaly detection based on the historical data and the prediction data. Based on the breakthrough of the bottleneck of the existing marine organism community monitoring technology, the multi-dimensional technical improvement and application value are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of marine biological community monitoring, and in particular, to a method and system for marine biological community monitoring based on memory hybrid prototypes and knowledge distillation. Background Art

[0002] Existing marine biological community monitoring technologies have significant bottlenecks in terms of detection accuracy and performance overhead. Traditional monitoring systems mostly rely on single-modal data (such as optical images or sonar signals) and face multiple challenges in complex marine environments. Water turbidity forms a natural barrier to optical monitoring devices. Suspended particles cause light scattering, making the captured biological images blurred and reducing the feature extraction accuracy of vision-based recognition algorithms. Lighting conditions severely restrict the effective monitoring period of optical devices. In deep-sea or night-time environments, insufficient lighting makes it difficult to capture the outlines of organisms, forcing the monitoring system to use artificial light sources, which not only disturbs marine organisms but may also introduce new optical interference. Sonar signal monitoring is also significantly affected by background noise. Natural and human noises such as ship propellers and underwater earthquakes in the ocean are extremely likely to submerge the weak acoustic features of organisms. According to statistics, such interference results in a misjudgment rate of biological recognition as high as 15%-30%.

[0003] In addition, existing marine biological community monitoring technologies have significant deficiencies in the deep fusion and knowledge transfer of multi-modal data. Traditional methods mostly rely on single-modal data or simple feature splicing, making it difficult to cope with noise interference and data heterogeneity problems in complex marine environments and lacking the effective use of professional ecological knowledge. The present invention proposes to construct a cross-modal teacher model by integrating a calibrated language model and privileged knowledge distillation technology, which can dynamically combine professional knowledge in marine scientific literature with real-time monitoring data to generate high-quality feature representations containing deep semantic associations. At the same time, the design of the subtraction cross-attention mechanism can optimize the heterogeneous characteristics of multi-modal data, suppress redundant information, and strengthen the fusion of complementary features, thereby improving the model's ability to extract multi-dimensional features of biological communities under complex conditions such as low light and high turbidity, and solving the problems of insufficient cross-modal association and weak anti-interference ability in traditional methods.

[0004] The problem of high - performance model adaptation during the deployment of edge devices is another bottleneck in the existing technology. Due to high complexity and large computational volume, mainstream monitoring models often face the contradiction of excessive inference latency and high energy consumption when running on terminal devices such as ocean buoys and underwater robots, making it difficult to meet the requirements of real - time monitoring and long - term battery life. Through the privileged knowledge distillation mechanism, the present invention lightweightens the complex teacher model containing rich ecological knowledge, enabling the student model to efficiently reproduce the core behaviors of the teacher model while significantly reducing the number of model parameters and computational overhead. This lightweight design is optimized in combination with the computing power characteristics of edge devices. On the premise of maintaining monitoring accuracy, it effectively improves the running efficiency of the model on low - power terminals, breaks through the technical barrier between "model performance" and "device capabilities", and provides a feasible solution for the long - term continuous operation of ocean monitoring devices.

[0005] In terms of the integration of temporal dynamic features and ecological expertise, the existing technology lacks in - depth modeling of the multi - scale evolution laws of marine biological communities and does not fully utilize domain expertise to guide model learning. The present invention constructs a hybrid prototype learning mechanism based on reconstruction, extracts normal prototypes and periodic prototypes from monitoring data at different time scales, and can effectively capture short - term fluctuations, long - term trends, and periodic patterns in the dynamic changes of the community. Combining the parsing ability of the calibration language model for expertise in ecological literature, knowledge in fields such as species habits and environmental response mechanisms is incorporated into the prototype construction process, enabling the model to not only learn temporal features based on data - driven methods but also use prior knowledge to enhance the understanding of complex dynamic scenarios, improve the accuracy of anomaly pattern recognition and trend prediction, and solve the problems of lack of knowledge guidance and insufficient generalization ability in the temporal analysis of traditional models. Summary of the Invention

[0006] To solve the above - mentioned problems, the present invention provides a method and system for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation.

[0007] In the first aspect, a method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation provided by the present invention adopts the following technical solutions: A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation includes: Obtain multi - source heterogeneous monitoring data; Perform data pre - processing on the obtained multi - source heterogeneous monitoring data; Construct multi - modal joint features using the pre - processed data; Perform cross - modal teacher model training using the multi - modal joint features; Train the student model based on historical marine ecological sequence data; Perform enhanced knowledge distillation on the student model using the cross - modal teacher model; Using the trained student model for predicting the time series of marine ecological communities; Performing anomaly detection based on historical data and predicted data.

[0008] Furthermore, the data preprocessing of the obtained multi-source heterogeneous monitoring data includes denoising the original bi-optical image by using a bilateral filtering algorithm for the bi-optical image, suppressing the scattering noise caused by water turbidity through the joint weighting of spatial proximity and pixel similarity; performing short-time Fourier transform on the time-domain echo signal and generating a time-frequency matrix by intercepting a local signal segment through a Hann window function ; performing normalization processing on numerical data such as temperature and salinity; encoding the ecological relationship triple text into a 768-dimensional semantic embedding vector by using the SBERT model for text data ; ; .

[0009] Furthermore, the training of the cross-modal teacher model by using the preprocessed data includes adding learnable positional encoding PE to the multi-modal joint features to obtain an initial state , and then processing by CLMs to generate context embeddings, including corrective attention, layer normalization, and feed-forward operations. Finally, using the Transformer encoder PT encoder(·), the true value in the time series modality is reconstructed through in the text modality, to achieve cross-modal reconstruction, where the LN and FFN in CLMs are defined as follows: where is the output after passing through the second layer , LN and are learnable scaling and translation parameters, and represent the mean and standard deviation respectively, represents element-wise multiplication. ;

[0010] Furthermore, the training of the student model based on historical marine ecological sequence data includes given historical data , the inverse embedding layer converts into a learnable matrix to capture the time dependence across multiple variables, and first performing normalization processing through a reversible instance to mitigate the impact of distribution shift. Subsequently, the normalized embedding representation Subsequently, it is fed into a Transformer encoder, the TST encoder(·), which is used to model the temporal dependencies among multiple variables, and finally outputs to fuse the features of time series and cross-variables.

[0011] Furthermore, the use of the cross-modal teacher model to perform enhanced knowledge distillation on the student model includes promoting the student model to imitate the behavior of the teacher to achieve positive correlation distillation by aligning the attention maps between the enhanced ET encoder(·) in the teacher network and the student time series TST encoder(·); achieving feature distillation by aligning the embedding spaces of the teacher model and the student model; the overall distillation loss combines the correlation distillation loss and the feature distillation loss to guide the learning process of the student. The total knowledge distillation loss is defined as: where and are hyperparameters used to balance the contributions of the correlation distillation loss and the feature distillation loss .

[0012] Furthermore, the use of the trained student model to perform time series prediction of the marine ecological community includes performing inference using the student model, where is the time series embedding representation from the distilled student model . Input into the projection function to complete future prediction, which is expressed as: where is the input of the historical data of the marine ecological community, represents the prediction result of the marine ecological community, and are learnable parameters. Finally, output is normalized, and the prediction loss uses SL the 1 loss function: where represents the number of prediction samples.

[0013] Furthermore, the anomaly detection based on historical data and prediction data includes concatenating the historical data of the marine biological community and the prediction data along the time axis into the original sequence data , dividing the original sequence data into multiple equal-length subsequences , where each subsequence All are regarded as an independent training time series. By extracting the fragment features at different scales from the normal marine biological community data sequence, a prototype set with size differences is constructed. Among them, by dividing the given sequence into non-overlapping fragments and using the generated sequences are embedded into the high-dimensional feature space through the Transformer encoder for fragment prototype update, and the updated fragment prototypes are used to reconstruct the query sequence

[0014] to obtain the reconstructed sequence, realizing the update of the fragment prototype.

[0014] Furthermore, for anomaly detection based on historical data and predicted data, it also includes learning the periodic prototypes of the normal time series to model the characteristics of different periodic patterns. The time series is transformed to the frequency domain to calculate the average amplitude value through the fast Fourier transform (FFT). After selecting the top k amplitude values for period division, the periodic prototypes are learned for each variable respectively. For each observed variable of the time series, a periodic prototype is learned from the period division and after random initialization

[0015] the prototype is updated by weighted aggregation of each periodic fragment, realizing the update of the periodic prototype. Furthermore, for anomaly detection based on historical data and predicted data, it includes integrating the sequences after update and reconstruction by fragment prototypes and periodic prototypes, constructing the reconstruction loss and entropy loss respectively, and at the same time constructing the periodic loss in the feature space based on the distance between the periodic prototype and the query vector. Finally, the total loss function is obtained according to the reconstruction loss, entropy loss and periodic loss.

[0016] In a second aspect, a marine biological community monitoring system based on memory hybrid prototypes and knowledge distillation includes: A data acquisition module, configured to acquire multi-source heterogeneous monitoring data; A preprocessing module, configured to perform data preprocessing on the acquired multi-source heterogeneous monitoring data; A joint module, configured to construct multi-modal joint features using the preprocessed data; A teacher training module, configured to perform cross-modal teacher model training using the multi-modal joint features; A student training module, configured to train the student model based on historical marine ecological sequence data; A knowledge distillation module, configured to perform enhanced knowledge distillation on the student model using the cross-modal teacher model; A prediction module, configured to perform time series prediction of the marine ecological community using the trained student model; Anomaly detection module, configured to perform anomaly detection based on historical data and predicted data.

[0017] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions adapted to be loaded and executed by a processor of a terminal device to implement the method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation.

[0018] In a fourth aspect, the present invention provides a terminal device including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions adapted to be loaded and executed by the processor to implement the method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation.

[0019] In summary, the present invention has the following beneficial technical effects: Based on the breakthrough of the bottleneck of the existing marine biological community monitoring technology, the present invention realizes multi-dimensional technical improvement and application value. At the level of multi-modal data processing, the cross-modal teacher model constructed by the present invention uses the calibration language model and privileged knowledge distillation technology to deeply integrate the professional knowledge in marine scientific literature into real-time monitoring data. Through the dynamic optimization of heterogeneous data such as optical images and sonar signals by the subtraction cross-attention mechanism, the problem of feature distortion caused by environmental interferences such as water turbidity, insufficient light, and background noise is effectively suppressed, and the extraction accuracy of multi-dimensional features of biological communities in complex scenarios is greatly improved. Compared with traditional single-modal or simple splicing methods, significant breakthroughs have been achieved in both the biometric accuracy and the integrity of community structure analysis, providing a solid foundation for the accurate interpretation of marine ecological data.

[0020] In the field of edge device deployment, the privileged knowledge distillation mechanism of the present invention successfully breaks the contradiction between "model performance" and "device capabilities". By efficiently migrating the core knowledge of the complex teacher model to the lightweight student model, while maintaining the monitoring accuracy, the model calculation complexity and energy consumption are greatly reduced. This innovation enables terminal devices such as marine buoys and underwater robots to significantly shorten the inference latency when running the monitoring model, and effectively improves the device battery life. It not only meets the real-time monitoring requirements in the dynamic change scenarios of marine organisms, but also provides a low-cost and sustainable solution for long-term and continuous marine ecological monitoring, strongly promoting the transformation of marine monitoring devices from cloud dependence to edge autonomy.

[0021] In terms of integrating temporal dynamic features and ecological knowledge, the hybrid prototype learning mechanism accurately extracts normal prototypes and periodic prototypes from multi-scale time series data, combines the ecological expertise parsed by the calibration language model, and constructs a monitoring system with deep semantic understanding ability. This mechanism can effectively capture short-term fluctuations, long-term trends, and periodic patterns in the evolution of marine biological communities, significantly improving the accuracy of anomaly pattern recognition and trend prediction in complex dynamic scenarios such as red tide early warning and fish migration prediction. It makes up for the deficiencies of traditional models in lacking knowledge guidance and insufficient generalization ability in time series analysis, provides a scientific and reliable decision-making basis for applications such as marine ecological protection and fishery resource management, and realizes the leap from data perception to knowledge-driven in marine biological community monitoring. Brief Description of the Drawings

[0022] Figure 1 is a schematic diagram of a method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation according to Embodiment 1 of the present invention; Figure 2 is a schematic diagram of the prediction result according to Embodiment 1 of the present invention. Detailed Description of the Preferred Embodiments

[0023] The present invention will be further described in detail below with reference to the accompanying drawings.

[0024] Embodiment 1 Referring to Figure 1 , a method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation in this embodiment includes: Obtain multi-source heterogeneous monitoring data; Perform data preprocessing on the obtained multi-source heterogeneous monitoring data; Construct multi-modal joint features using the preprocessed data; Train a cross-modal teacher model using the multi-modal joint features; Train a student model based on historical marine ecological sequence data; Perform enhanced knowledge distillation on the student model using the cross-modal teacher model; Perform time series prediction of the marine ecological community using the trained student model; Perform anomaly detection based on historical data and predicted data.

[0025] Specifically: S1. Data input and preprocessing, The data input module of the present invention integrates multi-source heterogeneous monitoring data, including marine bio-optical images, high-frequency sonar echo signals, environmental sensor time-series data, and structured knowledge from historical literature. In view of the characteristics and noise interference problems of different modal data, in the preprocessing stage, through multi-dimensional cleaning, feature fusion, and time-series alignment technologies, a standardized data set suitable for pre-training of large language models (LLMs) and training of anomaly detection models is constructed.

[0026] S1-1. Optical image processing, First, the bilateral filtering algorithm is used to denoise the original image By jointly weighting the spatial proximity and pixel similarity, the scattering noise caused by water turbidity is suppressed: Where is the spatial proximity kernel, is the pixel similarity kernel, i , j , k , l are coordinate values, is a hyperparameter controlling the kernel width. After denoising, semantic segmentation is performed through the U-Net network to output a pixel-level biological contour mask matrix , where C is the number of class channels. Further, a high-dimensional visual feature vector is obtained through global average pooling: Where is the activation value of the c th channel at the spatial position ( i , j ), is the c th element of the output vector, and the average values of all channels are concatenated into the visual feature vector : S1-2. Sonar signal processing, Perform short-time Fourier transform on the time-domain echo signal and intercept local signal segments through the Hanning window function to generate a time-frequency matrix: Where the window function satisfies: Where T is the window length, and Mel Frequency Cepstral Coefficients (MFCC) are extracted based on the time-frequency matrix. First, the linear frequency is converted to Mel frequency through the Mel filter bank, and then 13-dimensional basic features are obtained through the discrete cosine transform. After calculating the first-order and second-order differences, they are concatenated into a 39-dimensional voiceprint feature vector. : Among them is the combination of the Mel filter and logarithmic energy compression: Among them represents the weight of the m th Mel filter, and the input represents the i th frame's k th frequency component. Through the Mel filter bank weighted summation, and then taking the logarithm to compress the dynamic range, and then using the discrete cosine transform (DCT) for dimensionality reduction: This patent sets 40 Mel filters, that is, M = 40.

[0027] S1-3. Environmental parameter processing, Normalize numerical data such as temperature and salinity : Among them , is the mean and standard deviation of the training set, N is the total number of samples, To avoid a zero denominator, the normalized vector .

[0028] S1-4. Literature knowledge encoding, In the data preprocessing stage, the present invention uses the SBERT model to encode the ecological relationship triple text into a 768-dimensional semantic embedding vector . SBERT encodes the text through a siamese network architecture. First, the input text k is tokenized into a token sequence, and after being processed by the WordPiece tokenizer, it generates , and appends the position encoding p and segment encoding s to form the input representation: To obtain the output vector of the first token, and generate the final semantic embedding through the linear mapping layer: Among them W is the weight matrix of the mapping layer, which adjusts the expression space of semantic features by learning the weight relationship between different dimensions. b is the bias vector, which increases the fitting ability of the model and prevents the output from always being centered at the origin.

[0029] S1-5. Feature concatenation The finally generated semantic vector and the visual features , sonar features , environmental features are concatenated by dimension to form multi-modal joint features : This process realizes the numerical representation of ecological knowledge and its alignment with multi-modal monitoring data in the semantic space. It not only retains the original features of biological morphology, acoustic signals, and environmental parameters but also embeds domain knowledge such as species habits and ecological rules, providing an input basis driven by both "data - knowledge" for the subsequent pre-training of large language models. This standardized input enables the model to learn both the data distribution law and professional knowledge constraints during training, effectively improving the accuracy and interpretability of community structure analysis and anomaly pattern recognition in complex marine environments. It provides a standardized input integrating domain knowledge for the subsequent model.

[0030] S2. Cross-modal teacher model training S2-1. Divide the dataset The multi-modal joint features F are divided into historical data and the true values .

[0031] S2-2. Correct the attention mechanism First, learnable positional encoding PE is added to the multi-modal joint features to obtain the initial state : Then, these initial state representations are processed by CLMs to generate context embeddings. This process includes a series of transformation operations: corrective attention, layer normalization, and feed-forward operations, which gradually refine the token representations through a multi-layer network: Among them represents the intermediate representation after the i -th layer applies and , D represents the hidden dimension of the language model.

[0032] The patent designs a corrective attention mechanism, FocusAtt, to enhance the masked multi-head self-attention (MMSA) for processing multi-modal marine ecological data in large language models (LLMs). Traditional MMSA often has difficulty distinguishing the importance of cross-modal and intra-modal interactions, leading to data entanglement problems. The calibrated attention mechanism (top) strengthens intra-modal interactions by reducing cross-modal interaction weights (e.g., the interaction between the time series token "10" and the text token "were"). The formulation of FocusAtt is as follows: where is the output after the first LN , and the superscript i represents the i th sample in the batch, i.e., the index for distinguishing different input data when processing multiple samples. represents the attention mechanism, which calculates the attention-weighted combination of and key , and the query, key, and value are obtained from the input through linear transformations , , and and respectively. represents the dimension of the key vector, which is used to scale the dot-product attention scores to ensure numerical stability. adjusts the attention scores by suppressing cross-modal interactions, i represents the i th row in the attention matrix, corresponding to the position index of the i th token in the sequence.

[0033] In CLMs, LN and FFN are defined as follows: where is the output after the second layer LN . and are learnable scaling and translation parameters, and represent the mean and standard deviation respectively, represents element-wise multiplication.

[0034] The output of the CLM module is denoted as and , Subsequently, the embedding representation of the last token is extracted from these outputs. Due to the masked attention mechanism in LLMs, the token at the end of the prompt sequence can aggregate the most comprehensive knowledge. Specifically, the representation of the last token at a given position is only affected by the tokens in front of it.

[0035] To utilize this property and reduce computational overhead for efficient knowledge distillation, the last token embeddings are extracted from and respectively: and . This strategy retains the core information for subsequent distillation while ensuring computational efficiency.

[0036] S2-3. Negative Cross-Attention, A Negative Cross-Attention (NCA) mechanism is designed to eliminate the text information embedded in the last token representation, ensuring that the remaining embeddings are still highly relevant to time series prediction.

[0037] In NCA, first, layer normalization and projection functions and are applied to the projected embeddings of the ground truth , and . Then, the channel-wise similarity matrix is calculated through matrix multiplication: where represents matrix multiplication.

[0038] Next, channel-level feature aggregation is performed by multiplying with . Subsequently, by subtracting this intersection from and , and then passing through layer normalization and a feed-forward layer, the refined ground truth prompt embedding is obtained: where represents the refined ground truth embedding vector. Here is a linear layer, represents subtraction. NCA retains the privileged knowledge for subsequent distillation by refining the ground truth prompt embedding. Additionally, to avoid redundant processing with frozen CLMs, we store the embeddings after the subtraction operation for efficient reconstruction.

[0039] S2-4. Cross-modal Reconstruction, This task utilizes the Transformer encoder PT encoder(·) to reconstruct the ground truth in the time series modality through in the text modality. Inside the PT encoder(·), first undergoes layer normalization processing at the layer: where represents the normalized intermediate embedding. has the same structure as before.

[0040] The normalized embedding is then processed through the multi - head attention layer (denoted as ) in the PT encoder(·). The output result is combined with the input through a residual connection: where , , and are learnable linear projections. The attention mechanism calculates the dependencies between feature dimensions.

[0041] The output is then normalized through another layer of and processed by the feed - forward network . The final result is combined with the input through a residual connection: where, represents the output of the PT encoder(·), denoted as for simplicity, while has the same structure as before.

[0042] The reconstructed time series is then generated by applying a projection function to : where represents the reconstructed time series, and are learnable parameters. The reconstruction loss is defined using the smooth L1 loss ( ): where, G represents the total number of reconstructed samples. SL The L1 loss function ensures robustness to external outliers while maintaining sensitivity to small errors.

[0043] S3. Student model training The student model processes historical marine ecological sequence data through invertible instance normalization, an inverse embedding layer, a time series Transformer, and a projection layer. Given historical data , the inverse embedding layer converts into a learnable matrix to capture the temporal dependencies across multiple variables. First, it is normalized through invertible instance normalization to mitigate the impact of distribution shift. Subsequently, the normalized data is embedded as follows: where represents the output of the inverse embedding layer. T represents the hidden dimension of the embedding, and and are learnable parameters.

[0044] The embedded representation is then fed into a Transformer encoder TST encoder(·), which is used to model the temporal dependencies between multiple variables. The encoder consists of multiple layers, and each layer follows the following processing flow: where represents the intermediate embedding after layer normalization, and is the output after being combined with the input through a residual connection. The multi-head self-attention mechanism in the encoder calculates the attention-weighted representation: where , , and are learnable projection matrices, is the dimension of the key vector. The final output from TST encoder(·) fuses the temporal and cross-variable features, which will be used for downstream tasks.

[0045] S4. Reinforcement knowledge distillation Reinforcement knowledge distillation focuses on transferring the privileged representations in the cross-modal teacher model to the lightweight student model through correlation and feature distillation.

[0046] S4-1. Correlation distillation Correlation distillation prompts the student model to mimic the teacher's behavior by aligning the attention maps between the Enhanced Transformer (ET encoder(·)) in the teacher network and the Temporal Series Transformer (TST encoder(·)) in the student. The rectified attention map of the teacher guides the construction of the student's attention distribution to maintain the correlation dependencies between features. Specifically, the attention matrix obtained from the ET encoder(·) and the one obtained from the TST encoder(·) are averaged over all attention heads in the last encoder layer to form a unified representation. The correlation distillation loss is defined as: where, ensures that the student reduces the sensitivity to outliers while replicating the teacher's contextual understanding of feature features.

[0047] S4-2. Feature distillation, Feature distillation aims to align the embedding spaces of the teacher model and the student model. The teacher model generates privileged embeddings through its ET encoder(·) , while the student model replicates this process through the TST encoder(·) of its temporal series transformer to produce embeddings . The feature distillation loss is also implemented using the smooth SL 1 loss function: This process ensures that the student model effectively captures the rich features learned by the teacher model, thus achieving robust and accurate knowledge transfer.

[0048] The overall distillation loss combines the correlation distillation loss and the feature distillation loss to guide the learning process of the student. The total knowledge distillation loss is defined as: where, and are hyperparameters used to balance the contributions of the correlation distillation loss and the feature distillation loss .

[0049] S5. Marine ecological community time series prediction, (1) The time series prediction module uses the well-trained student model for efficient prediction. During the test process, only the student model is used for inference. Specifically, is the time series embedding representation from the distilled student model . Subsequently, Input into the projection function to complete future predictions. Its mathematical expression is as follows: Where is the input of historical data of the marine ecological community, represents the prediction result of the marine ecological community. and are learnable parameters. The final output will be normalized. The prediction loss uses SL 1 loss function: Where, represents the number of prediction sample sizes. The loss function of the proposed TimeKD method consists of four parts: reconstruction loss , correlation distillation loss , feature distillation loss and prediction loss . We combine these losses to obtain the overall loss function as shown below: Where, , and are hyperparameters for balancing.

[0050] (2) Processing of marine biological community data sequences, This patent realizes the dynamic assessment of community status through the anomaly detection of fusing historical data and real-time data, and constructs an ecological early warning mechanism based on the anomaly detection of prediction data. Specifically, the system first establishes a normal mode benchmark for historical observation data, then detects the current community anomaly status through real-time data stream analysis, and finally realizes the early warning of ecological risks in advance by combining the deviation degree of prediction data.

[0051] First, the historical data of the marine biological community and the prediction data are concatenated along the time axis into the original sequence data : After that, the original sequence data is divided into multiple equal-length subsequences , where each subsequence is regarded as an independent training time series. Let , where L is the sequence length, represents the observation vector at time t . The goal of the marine biological community monitoring system is to learn knowledge from known normal data sequences and for unknown data sequences Generate status labels , where , 0 indicates normal and 1 indicates abnormal. For a reconstruction-based model, the anomaly score is obtained through: where is reconstructed through learned knowledge observation value, represents the L2 norm. Finally, the anomaly label is determined by the threshold : If then , otherwise .

[0052] S6. Prototype learning for marine biological community fragments To capture multi-scale temporal information, the model constructs a set of prototypes with size differences by extracting fragment features at different scales from the normal marine biological community data sequence. This method encodes local information with different time spans into prototypes of corresponding sizes, thereby realizing the embedding of multi-perspective representations of normal features.

[0053] S6-1. Average pooling Given the input data , it is divided into several non-overlapping fragments of length . All values within each fragment are averaged according to the formula: where . The newly generated sequence is , where . In particular, when z = 1, is the original time series X . After passing through the pooling layer with a kernel size of 1, the original information is retained. Through the average pooling operation, different z values of Xᶻ can capture local information with different time spans, thereby learning temporal dependencies with different durations.

[0054] S6-2. Update fragment prototypes To learn prototypes in the feature space, the generated sequence is embedded into a high-dimensional feature space through a Transformer encoder: where , which is used to learn the segment prototypes at different scales. The sequences at different scales are embedded into a high-dimensional feature space through a Transformer encoder to learn the feature prototypes. The working principle of the encoder part is elaborated below: Specifically, the pooled sequence is first embedded into a high-dimensional space through a linear layer to represent more complex feature information and potential laws; then, through multiple Transformer modules, the long-term temporal dependencies are captured, and finally, a query sequence for learning multi-scale segment prototypes is generated. . This feature will be fed into the time segment query update module to learn the corresponding time segment prototypes for each variable. It should be noted that since the time series exhibits different variable correlation characteristics at different time segments z , this architecture separately configures an encoder for each time segment.

[0055] The randomly initialized segment prototypes are updated through an update gate based on the query sequence. The similarity matrix between the segment prototypes and the query sequence can be calculated as follows: where is the temperature parameter. It should be noted that the segment prototypes need to contain not only new information but also historical information. Therefore, the proposed update gate is used to update , and its update formula is: where "◦" represents the element-wise multiplication operation, and the calculation formula is: where is the sigmoid activation function. and are randomly initializable learning matrices used to adjust the retention and removal degrees of historical prototypes and new information.

[0056] S6-3. Query Reconstruction During the reconstruction process, the segment prototypes need to effectively reconstruct the query sequence. At each scale z , the updated segment prototypes are used to reconstruct the query sequence to obtain the reconstructed sequence : Indicates a query For the segment prototype of the attention weight matrix.

[0057] Reconstruction query sequences of different scales Need to be appropriately integrated to reconstruct considering different temporal dependence lengths . Recall the average pooling strategy, The local information of is included in . In other words, at the scale, contains the normal information. Therefore, it is natural to consider ( j = 1, 2,..., m ) to reconstruct , All relevant are integrated through a neural network to reconstruct : Finally, the reconstructed sequence is concatenated with the original scale features and decoded back to the original space through the normal segment prototype to obtain the reconstruction result X of the input sequence .

[0058] S7. Marine biological community periodic prototype learning, Most time series in the real world have multiple periodicities, which interact with each other and jointly present the overall change trend of the time series. In addition to the regular segment prototypes, this patent also models the characteristics of different periodic patterns by learning the periodic prototypes of normal time series.

[0059] S7-1. Period division, The period division of a time series depends on the frequency information of the sequence in the frequency domain. To this end, the time series is transformed to the frequency domain through the fast Fourier transform (FFT), and the average amplitude value is calculated according to the following formula: .

[0060] Considering the sparsity of the frequency domain, only the first k amplitude values are used for period division, where: argTopK( )Indicates selecting the first k amplitudes, given a period p ∈ P , the time series X can be divided into segments (padding zeros at the end). Since each observed variable has a unique variation period, it is necessary to learn the periodic prototype for each variable separately. Therefore, extract the univariate sequence from , divide it into N segments of length p and reshape it into . Finally, X is recombined into to learn the periodic prototype.

[0061] S7-2. Update the periodic prototype, Embed through the Transformer encoder to obtain . For each observed variable of the time series, learn a periodic prototype from the period division , where. After randomly initializing , update the prototype by weighted aggregation of each periodic segment: where: where and are learnable matrices. Finally, obtain the set of periodic prototypes C for variables, which is used for query sequence reconstruction.

[0062] S7-3. Reconstruct the query Query reconstruction. Considering the correlation between variables, each query vector can be reconstructed through the updated periodic prototype to obtain , and its calculation formula is: Concatenate the reconstructed periodic segments with to generate , where each . After summarizing the reconstructed segments of all C variables, is reconstructed into through the decoder, representing the reconstructed data obtained from the i th period. Finally, the reconstruction result of the periodic branch is 。

[0063] S8. Data integration After the reconstruction is completed through the segment prototype and the period prototype respectively, finally, by and performing a weighted sum to reconstruct X to obtain : where is a hyperparameter.

[0064] S9. Abnormal detection of marine biological communities, S9-1. Loss function and anomaly score, Generally, in the training stage, the reconstruction loss is undoubtedly a component of the loss function, and its formula is: where represents the Frobenius norm. In addition, having too many prototypes for reconstruction may over-interpret the normal information in the training time series. To ensure that only the most relevant prototypes are presented during the reconstruction, a sparsity constraint needs to be imposed on the reconstruction weights to reduce the likelihood of overfitting problems. In this paper, the entropy loss is used as the sparsity constraint on the reconstruction weights: because all period prototypes are required for the reconstruction process. Specifically, the period prototypes must appropriately characterize the periodicity of the training time series as much as possible. Therefore, based on the distance between the period prototype and the query vector a period loss in the feature space is designed: where represents the th variable in the feature space i with a period of p and the j th query vector. To sum up, the total loss function LOSS is , a weighted combination of and where , and are the adaptive parameters of different loss terms.

[0065] To detect anomalies, significant deviations in the input space and feature space are designed as anomaly scores. In the input space, the reconstruction error is usually regarded as an anomaly score. Finally, the anomaly score is calculated by a preset threshold. Determine sample category: in yes At the moment t The specific value of y =1 indicates abnormality, y =0 means normal, and accurate identification of abnormal points in the marine biological community data sequence is achieved through comprehensive multi-dimensional deviations.

[0066] Experimental verification: Table 1 Comparison of monitoring effects The comprehensive analysis results show that the proposed method has significant advantages in both prediction accuracy and stability. Figure 2 shown.

[0067] The prediction curve and the true value are highly consistent in overall trend, especially during critical periods of drastic data fluctuations. This method can still accurately track the actual trend of changes, avoiding the common prediction lag or over-smoothing phenomenon of baseline models. This excellent fitting effect is fully verified in the quantitative indicators in Table 1: the precision (precision, measuring the accuracy of predicted positive examples) is improved by 7.9%, the recall (recall, measuring the ability to detect true positive examples) is improved by 8.5%, and finally the comprehensive evaluation indicator F1 value (F1-score, the harmonic mean of precision and recall) reaches 92.76%, an increase of 8.3% over the optimal baseline model. It is worth noting that the simultaneous improvement of the three indicators reveals the advantages of the method in many aspects: the improvement of precision indicates that the false alarm rate of the prediction results is significantly reduced; the greater increase in recall indicates that the method has a stronger ability to capture positive samples; the combined effect of the two results in a balanced improvement in the F1 value. Figure 2 This is manifested as: the prediction curve can respond to sudden changes in the true value in a timely manner (high recall rate characteristic) without generating false fluctuations (high precision rate characteristic). The experiment also found that this method maintains stable performance in different data distribution scenarios. As shown in Table 1, in the extreme case of insufficient training data (dataset HMAP), the F1 value can still be maintained above 0.90, surpassing the baseline model's 0.89. This effectively overcomes the defect that traditional models are prone to overfitting in small sample scenarios.

[0068] Example 2 This embodiment provides a marine biological community monitoring system based on memory hybrid prototypes and knowledge distillation, including: A data acquisition module, configured to: A computer-readable storage medium storing multiple instructions, which are adapted to be loaded and executed by a processor of a terminal device for the method for monitoring a marine biological community based on memory hybrid prototypes and knowledge distillation.

[0069] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the method for monitoring a marine biological community based on memory hybrid prototypes and knowledge distillation.

[0070] The above are all preferred embodiments of the present invention. Without limiting the protection scope of the present invention accordingly, therefore: Any equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation, characterized in that Including: Obtain multi-source heterogeneous monitoring data; Perform data preprocessing on the obtained multi-source heterogeneous monitoring data; Construct multi-modal joint features using the preprocessed data; Train a cross-modal teacher model using the multi-modal joint features; Train a student model based on historical marine ecological sequence data; Perform enhanced knowledge distillation on the student model using the cross-modal teacher model; Perform time series prediction of marine ecological communities using the trained student model; Perform anomaly detection based on historical data and predicted data.

2. The marine biological community monitoring method based on memory hybrid prototype and knowledge distillation according to claim 1, wherein The data preprocessing of the obtained multi-source heterogeneous monitoring data includes denoising the original biological optical image by using a bilateral filtering algorithm, suppressing the scattering noise caused by water turbidity through the joint weighting of spatial proximity and pixel similarity; for the time-domain echo signal performing a short-time Fourier transform, and intercepting a local signal segment through a Hanning window function to generate a time-frequency matrix; performing normalization processing on the temperature and salinity numerical data ; For the text data, the SBERT model is used to encode the ecological relationship triple texts into 768-dimensional semantic embedding vectors .

3. The marine biological community monitoring method based on memory hybrid prototype and knowledge distillation according to claim 2, wherein The cross-modal teacher model training using the preprocessed data includes adding learnable positional encoding PE to the multi-modal joint features to obtain the initial state , and then processing by CLMs to generate context embeddings, including corrective attention, layer normalization, and feed-forward operations. Finally, using the Transformer encoder PT encoder(·), through the in the text modality to reconstruct the true value in the time series modality, realizing cross-modal reconstruction. Among them, the LN and FFN in CLMs are defined as follows: Among them is the output after passing through the second layer LN , and are learnable scaling and translation parameters and represent the mean and standard deviation respectively represents element-wise multiplication 4. The marine biological community monitoring method based on memory hybrid prototype and knowledge distillation according to claim 3, wherein, Training the student model based on historical marine ecological sequence data includes providing historical data , the inverse embedding layer converts into a learnable matrix to capture the temporal dependencies across multiple variables. First, normalization is performed through invertible instances to mitigate the impact of distribution shift. Subsequently, the normalized embedding representation is then fed into a Transformer encoder TST encoder(·) for modeling the temporal dependencies between multiple variables, and the final output is used to fuse the temporal and cross-variable features.

5. A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation according to claim 4, characterized in that, The use of the cross-modal teacher model to perform enhanced knowledge distillation on the student model includes promoting the student model to imitate the behavior of the teacher to achieve positive correlation distillation by aligning the attention maps between the enhanced ET encoder (·) in the teacher network and the student time series TST encoder (·); achieving feature distillation by aligning the embedding spaces of the teacher model and the student model; the overall distillation loss combines the correlation distillation loss and the feature distillation loss , to guide the learning process of the student, and the total knowledge distillation loss is defined as: Among them, and are hyperparameters used to balance the contribution of the relevance distillation loss and the feature distillation loss .

6. The marine biological community monitoring method based on memory hybrid prototype and knowledge distillation according to claim 5, characterized in that, The time series prediction of the marine ecological community using the trained student model includes performing inference using the student model, where is the time series embedding representation from the distilled student model , which is input into the projection function to complete future prediction, expressed as: ​ Among them is the input of historical data of marine ecological communities represents the prediction results of marine ecological communities and is a learnable parameter, and the final output is normalized, and the prediction loss uses SL 1 loss function where represents the number of prediction samples 7. A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation according to claim 6, characterized in that, The anomaly detection based on historical data and predicted data includes the historical data of marine biological communities X h and predicted data X f being spliced along the time axis into original sequence data , dividing the original sequence data into multiple equal-length subsequences , where each subsequence is regarded as an independent training time series. By extracting the fragment features at different scales in the normal marine biological community data sequence, a prototype set with size differences is constructed. Among them, by dividing the given sequence into non-overlapping fragments and embedding the generated sequence into a high-dimensional feature space through a Transformer encoder for fragment prototype update, and using the updated fragment prototype to reconstruct the query sequence to obtain a reconstructed sequence, thereby realizing the update of the fragment prototype.

8. A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation according to claim 7, characterized in that, The anomaly detection based on historical data and predicted data further includes modeling the characteristics of different periodic patterns by learning the periodic prototypes of normal time series, and converting the time series to the frequency domain by fast Fourier transform (FFT) to calculate the average amplitude value. After selecting the first k amplitude values for period division, the periodic prototypes are learned for each variable respectively. For each observed variable of the time series, a periodic prototype is learned from the period division and after randomly initializing , the prototype is updated by weighted aggregation of each period segment to achieve the update of the periodic prototype. ​ 9. A method for monitoring marine biological communities based on memory hybrid prototypes and knowledge distillation according to claim 8, characterized in that, The anomaly detection based on historical data and predicted data includes constructing a reconstruction loss and an entropy loss respectively after integrating the sequences updated and reconstructed through the segment prototype and the period prototype, and simultaneously constructing a period loss in the feature space based on the distance between the period prototype and the query vector ; finally, obtaining the total loss function according to the reconstruction loss, the entropy loss, and the period loss.

10. A marine biological community monitoring system based on memory hybrid prototypes and knowledge distillation, characterized in that, Including: A data acquisition module configured to obtain multi-source heterogeneous monitoring data; A preprocessing module configured to perform data preprocessing on the obtained multi-source heterogeneous monitoring data; A joint module configured to construct multi-modal joint features using the preprocessed data; A teacher training module configured to train a cross-modal teacher model using the multi-modal joint features; A student training module configured to train a student model based on historical marine ecological sequence data; A knowledge distillation module configured to perform enhanced knowledge distillation on the student model using the cross-modal teacher model; A prediction module configured to perform time series prediction of marine ecological communities using the trained student model; An anomaly detection module configured to perform anomaly detection based on historical data and predicted data.

Citation Information

Patent Citations

  • Zero sample cross-modal retrieval method based on Transform network selective distillation

    CN115563327A

  • Track target point prediction method based on knowledge distillation

    CN116579423A

  • Underwater target detection method and system based on multi-knowledge distillation

    CN116612379A

  • Knowledge distillation method and system based on multi-teacher multi-modal model

    CN117669693A

  • Web service quality prediction method based on graph knowledge distillation

    CN117827611A

Cited By

  • Marine ecological abnormity early warning method and system based on multimode sensing and spatio-temporal reasoning

    CN120541729A

  • Floating type storm power generation system control method based on knowledge distillation and zoning strategy

    CN121165594A

  • Floating wind wave power system control method based on knowledge distillation and partition strategy

    CN121165594B

  • Industrial equipment, progressive distillation unsupervised anomaly detection method and device thereof and medium

    CN121808643A

  • Industrial equipment and progressive distillation unsupervised anomaly detection method, device and medium thereof

    CN121808643B