Phytoplankton chromatographic sequence identification method and model building method thereof

By using the PPSA-Net model and employing a twin multi-scale feature encoder and a physical sharpness prior module, the problems of insufficient depth of field and out-of-focus noise interference in phytoplankton identification are solved, achieving efficient full depth of field information aggregation and accurate identification.

CN121838155BActive Publication Date: 2026-05-05OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OCEAN UNIV OF CHINA
Filing Date
2026-03-12
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for phytoplankton identification suffer from feature truncation due to insufficient depth of field in microscopic imaging, as well as computational redundancy and out-of-focus noise interference caused by blindly applying video models, making it difficult to achieve high-precision aggregation and identification of phytoplankton panoramic depth information.

Method used

The PPSA-Net model is adopted, and features are extracted through a Siamese multi-scale feature encoder. It combines a physical sharpness prior module and a depth-aware sequence aggregation module, uses the Laplacian operator and Fourier transform to evaluate image sharpness, and introduces focal length context modeling and physical gating mechanism to perform end-to-end weakly supervised learning.

Benefits of technology

It significantly improves the fine-grained classification accuracy of phytoplankton identification, reduces computational costs, enhances the model's robustness in complex underwater environments, and can effectively aggregate full depth information while suppressing out-of-focus noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838155B_ABST
    Figure CN121838155B_ABST
Patent Text Reader

Abstract

This invention provides a method for identifying phytoplankton tomographic sequences and a model building method thereof, belonging to the field of image enhancement and recognition technology. First, microscopic tomographic sequence data of phytoplankton is acquired, and field of view extraction and serialization reconstruction are performed to construct a stereo dataset. Next, a stereo recognition model incorporating physical perception and sequence aggregation mechanisms is constructed. The model uses a parameter-shared Siamese network to extract semantic features of single frames and introduces a physical sharpness prior module to calculate the spatial frequency domain quality score of the slices. Then, a depth-sensing sequence aggregation module is designed, using the sharpness score as a gating signal and combining it with spatial context information between slices to adaptively aggregate key features with high signal-to-noise ratio. Finally, the model is trained and optimized based on image-level weakly supervised labels to obtain the optimal model. This invention overcomes the problems of information truncation and defocus noise interference caused by the extremely shallow depth of field in high-magnification microscopic imaging, enabling full-depth stereo perception without the need for frame-by-frame detailed annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image enhancement and recognition technology, and in particular relates to a method for identifying phytoplankton tomographic sequences and a method for building a model thereon. Background Technology

[0002] As primary producers in marine ecosystems, phytoplankton community structure dynamics are a core indicator for assessing marine environmental health and carbon sequestration potential. Utilizing underwater microscopy for high-throughput, automated phytoplankton monitoring has become a key technological means for building a "smart ocean" observation network, achieving early warning of red tide disasters, and scientific management of fishery resources. However, phytoplankton are diverse and exhibit subtle morphological differences, especially since many harmful red tide algae are highly morphologically similar to harmless algae. This necessitates monitoring systems with extremely high fine-grained identification capabilities.

[0003] In practical in-situ observation missions, acquiring high-quality images faces severe optical and physical limitations. To accurately capture key micron-level distinguishing features (such as the spiny protrusions of diatoms and the transverse and longitudinal grooves of dinoflagellates), high-magnification microscope objectives with high numerical apertures are typically required. According to the principles of optical imaging, the improvement in spatial resolution inevitably leads to a sharp compression of the depth of field. This means that a single-frame two-dimensional image can only record cross-sectional information of the organism at a specific focal plane, while a large number of three-dimensional structural features distributed outside the focal plane become blurred due to defocus. This "information truncation" phenomenon makes it difficult for traditional single-frame image-based recognition methods to capture the full picture of the organism, resulting in the loss of key classification features and severely limiting the improvement of recognition accuracy.

[0004] To address the information loss caused by insufficient depth of field, existing technologies primarily employ two approaches: First, using Extended Depth of Field (EDoF) technology, which combines pixels from multiple frames at different focal planes into a single fully focused image using image fusion algorithms. However, this method is highly susceptible to introducing non-biological artifacts such as halos, ghosting, or jagged edges during the synthesis process. This not only destroys the original texture details of the organism but also introduces persistent noise interference for subsequent feature extraction. Second, directly transplanting deep learning models from video classification (such as LSTM and 3D-CNN) and treating the microscopic tomography sequence (Z-Stack) as a video stream. However, a microscopic tomography sequence is essentially a slice of a stationary object in spatial depth; there is no temporal causal evolution relationship between frames as in the video stream, and the sequence is filled with numerous invalid off-focus background frames. Blindly applying video models based on temporal assumptions not only introduces significant computational redundancy, making deployment difficult on computationally limited underwater edge devices, but also prevents the model from physically distinguishing between blurred and sharp frames, making it highly susceptible to misleading off-focus noise and unable to accurately aggregate sparsely distributed key three-dimensional morphological features.

[0005] Therefore, developing a stereo recognition method that can adapt to the physical characteristics of microscopic tomography sequences, has the ability to perceive clarity, and is computationally lightweight is of great academic research value and urgent engineering need for overcoming the bottleneck of "unclear visibility and inaccurate identification" in underwater in-situ monitoring and achieving efficient aggregation and accurate identification of panoramic depth information of phytoplankton. Summary of the Invention

[0006] To address the above problems, the first aspect of this invention provides a method for constructing a phytoplankton chromatography sequence identification model, comprising the following steps:

[0007] Step 1: Obtain several sets of phytoplankton micro-chromatographic sequence data containing multiple consecutive focal plane slices;

[0008] Step 2: Perform stereo preprocessing on the tomographic sequence data to construct training and testing sets containing physical depth information;

[0009] Step 3: Build a stereo recognition model PPSA-Net that integrates physical prior and sequence aggregation within a deep learning framework. The model adopts a three-level cascaded architecture: First, the front end uses a parameter-shared Siamese multi-scale feature encoder, which extracts features independently from each frame slice in the sequence through a convolutional network that shares weights in the time dimension, generating a high-dimensional semantic feature sequence. Second, a physical sharpness prior module is set up in parallel, which extracts objective imaging quality indicators from the spatial and frequency domains of the image based on the Laplacian operator and Fourier transform, generating a physical prior vector. Finally, the back end is connected to the depth-sensing sequence aggregation module, which introduces context modeling and physical gating mechanisms. The physical prior vector is used as an attention gating signal to guide the adaptive weighted fusion of semantic features in the sequence dimension, generating a global feature descriptor representing the full depth stereo morphology of the organism, which is then input into the classifier to obtain the phytoplankton category.

[0010] Step 4: Based on the weakly supervised learning strategy, the PPSA-Net model is jointly trained end-to-end and its parameters are optimized using a training set containing only bag-level labels. The model with the best performance is selected as the final model.

[0011] Preferably, the stereoscopic preprocessing in step 2 includes field-of-view extraction, sequence normalization, and tensor reconstruction. First, field-of-view extraction based on the region of interest (ROI) is performed. A threshold segmentation algorithm is used to calculate a binary mask of the organism's outline in the original microscopic field of view, and a minimum bounding rectangle is generated based on this mask to crop out the central region containing complete biological information. Second, equally spaced sequence sampling is performed, assuming the original sequence length is... The target input length is , with step size Perform uniform sampling to construct a standardized input sequence. Finally, batch-folding-based tensor reconstruction is performed, transforming the standardized four-dimensional sequence tensors before inputting them into the network to convert the batch dimensions. With sequence length dimension Merge to generate folded tensors This allows the model to process all slices in parallel during a forward propagation process.

[0012] Preferably, the twin multi-scale feature encoder in step 3 consists of several stacked multi-scale sensing units, for the input feature map Each unit contains three parallel paths:

[0013] The first path uses a standard 3×3 convolutional kernel to extract high-frequency fine texture features; the second path uses a 3×3 dilated convolutional kernel with a dilation rate of d=2 to expand the receptive field and extract large-scale geometric configurations without increasing the number of parameters; the third path uses a 1×1 point convolution for channel dimensionality reduction and feature interaction; the outputs of the three paths are first concatenated to generate intermediate joint features. The calculation formula is as follows:

[0014]

[0015] Subsequently, the joint features are adaptively weighted using a channel attention mechanism to obtain the final output. The calculation formula is as follows:

[0016]

[0017] in, This indicates a channel splicing operation. It is the Sigmoid activation function. For element-wise multiplication, For global average pooling, It is a multilayer perceptron; the first in the sequence Frame Image The feature extraction process satisfies ,in For shared parameter weights.

[0018] Preferably, the physical sharpness prior module in step 3 includes two parallel branches: spatial gradient analysis and frequency spectrum analysis. The specific calculation process is as follows:

[0019] S1, Spatial clarity score calculation; using the Laplace operator. Extract the first Frame Image The second-order differential edge information is obtained, and its variance is calculated as a spatial sharpness score. The calculation formula is:

[0020]

[0021] in, This represents the convolution operation. The Laplacian convolution kernel is defined as follows:

[0022]

[0023] S2, Frequency Domain Sharpness Score Calculation; Perform a two-dimensional Fast Fourier Transform (FFT) on the image to calculate the high-frequency region in the logarithmic amplitude spectrum. Average energy as a frequency domain sharpness score The calculation formula is:

[0024]

[0025] in, This represents the total number of pixels in the high-frequency region. For frequency domain coordinates, This represents the result of the Fourier transform of the image. This indicates a summation operation;

[0026] S3, Physical Prior Fusion Mapping; concatenates the normalized spatial domain score with the frequency domain score, and passes it through a multilayer perceptron. Mapped to physical prior vectors aligned with semantic feature dimensions The calculation formula is:

[0027]

[0028] in, This indicates a normalization operation.

[0029] Preferably, the depth-sensing sequence aggregation module in step 3 introduces focal length context modeling and physical prior gating mechanisms, and its specific execution steps include:

[0030] S1, Focal Context Modeling; The feature sequence output by the twin encoder... The input focal length context submodule is used to extract inter-layer complementary information using a one-dimensional temporal convolutional network, resulting in an enhanced feature sequence. ;

[0031] S2, Physical Gated Weight Generation; Based on Physical Prior Vector Calculate the attention weights for each frame. The calculation formula is:

[0032]

[0033]

[0034] in, and These are the projection matrices for semantic features and physical priors, respectively. For the rating vector, For bias terms, It is the hyperbolic tangent activation function;

[0035] S3, Global Feature Weighted Fusion: The generated weights are used to linearly weight the enhanced feature sequences, generating a global 3D feature descriptor. The calculation formula is:

[0036]

[0037] in, This represents matrix or scalar multiplication. This indicates that for all in the sequence The frames are accumulated.

[0038] S4, Feature Mapping and Probabilistic Classification: A fully connected classification layer is constructed as the output head to map the global 3D feature descriptor to a high-dimensional semantic category space. The Softmax normalization exponential function is used to calculate the posterior probability distribution of the input sample belonging to each phytoplankton category, and the final phytoplankton classification result is output based on the maximum probability. The specific calculation formula is as follows:

[0039]

[0040]

[0041] in, This is the global stereo feature descriptor output in step S3; and These are the learnable weight matrix and bias term of the classifier, respectively; The total number of phytoplankton species (in this embodiment) =18); This is the output predicted probability vector. Indicates that the sample belongs to the first Class confidence level; This is the index for the final determination of phytoplankton categories.

[0042] Preferably, the weakly supervised training strategy in step 4 is to use only the bag-level species category labels to supervise the entire sequence during the model training stage, without providing additional fine-grained annotation information for the clarity or key feature positions of each frame slice in the sequence.

[0043] The cross-entropy loss function is used as the objective function for global optimization. This loss function updates the network parameters by minimizing the difference between the predicted probability distribution and the true label distribution. Its mathematical expression is as follows:

[0044]

[0045] in, Indicates the total loss of the batch. Indicates batch size, This indicates the total number of phytoplankton species. For the first The sample belongs to the first The one-hot encoding of the class's true label is 1 when the sample belongs to that class, and 0 otherwise. The output of the model classification layer The sample belongs to the first The predicted probability value of the class.

[0046] Preferably, a spatiotemporal consistency data augmentation strategy is introduced during training, which mandates that for all data in the same input sequence... Frame images, with the exact same random geometric transformation parameters applied. To maintain the spatial alignment within the sequence, the transformation process satisfies:

[0047]

[0048] in, to This represents the images in the sequence. This indicates specific geometric transformation operations, including random rotation and horizontal flipping.

[0049] A second aspect of the present invention provides a method for identifying phytoplankton chromatographic sequences, comprising the following steps:

[0050] S1, real-time acquisition of high-resolution phytoplankton microchromatographic sequence data to be identified;

[0051] S2, input the microscopic tomographic sequence data of the phytoplankton to be identified into the PPSA-Net stereo recognition model constructed by the construction method described in the first aspect;

[0052] S3 outputs the real-time phytoplankton category.

[0053] A third aspect of the present invention provides a phytoplankton chromatography sequence identification device, the device comprising at least one processor and at least one memory, the processor and the memory being coupled together; the memory storing a computer-executable program of a stereoscopic identification model PPSA-Net constructed by the construction method described in the first aspect; when the processor executes the computer-executable program stored in the memory, the processor executes a phytoplankton chromatography sequence identification method.

[0054] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer-executable program of a stereo recognition model PPSA-Net constructed by the construction method described in the first aspect, wherein when the computer-executable program is executed by a processor, the processor performs a phytoplankton chromatography sequence recognition method.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] This invention addresses the feature truncation problem caused by the extremely shallow depth of field in microscopic imaging in existing phytoplankton monitoring technologies, as well as the computational redundancy and out-of-focus noise interference caused by blindly applying video models. It proposes a tomographic sequence stereo recognition model, PPSA-Net, which integrates physical priors. Building upon the advantages of deep learning feature extraction, this model further incorporates optical imaging mechanisms, improving the accuracy of stereoscopic morphology perception.

[0057] First, the PPSA-Net model innovatively constructs a physical sharpness prior module based on the principles of physical optics. Unlike traditional black-box models that rely entirely on data-driven approaches, this module directly utilizes the Laplacian operator and Fourier transform to quantitatively calculate the physical imaging quality of each frame slice from two dimensions: spatial gradient and frequency energy. This physical prior information, acting as a deterministic gating signal, effectively guides the model to distinguish between high signal-to-noise ratio focused frames and low-information out-of-focus background frames. This allows for precise suppression of out-of-focus noise interference during feature aggregation, significantly improving the model's robustness in complex underwater environments.

[0058] Secondly, the model employs a depth-sensing sequence aggregation mechanism, enabling effective utilization of panoramic depth-based 3D information about phytoplankton. By using a Siamese network to extract features in parallel from the sequence and combining this with one-dimensional focal length context modeling, the model can not only capture texture details within a single frame but also perceive morphological evolution between adjacent slices. A physically-prior-based gating aggregation strategy allows PPSA-Net to automatically filter and fuse key discriminative features distributed across different focal planes (such as spiky protrusions or groove structures at different depths), fundamentally solving the "blind men and the elephant" problem of missing information in single-frame imaging and significantly improving fine-grained classification accuracy.

[0059] Finally, this invention employs efficient weakly supervised learning and spatiotemporal consistency strategies. The model only requires image-packet level category labels for end-to-end training, eliminating the need for expensive and time-consuming frame-by-frame sharpness annotation or keyframe localization, significantly reducing data annotation costs and facilitating large-scale deployment in practical engineering. Simultaneously, the stereoscopic preprocessing and spatiotemporal consistency enhancement strategies ensure the physical alignment of sequence data under geometric transformations, effectively preventing training bias introduced by non-physical deformations.

[0060] In summary, this invention provides a tomographic sequence modeling method that combines physical perception capability with stereo recognition accuracy. It can effectively address practical problems such as insufficient depth of field and severe defocus interference in phytoplankton microscopic imaging, and has significant theoretical innovation and engineering application value. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating the phytoplankton chromatography sequence identification method based on physical priors proposed in this invention.

[0062] Figure 2 This is a schematic diagram of an example of a phytoplankton microchromatographic sequence.

[0063] Figure 3 This is a schematic diagram of the overall structure of the PPSA-Net model.

[0064] Figure 4 This is a schematic diagram of the tensor transformation and parallel computing process based on batch folding.

[0065] Figure 5 This is a schematic diagram of the physical clarity prior extraction and gated injection mechanism.

[0066] Figure 6 A schematic diagram of the module structure for focal length context.

[0067] Figure 7 This is a schematic diagram of the convergence curve of the loss function during the training process.

[0068] Figure 8 This diagram illustrates the impact of input sequence length on model classification accuracy and inference latency.

[0069] Figure 9 This is a schematic diagram illustrating the correlation between model attention weights and image physical sharpness scores.

[0070] Figure 10 A simplified structural diagram of a phytoplankton chromatography sequence identification device that incorporates physical priors. Detailed Implementation

[0071] The invention will be further described below with reference to specific embodiments.

[0072] This embodiment further illustrates the method proposed in this invention through specific experimental procedures. The overall process is as follows: Figure 1 As shown.

[0073] 1. Raw data acquisition and preprocessing

[0074] The microscopic tomographic sequence data used in this embodiment comes from real marine phytoplankton samples collected by a high-throughput in-situ microscopic imaging system. The dataset covers 18 representative phytoplankton categories, totaling 16,029 tomographic sequence samples. It includes not only common dominant species such as *Pseudo-necklea* and *Pterygota*, but also species with specific ecological indicative significance such as *Gnaphalium* and *Chaetoceros*. The detailed category distribution and sample quantity of the dataset are shown in Table 1. Unlike traditional single-frame datasets, each sample in this embodiment is a Z-Stack sequence containing multiple consecutive focal plane slices (tomographic sequence data refers to a set of images continuously captured on the same phytoplankton target by moving the focal plane vertically in the Z-axis direction using an automated microscope; i.e., a multi-focal plane slice sequence). Typical sequence imaging effects are shown in... Figure 2 As shown, this figure demonstrates the tomographic performance of phytoplankton under actual imaging challenges such as tilted posture, complex structure, or extremely shallow depth of field, and fully records the three-dimensional morphological information of organisms at different physical depths.

[0075] Table 1. Details of the category distribution of phytoplankton chromatography sequence datasets

[0076]

[0077] The acquired raw microscopic image sequences need to undergo rigorous stereoscopic preprocessing before being used in model training. The specific process is as follows:

[0078] First, the raw data undergoes field-of-view extraction and cleaning. Since in-situ imaging has a large field of view and may contain suspended particle noise in the background, this embodiment utilizes the Otsu adaptive threshold segmentation algorithm to calculate the binarized contour mask of the organism in the center frame of the sequence, and generates a minimum bounding rectangle based on this mask. Subsequently, the central region containing complete biological information is cropped from the redundant background, and invalid samples with abnormal aspect ratios (e.g., greater than 5:1 or less than 1:5) or excessively small pixel areas (less than 64×64 pixels) are removed to ensure data quality.

[0079] Secondly, sequence length normalization is performed. Considering the impact of water flow disturbances or equipment settings during actual sampling, the number of Z-axis scan steps in the original Z-Stack sequence varies (typically ranging from 10 to 50 frames). To meet the uniform requirements of deep neural networks for input tensor dimensions, this embodiment sets the target input sequence length. It is 16. For the original length is The sequence is sampled using a uniformly spaced sampling strategy with a step size of [missing information]. Keyframes are extracted to construct a standardized input sequence. This operation preserves the physical depth span while removing redundant information caused by excessively high sampling rates.

[0080] Finally, to enrich the diversity of training samples and improve the model's generalization ability, this embodiment introduces a spatiotemporal consistency data augmentation strategy. Specific augmentation operations include: performing random cropping on the image sequence, with a cropping ratio ranging from 80% to 100%; applying random rotations within a range of ±30 degrees to the sequence; and performing a random horizontal flip with a 50% probability. Unlike single-frame image augmentation, this embodiment mandates that all images within the same input sequence undergo this process. Each frame image must be subjected to the exact same random geometric transformation parameters, namely the same rotation angle, the same cropping region coordinates, and the same flip direction. This constraint ensures that the enhanced sequence maintains physical continuity and avoids disrupting the contextual relationships between slices due to independent frame-by-frame enhancement. Furthermore, color dithering was applied to the sequence, with brightness and contrast adjusted within a range of ±0.2 to adapt to varying lighting conditions in different aquatic environments.

[0081] In this embodiment, the preprocessed dataset is randomly divided into a training set, a validation set, and a test set in a 6:2:2 ratio. The training set is used for gradient descent updates of model parameters, the validation set is used for hyperparameter monitoring, learning rate adjustment, and selection of optimal model checkpoints during training, and the test set is used only for final model performance evaluation to ensure the objectivity and rigor of experimental results and prevent evaluation bias caused by data leakage.

[0082] 2. Model Building

[0083] This embodiment, based on the aforementioned phytoplankton microscopic tomography sequence dataset, designs and builds a stereo recognition model, PPSA-Net (Physics-PriorSequence Aggregation Network), within a deep learning framework, integrating physical priors and sequence aggregation. The overall architecture of this model is as follows: Figure 3 As shown, its design is inspired by the cognitive process of biologists observing the three-dimensional structure of organisms by adjusting the focal length of a microscope. The main body of the model consists of three tightly coupled functional modules: a twin multi-scale feature encoder, a physical clarity prior module, and a depth perception sequence aggregation module.

[0084] Before inputting sequence data into the network for feature extraction, in order to adapt to the parallel processing requirements of two-dimensional convolutional neural networks for batch data, the following steps are first performed: Figure 4 The diagram illustrates a batch-folding-based tensor transformation. This process reduces the dimension of the preprocessed input sequence tensor from... Reorganized into ,in For batch size, For sequence length, These represent the number of channels, height, and width of the image, respectively. Through this dimensional folding, each frame slice in the sequence can be fed into the encoder as an independent sample in parallel, thus significantly improving computational efficiency without destroying the internal information of the sequence.

[0085] To address the characteristics of phytoplankton samples, which simultaneously exhibit micron-level fine textures (such as diatom bristles and flagella) and larger overall geometric configurations (such as cell outlines and chain-like structures), this embodiment designs a twin multi-scale feature encoder. This encoder employs a parameter-sharing mechanism, meaning that for each slice in the input sequence, it processes the same network structure and shares weight parameters to generate corresponding semantic features. This design not only significantly reduces the number of model parameters but also ensures the translation invariance of the feature extraction process to the absolute position of the slice in the sequence. The core component of the encoder is a multi-scale perceptual unit, each containing three parallel computation paths, each used to capture image information at different frequencies. Specifically, the first path uses a standard 3×3 convolutional kernel with a stride of 1 and padding of 1, focusing on extracting high-frequency fine texture features from the surface of organisms. The second path uses a 3×3 dilated convolutional kernel with a dilation rate of d=2, expanding the effective receptive field to 5×5 without increasing the number of parameters, thus capturing a wider range of contextual information, suitable for extracting the overall contour and geometric configuration of organisms. The third path uses a 1×1 pointwise convolution, mainly used for information exchange between channels and feature dimensionality reduction, acting as a bottleneck layer to reduce computational redundancy. For the input feature map... The output feature maps of the three paths are first concatenated along the channel dimension to generate intermediate joint features containing rich multi-scale information. The calculation process is expressed as follows:

[0086]

[0087] Subsequently, to adaptively select key features, the encoder introduces a channel attention mechanism to reweight the joint features. Specifically, firstly, for... Global average pooling is performed to compress the channel descriptors into a one-dimensional dimension. Then, a multilayer perceptron (MLP) is used to learn the dependencies between channels. This MLP contains two fully connected layers. The first layer compresses the number of channels to 1 / r of the input dimension (r=16 in this embodiment to balance performance and parameter count). After ReLU activation, the second layer restores the number of channels to the original dimension. Finally, normalized channel weights are generated using the sigmoid function and multiplied element-wise back. To obtain the final output The formula is as follows:

[0088]

[0089] in, This indicates a channel splicing operation. It is the Sigmoid activation function. For element-wise multiplication, For global average pooling, It is a multilayer perceptron.

[0090] While performing semantic feature extraction, the model runs the physical clarity prior module in parallel, the structure of which is as follows: Figure 5 As shown. Existing video classification models typically treat blurred frames as ordinary texture features for learning, making them highly susceptible to being misled by defocus noise. To address this issue, this embodiment innovatively introduces a physical sharpness prior module that does not participate in gradient backpropagation. It directly utilizes optical imaging principles to objectively evaluate the physical imaging quality of each frame slice from both spatial and frequency domains. In terms of spatial domain analysis, based on the principle that the image on the focal plane has the sharpest edge gradient, the Laplacian operator is used... Extract the first Frame Image The second-order differential edge information is obtained, and its variance is calculated as a spatial sharpness score. The larger the variance, the richer the high-frequency edges in the image, i.e., the clearer it is. The calculation formula is:

[0091]

[0092] in, This represents the convolution operation. For a standard 3×3 Laplacian convolution kernel, its numerical matrix is ​​defined as follows:

[0093]

[0094] In frequency domain analysis, based on the physical property that defocusing blur is equivalent to a low-pass filter causing high-frequency energy attenuation, a two-dimensional Fast Fourier Transform (FFT) is performed on each frame of the image to transform it from the spatial domain to the frequency domain. Subsequently, the high-frequency region in the logarithmic amplitude spectrum is statistically analyzed. Average energy as a frequency domain sharpness score ,in Defined as the outer region in the spectrum after removing the low-frequency components in the central 1 / 4 radius, the calculation formula is:

[0095]

[0096] in, This represents the total number of pixels in the high-frequency region. For frequency domain coordinates, This represents the result of the Fourier transform of the image. This indicates a summation operation.

[0097] Finally, in order to integrate these two physical scalars into the deep neural network, this embodiment first uses the Min-Max normalization method to... and The scores are mapped to the [0,1] interval to eliminate dimensional differences. Then, the normalized scores are concatenated and mapped through a lightweight MLP (containing one hidden layer with ReLU activation function) to a physical prior vector aligned with the semantic feature dimension. :

[0098]

[0099] Finally, the extracted semantic feature sequence and physical prior vectors are fed into the deep-aware sequence aggregation module. This is the core backend component of the model, such as... Figure 6 As shown, this module aims to combine physical priors and contextual information to fuse discrete feature sequences into a global three-dimensional representation. Considering the significant spatial continuity of microscopic tomography sequences along the Z-axis and the gradual changes in biological morphology between adjacent slices, this module first integrates the feature sequences output by the twin encoder. Input a one-dimensional temporal convolutional network (1D-CNN). This network slides along the sequence dimension, with a kernel size of 3 and padding of 1, allowing the features of each frame to perceive the contextual information of its preceding and following focal planes. This leverages inter-layer complementarity to repair information loss caused by local occlusion or artifacts, resulting in an enhanced feature sequence. .

[0100] Based on this, the aforementioned generated physical prior vectors are used As a gating signal, the attention weights for each frame are calculated. The calculation process integrates semantic features and physical priors, forcing the model to focus on keyframes that are both semantically rich and physically clear. The specific calculation formula is as follows:

[0101]

[0102]

[0103] in, and These are the projection matrices of semantic features and physical priors, respectively (implemented as 1×1 convolutions with feature map input). For the rating vector, For bias terms, The hyperbolic tangent activation function is used. For the first Unnormalized attention scores for each frame. Weight distribution for all frames is obtained through Softmax normalization. .

[0104] Finally, the enhanced feature sequences are linearly weighted and summed using the generated weights to generate a global 3D feature descriptor. :

[0105]

[0106] This global descriptor highly condenses the effective stereo information in the entire Z-Stack sequence and is then fed into a fully connected classification layer for final species category prediction.

[0107] 3. Model Training

[0108] In this embodiment, the phytoplankton tomography sequence identification method integrating physical priors is implemented on the Ubuntu 22.04 LTS operating system, using Python 3.10 as the programming language, PyTorch 2.1.0 as the deep learning framework, and CUDA 12.1 as the parallel computing acceleration library. Considering the high dimensionality of the tomographic sequence data (the input tensor shape is...), the implementation platform is designed to support this method. To meet the memory requirements of large-scale tensor operations, model training was performed on a high-performance computing workstation equipped with an NVIDIA GeForce RTX 4090 GPU (24GB of video memory).

[0109] During the experiment, to ensure that the PPSA-Net model could converge fully and avoid overfitting, various training parameters were set in detail. The specific configuration is shown in Table 2.

[0110] Table 2 Detailed Experiment Configuration

[0111]

[0112] In terms of optimization strategy, this embodiment selects the AdamW optimizer, utilizing its decoupled weight decay mechanism to improve the model's generalization ability. Simultaneously, to dynamically adjust the optimization step size, the ReduceLROnPlateau learning rate scheduling strategy is adopted. This strategy monitors the changes in classification loss on the validation set in real time, setting a patience value of 10. When the validation set loss does not show a significant decrease within 10 consecutive epochs, the current learning rate is automatically decayed to 0.5 times its original value (Factor=0.5), until the learning rate drops to the minimum threshold 1e-6. Figure 7 The diagram shows the convergence curve of the training set loss function and the upward trend of the validation set accuracy during a complete training process. It can be seen that the model reaches a convergent and stable state at around 120 rounds.

[0113] In constructing the loss function, considering the weakly supervised learning scenario used in this embodiment—that is, only having bag-level species labels for the entire tomographic sequence, but lacking fine-grained annotations for each frame slice—the model employs the Cross-Entropy Loss function as the global optimization objective function. This loss function directly applies to the global feature descriptor output by the deep-aware sequence aggregation module. Between the predicted and true labels, the parameter updates of the entire network (including the front-end encoder, physical prior module, and aggregation module) are driven by minimizing the difference between the predicted probability distribution and the true distribution.

[0114] For including The total loss for a batch of sequence samples The calculation formula is as follows:

[0115]

[0116] in, Indicates batch size, This represents the total number of phytoplankton species (in this embodiment) (18) For the first The sample belongs to the first The one-hot encoding of the class's true label means that the value is 1 when the sample belongs to that class, and 0 otherwise; The output of the model classification layer The sample belongs to the first The predicted probability value of the class (normalized by Softmax). Through this end-to-end supervision signal, the error gradient can be backpropagated to the feature extractor of each frame through the sequence aggregation module, forcing the model to automatically learn how to use physical priors to select high-value slices, thereby achieving high-precision sequence recognition without frame-by-frame annotation.

[0117] 4. Experimental Results

[0118] To fully verify the effectiveness of the proposed method, two parts of experiments were designed and completed. The first part was a comparative experiment, which evaluated the comprehensive advantages of PPSA-Net in the phytoplankton tomography sequence identification task by comparing its performance with existing mainstream image classification models and video sequence analysis models. The second part was an ablation experiment, which gradually introduced core functional modules to quantify the specific contributions of physical priors and sequence aggregation mechanisms to the model performance.

[0119] The experiment employed three classification evaluation metrics: classification accuracy (Acc), Matthews correlation coefficient (MCC), and F1 score. Classification accuracy measures the model's overall discriminative ability on the test set; the Matthews correlation coefficient comprehensively considers true positives, true negatives, false positives, and false negatives, making it particularly suitable for class imbalance scenarios common in biological data; and the F1 score is the harmonic mean of precision and recall, reflecting the model's robustness.

[0120] Comparative experiment:

[0121] To highlight the superiority of this invention in processing microscopic tomography sequence data, this embodiment constructs a comparative benchmark that includes different technical approaches and selects two representative deep learning models for performance evaluation.

[0122] First, ResNet-50, EfficientNet-B0, and ViT-B / 16 were selected as representatives of mainstream 2D ​​image classification models. In this comparative experiment, the model ignores the interlayer structure of the sequence and treats each frame slice in the tomographic sequence as an independent two-dimensional sample for feature extraction and prediction. Finally, the average of the prediction results of all frames is calculated as the classification result of the sequence. This set of experiments aims to verify the necessity of introducing interlayer correlation and stereo information.

[0123] Secondly, LSTM (Long Short-Term Memory), C3D (3D Convolutional Network), and TimeSformer were selected as representative video sequence analysis models. These models directly treat Z-Stack tomography sequences as video streams with a time dimension for spatiotemporal feature extraction. This set of experiments aims to compare and verify the significant advantages of the physical sharpness prior mechanism proposed in this invention compared to traditional black-box temporal modeling in suppressing defocus noise. The experiments were conducted on a self-built dataset of 18 types of phytoplankton tomography sequences, and the specific comparison results are shown in Table 3.

[0124] Table 3. Performance metrics comparison of different models on the tomography sequence dataset

[0125]

[0126] Analysis of the experimental data in Table 3 shows that traditional 2D models (such as ResNet-50) suffer from insufficient effective feature extraction and have the lowest overall accuracy of only 86.42% because they ignore the correlation information between slices in the tomographic sequence and are highly susceptible to defocus blurring in a single frame. In contrast, video analysis models (such as TimeSformer) can improve accuracy to over 90% by utilizing temporal information, but their parameter count is as high as 121.3M, resulting in extremely high computational costs. Furthermore, due to a lack of physical perception of "blur," they are prone to misclassifying defocus noise in the background as texture features. In stark contrast, the PPSA-Net model proposed in this invention achieves the highest classification accuracy (94.28%) with a lightweight parameter count of only 18.5M. This significant advantage is mainly attributed to the physical sharpness prior module effectively filtering out defocus background noise and the depth perception aggregation module accurately capturing key stereo features, thus achieving a dual breakthrough in recognition performance and operating efficiency with extremely low computational load.

[0127] Ablation experiment:

[0128] To verify the contribution of each innovative structural module in this invention to the overall model performance, multiple ablation experiments were designed and conducted. A twin encoder with physical priors and context aggregation removed, plus average pooling, was used as the baseline model. Focal length context modeling and physical sharpness priors were gradually added. The experimental results are shown in Table 4.

[0129] Table 4. Ablation Experiment Results of Key Modules in PPSA-Net

[0130]

[0131] Table 4 clearly reveals the contributions of each core module in the ablation experiment. First, compared to the simple average pooling strategy of the baseline model, introducing one-dimensional convolution for focal length context modeling (variant A) improved the model accuracy by approximately 2.28%, indicating that utilizing complementary information between adjacent slices to repair missing local features is highly effective. Second, introducing only the physical prior gating mechanism (variant B) brought a significant performance improvement of approximately 2.91%, fully demonstrating that explicitly calculating the Laplacian gradient and frequency domain energy can effectively guide the model to "discard false information and retain true information," accurately suppressing interference from out-of-focus frames. Finally, when both context modeling and the physical prior mechanism are enabled simultaneously (the complete model), all evaluation metrics reach their optimal levels. The two mechanisms complement each other; the former is responsible for filling in missing information, while the latter is responsible for filtering high-quality information, together constituting PPSA-Net's powerful stereo recognition capability.

[0132] 5. In-depth analysis and mechanism verification

[0133] Based on the quantitative evaluation of the above indicators, this embodiment further analyzes the internal working mechanism and engineering characteristics of the PPSA-Net model in processing microscopic tomography sequences through hyperparameter sensitivity analysis and interpretability verification.

[0134] First, regarding the key hyperparameter of sequence sampling density, this embodiment explores the impact of input sequence length on the model's classification accuracy and inference efficiency. For example... Figure 8 As shown in the biaxial line graph, the changes in classification accuracy and single inference latency on the test set as the input sequence length increases from 4 to 32 are reflected. The curves reveal significant stage-specific response characteristics in the model's performance. When the sequence length is short (e.g., 4 or 8 frames), the complete three-dimensional structure of the organism is not effectively covered due to sparse Z-axis sampling, resulting in information truncation and low accuracy. The accuracy curve peaks at 16 frames, indicating that this sampling density is sufficient to cover the effective depth of field of most phytoplankton under a microscope. However, when the sequence length further increases to 24 or 32, the accuracy curve flattens out or even slightly decreases, indicating that excessive slices do not provide additional semantic increments but rather dilute the focusing ability of the attention mechanism; simultaneously, the inference latency exhibits a strictly linear growth trend with the increase in the number of frames. Based on this analysis, this embodiment ultimately selects 16 frames as the optimal implementation parameter, achieving optimal computational efficiency while ensuring high accuracy.

[0135] Secondly, to verify the interpretability of the model's internal attention mechanism, this embodiment statistically analyzes the correlation between the attention weights generated by the model and the image physical sharpness score. The results are as follows: Figure 9 As shown in the figure, the scatter distribution of the sequence frames in the test set and their fitting trend line are illustrated. The analysis results show a significant positive correlation between attention weights and physical sharpness (Pearson correlation coefficient of 0.85). This means that in most cases, the high-weight frames learned autonomously by the model correspond precisely to physically sharp, in-focus frames. This high degree of synchronicity confirms that PPSA-Net is not a black-box model, but has successfully learned cognitive logic similar to that of human experts: that is, in the deep semantic space, it can automatically identify and lock in in-focus slices containing rich texture details, while decisively suppressing the weight allocation of out-of-focus and blurred frames, thereby achieving intelligent stereo perception that surpasses simple physical optics.

[0136] like Figure 10As shown, this invention also provides a phytoplankton chromatography sequence identification device. The device includes at least one processor and at least one memory, as well as a communication interface and an internal bus. The memory stores a computer-executable program. The memory stores a computer-executable program for the 3D identification model PPSA-Net constructed by the construction method described above. When the processor executes the computer-executable program stored in the memory, it can execute a phytoplankton chromatography sequence identification method. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus in the accompanying drawings is not limited to only one bus or one type of bus. The memory may include high-speed RAM memory, and may also include non-volatile memory (NVM), such as at least one disk storage device, or it can be a USB flash drive, portable hard drive, read-only memory, disk, or optical disk, etc.

[0137] The device may be provided as a terminal, server, or other type of device. In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0138] The present invention also provides a computer-readable storage medium storing a computer executable program of a stereo recognition model PPSA-Net constructed by the construction method described above. When the computer executable program is executed by a processor, it enables the processor to execute a phytoplankton chromatography sequence recognition method.

[0139] Specifically, a system, apparatus, or device may be provided equipped with a readable storage medium on which software program code implementing the functions of any of the embodiments described above is stored, and the computer or processor of the system, apparatus, or device reads and executes the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any of the embodiments described above, therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0140] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0141] While the specific embodiments of the present invention have been described above, they are not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for constructing a phytoplankton chromatography sequence identification model, characterized in that, The process includes the following: Step 1: Obtain several sets of phytoplankton micro-chromatographic sequence data containing multiple consecutive focal plane slices; Step 2: Perform stereo preprocessing on the tomographic sequence data to construct training and testing sets containing physical depth information; Step 3: Build a stereo recognition model PPSA-Net that integrates physical prior and sequence aggregation within a deep learning framework. The model adopts a three-level cascaded architecture: First, the front end uses a parameter-shared Siamese multi-scale feature encoder, which extracts features independently from each frame slice in the sequence through a convolutional network that shares weights in the time dimension, generating a high-dimensional semantic feature sequence. Second, a physical sharpness prior module is set up in parallel, which extracts objective imaging quality indicators from the spatial and frequency domains of the image based on the Laplacian operator and Fourier transform, generating a physical prior vector. Finally, the back end is connected to the depth-sensing sequence aggregation module, which introduces context modeling and physical gating mechanisms. The physical prior vector is used as an attention gating signal to guide the adaptive weighted fusion of semantic features in the sequence dimension, generating a global feature descriptor representing the full depth stereo morphology of the organism, which is then input into the classifier to obtain the phytoplankton category. The physical clarity prior module comprises two parallel branches: spatial gradient analysis and frequency spectrum analysis. The specific calculation process is as follows: S1, Spatial clarity score calculation; using the Laplace operator. Extract the first Frame Image The second-order differential edge information is obtained, and its variance is calculated as a spatial sharpness score. The calculation formula is: in, This represents the convolution operation. The Laplacian convolution kernel is defined as follows: S2, Frequency Domain Sharpness Score Calculation; Perform a two-dimensional Fast Fourier Transform (FFT) on the image to calculate the high-frequency region in the logarithmic amplitude spectrum. Average energy as a frequency domain sharpness score The calculation formula is: in, This represents the total number of pixels in the high-frequency region. For frequency domain coordinates, This represents the result of the Fourier transform of the image. This indicates a summation operation; S3, Physical Prior Fusion Mapping; concatenates the normalized spatial domain score with the frequency domain score, and passes it through a multilayer perceptron. Mapped to physical prior vectors aligned with semantic feature dimensions The calculation formula is: in, This indicates a normalization operation; Step 4: Based on the weakly supervised learning strategy, the PPSA-Net model is jointly trained end-to-end and its parameters are optimized using a training set containing only bag-level labels. The model with the best performance is selected as the final model.

2. The method for constructing a phytoplankton chromatography sequence identification model as described in claim 1, characterized in that: The stereoscopic preprocessing in step 2 includes field-of-view extraction, sequence normalization, and tensor reconstruction. First, field-of-view extraction based on the region of interest (ROI) is performed. A threshold segmentation algorithm is used to calculate a binary mask of the organism's outline in the original microscopic field of view, and a minimum bounding rectangle is generated based on this mask to crop out the central region containing complete biological information. Second, equally spaced sequence sampling is performed, assuming the original sequence length is... The target input length is , with step size Perform uniform sampling to construct a standardized input sequence. Finally, batch-folding-based tensor reconstruction is performed, transforming the standardized four-dimensional sequence tensors before inputting them into the network to convert the batch dimensions. With sequence length dimension Merge to generate folded tensors This allows the model to process all slices in parallel during a forward propagation process.

3. The method for constructing a phytoplankton chromatography sequence identification model as described in claim 1, characterized in that: The twin multi-scale feature encoder in step 3 consists of several stacked multi-scale sensing units for the input feature map. Each unit contains three parallel paths: The first path uses a 3×3 standard convolution kernel to extract high-frequency fine texture features; the second path uses a 3×3 dilated convolution kernel with a dilation rate of d=2 to expand the receptive field to extract large-scale geometric configurations without increasing the number of parameters; the third path uses a 1×1 point convolution to perform channel dimensionality reduction and feature interaction. The outputs of the three paths are first concatenated to generate intermediate joint features. The calculation formula is as follows: Subsequently, the joint features are adaptively weighted using a channel attention mechanism to obtain the final output. The calculation formula is as follows: in, This indicates a channel splicing operation. It is the Sigmoid activation function. For element-wise multiplication, For global average pooling, It is a multilayer perceptron; the first in the sequence Frame Image The feature extraction process satisfies ,in For shared parameter weights.

4. The method for constructing a phytoplankton chromatography sequence identification model as described in claim 1, characterized in that: The depth-sensing sequence aggregation module in step 3 introduces focal length context modeling and physical prior gating mechanisms, and its specific execution steps include: S1, Focal Context Modeling; The feature sequence output by the twin encoder... The input focal length context submodule is used to extract inter-layer complementary information using a one-dimensional temporal convolutional network, resulting in an enhanced feature sequence. ; S2, Physical Gated Weight Generation; Based on Physical Prior Vector Calculate the attention weights for each frame. The calculation formula is: in, and These are the projection matrices for semantic features and physical priors, respectively. For the rating vector, For bias terms, It is the hyperbolic tangent activation function; S3, Global Feature Weighted Fusion: The generated weights are used to linearly weight the enhanced feature sequences, generating a global 3D feature descriptor. The calculation formula is: in, This represents matrix or scalar multiplication. This indicates that for all sequences... The frames are accumulated. S4, Feature Mapping and Probabilistic Classification: A fully connected classification layer is constructed as the output head to map the global stereo feature descriptor to a high-dimensional semantic category space. The posterior probability distribution of the input sample belonging to each phytoplankton category is calculated using the Softmax normalization exponential function, and the final phytoplankton classification result is output based on the maximum probability.

5. The method for constructing a phytoplankton chromatography sequence identification model as described in claim 1, characterized in that: The weakly supervised training strategy in step 4 is to use only the bag-level species category labels to supervise the entire sequence during the model training stage, without providing additional fine-grained annotation information for the clarity or key feature positions of each frame slice in the sequence. The cross-entropy loss function is used as the objective function for global optimization. This loss function updates the network parameters by minimizing the difference between the predicted probability distribution and the true label distribution. Its mathematical expression is as follows: in, Indicates the total loss of the batch. Indicates batch size, This indicates the total number of phytoplankton species. For the first The sample belongs to the first The one-hot encoding of the class's true label is 1 when the sample belongs to that class, and 0 otherwise. The output of the model classification layer The sample belongs to the first The predicted probability value of the class.

6. The method for constructing a phytoplankton chromatography sequence identification model as described in claim 5, characterized in that: During training, a spatiotemporal consistency data augmentation strategy is introduced, which mandates that for all data in the same input sequence... Frame images, with the exact same random geometric transformation parameters applied. To maintain the spatial alignment within the sequence, the transformation process satisfies: in, to This represents the images in the sequence. This indicates specific geometric transformation operations, including random rotation and horizontal flipping.

7. A method for identifying phytoplankton chromatographic sequences, characterized in that, The process includes the following: S1, real-time acquisition of high-resolution phytoplankton microchromatographic sequence data to be identified; S2, input the microscopic chromatography sequence data of the phytoplankton to be identified into the PPSA-Net stereo recognition model constructed by the construction method described in any one of claims 1 to 6; S3 outputs the real-time phytoplankton category.

8. A phytoplankton chromatography sequence identification device, characterized in that: The device includes at least one processor and at least one memory, the processor and the memory being coupled together; the memory stores a computer-executable program of the stereo recognition model PPSA-Net constructed by the construction method as described in any one of claims 1 to 6; when the processor executes the computer-executable program stored in the memory, the processor executes a phytoplankton chromatography sequence recognition method.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer executable program for a stereo recognition model PPSA-Net constructed by the construction method as described in any one of claims 1 to 6. When the computer executable program is executed by a processor, the processor executes a phytoplankton chromatography sequence recognition method.

Citation Information

Patent Citations

  • Fine-grained phytoplankton microscopic image enhancement classification method and model building method thereof

    CN119360376A

  • Small sample plankton image enhancement recognition method and small sample plankton image model building method

    CN119992223A