A small-sample plankton image enhancement recognition method and its model construction method

Through the CD-FRN model, combined with cross-domain feature alignment, channel space gating and morphological perception modules, the problems of insufficient cross-domain generalization capabilities and fine-grained recognition in plankton image recognition are solved, and high-precision recognition is achieved in complex marine environments.

CN119992223BActive Publication Date: 2025-07-04OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457298.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-04
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Traditional deep learning methods have insufficient cross-domain generalization capabilities and fine-grained recognition problems in plankton image recognition, especially in the case of limited sample size and complex marine environments, it is difficult to achieve high-precision and robust automatic recognition.

Method used

The CD-FRN model based on feature map reconstruction network is adopted, combining cross-domain feature alignment, channel space gating and morphological perception modules, feature maps are extracted through deep convolutional neural networks, and trained under the meta-learning framework. The cross-domain feature alignment strategy is used to reduce the differences between domains, the channel space gating module enhances adaptability, and the morphological perception module extracts multi-scale structure information, and finally classify the feature map reconstruction.

Benefits of technology

The model's cross-domain adaptability and fine-grained classification capabilities are improved, and high-precision recognition of plankton images in complex marine environments is achieved, which enhances the robustness and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992223B_ABST
    Figure CN119992223B_ABST
Patent Text Reader

Abstract

The present invention provides a method for enhancing and recognizing small-sample plankton images and a method for building its model, belonging to the technical field of underwater image enhancement. First, plankton microscopic image data is acquired and sorted out to construct a standardized data set. Then, a small-sample plankton image enhancement and recognition model based on feature map reconstruction technology is designed. A deep convolutional neural network is used to extract features from the image, generate intermediate layer feature maps, and combine the feature map reconstruction mechanism to achieve class discrimination. The model integrates multiple key designs for cross-domain adaptation and structure perception, alleviates the distribution differences between different data domains, and also significantly improves the model's ability to represent the complex morphological structures of plankton. Finally, the model is trained and optimized end-to-end under the meta-learning framework to obtain the best model. Experimental results show that this method exhibits excellent classification performance on multiple plankton data sets and can achieve high-precision recognition under the condition of limited sample quantity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater image enhancement, and particularly relates to a method for enhancing and recognizing small-sample plankton images and a method for building a model thereof. Background Art

[0002] Plankton are key biological groups in aquatic ecosystems. They not only form the basis of the food chain but also play a crucial role in water quality assessment and marine resource management. Achieving the automatic recognition and monitoring of plankton helps to scientifically evaluate the ecological status of water bodies, predict environmental evolution trends, and provide data support and decision-making basis for fishery development and ecological protection.

[0003] However, in actual operation, obtaining plankton images faces many challenges. On the one hand, the acquisition process is limited by high equipment costs, complex and variable marine environments, and a high dependence on operation skills, resulting in a limited number of available samples; it is more difficult to construct a large-scale dataset for rare or species in specific sea areas. On the other hand, plankton have diverse morphologies and subtle intergeneric differences, posing higher requirements for the representational ability of the model; if the training data is insufficient, traditional deep learning methods are prone to overfitting, resulting in unstable recognition and difficulty in meeting the requirements of high precision and robustness in practical applications.

[0004] In recent years, a series of few-shot learning methods have emerged in the field of computer vision. These methods train the network under the meta-learning framework, enabling the model to still have the ability of rapid generalization and accurate classification when facing new categories with extremely few samples. Such methods have shown good application prospects in typical data-scarce scenarios such as medical image diagnosis and remote sensing image analysis, and have important research value and wide practical significance. Among them, the Feature Map Reconstruction Network (FRN) has attracted much attention for transforming the few-shot classification task into a feature reconstruction problem. Compared with traditional metric learning or meta-learning methods, FRN can fully retain the spatial structure information of the image and achieve efficient reconstruction of features through closed-form linear regression, thereby reducing the number of model parameters while improving the inference efficiency. This method has achieved excellent performance in multiple typical few-shot learning tasks, reaching or approaching the current state-of-the-art level, showing good generality and application potential.

[0005] Although FRN performs well in few-shot learning tasks, its cross-domain adaptation ability may still be limited in the task of plankton microscopic image recognition. Plankton images are affected by various factors such as morphology, background, and acquisition conditions. The source waters are diverse, and the visual features vary significantly, which easily leads to a decline in the generalization ability of the model when migrating to new scenarios. In addition, there are a large number of plankton species, and the morphological differences between some species are extremely small, which poses higher requirements for the fine-grained recognition ability of the model. Therefore, improving the cross-domain adaptability and fine-grained classification ability of the model in specific marine ecological scenarios has important research value and is worthy of in-depth exploration and application. Summary of the Invention

[0006] In view of the above problems, the first aspect of the present invention provides a method for enhancing the recognition of plankton images with few samples and a method for building its model, including the following processes:

[0007] Step 1, obtain a dataset of plankton microscopic images;

[0008] Step 2, preprocess the original image data and divide it into a base class and a new class, with no class overlap between the two;

[0009] Step 3, construct a few-shot image enhancement recognition model CD-FRN based on the feature map reconstruction network FRN; the model first processes the input image through a deep convolutional neural network to extract its intermediate layer feature map; the intermediate layer feature map is processed by a cross-domain feature alignment strategy, and then sequentially input into a channel-spatial gating module and a morphology-aware module for feature enhancement. The cross-domain feature alignment strategy first calculates the statistics of the feature maps in the source domain and the target domain in the spatial dimension, performs normalization and affine alignment processing on the feature maps according to the statistical information, and then performs non-linear calibration on the alignment result through an adjustable power-law transformation; the calibrated feature map is then input into the channel-spatial gating module, adjusts the feature response through the joint modeling of channel attention and spatial attention, and introduces a gating mechanism to fuse the calibrated feature and the domain-enhanced feature to generate a fused feature map; the fusion result is further input into the morphology-aware module, extracts the structural information through multiple convolutional branches of different scales, and fuses the extracted multi-scale features and the fused feature through a residual connection to generate an enhanced feature map. The enhanced feature map is fed into the feature map reconstruction process. The model constructs a reconstruction support matrix based on the support set in each few-shot task, calculates the linear mapping result of the query feature using closed-form ridge regression, and performs class discrimination based on the difference between the mapping result and the original query feature;

[0010] Step 4, training and validating the model under the meta-learning framework: In the training stage, several small-sample tasks are generated from the base classes and divided into a support set and a query set. The model is optimized in an end-to-end manner to minimize the value of the joint loss function, which is composed of a main classification loss function constructed based on the reconstruction error and a CMA loss function based on channel statistics alignment; in the validation stage, tasks are constructed from new classes to evaluate the model performance, and the model parameters with the optimal validation effect are saved for subsequent applications.

[0011] Preferably, a cross-domain feature alignment strategy is adopted in the CD-FRN model in step 3, and domain-invariant feature extraction is achieved through dynamic normalization and feature calibration. The specific implementation includes the following steps:

[0012] S1, feature normalization and domain alignment. The input feature map is normalized in the spatial dimension to suppress the visual style differences between different domains. Let the source domain feature map be , where represents the batch size, represents the number of channels, and respectively represent the height and width of the feature map. Calculate the spatial dimension mean and standard deviation of the feature map in the spatial dimension:

[0013]

[0014]

[0015] The normalized feature is:

[0016]

[0017] where, and are respectively the mean and standard deviation of the source domain feature, is a constant to avoid division by zero error, is the normalized source domain feature map;

[0018] If the target domain feature is available, the statistics of the target domain feature are calculated synchronously:

[0019]

[0020]

[0021]

[0022] where, and are respectively the mean and standard deviation of the target domain feature, is the target domain feature map after standardization; then, the standardized source domain feature map is affine-transformed to the target domain distribution space to obtain the aligned source domain feature map :

[0023]

[0024] S2, Feature response calibration. Through a learnable exponential factor apply a sign-preserving power-law non-linear transformation to the aligned feature map to obtain the calibrated feature :

[0025]

[0026] wherein, the initial value of is 0.95, which is updated through backpropagation during training and constrained to the interval (0,1) by the Sigmoid function to control the transformation intensity. The calibrated feature is passed on as an intermediate feature for subsequent fusion and enhancement operations.

[0027] Preferably, a channel-spatial gating module is adopted in the CD-FRN model in step 3 to enhance the adaptation ability of the model between different domains. Specifically, the implementation of the channel-spatial gating module includes the following steps:

[0028] S1, Channel attention calculation. Perform global average pooling operation on the input feature map to generate a feature map with a shape of . Then, use two 1×1 convolutional layers to transform the pooled feature: the first convolutional layer compresses the number of channels to 1 / 8 of the original number of channels and performs non-linear mapping through the ReLU activation function; the second convolutional layer restores the channel dimension to the original size and generates channel attention weights with a shape of through the Sigmoid function;

[0029] S2, Spatial attention calculation. Apply a 1×1 convolutional layer to the input feature map to generate a spatial attention map with a size of and generate weights through the Sigmoid function. This spatial attention map is used to weight each spatial position of the input feature map;

[0030] S3, Feature weighted fusion. Expand the channel attention and spatial attention to the same dimension respectively and then perform element-wise multiplication to obtain the joint attention weight . Subsequently, input into a 1×1 convolutional layer initialized as an identity matrix to generate a domain-specific response feature map Apply the joint attention weight to the domain feature map to obtain the modulated feature response To achieve the adaptive fusion of the calibrated feature and the domain-enhanced feature, a learnable scalar parameter is introduced and a gating factor is generated through the Sigmoid function as the fusion weight. The fusion operation is expressed as follows:

[0031]

[0032] where is the fused feature. Through the joint attention mechanism and the gating adjustment structure, this module realizes the effective modeling and weighted integration of cross-domain features

[0033] Preferably, a morphological perception module is introduced in the CD-FRN model in step 3. Different-scale morphological features are extracted through a multi-scale convolution structure to adapt to the characteristics of subtle inter-class differences and complex structures of plankton. The morphological perception module includes three parallel convolution branches, which use 1×1, 3×3, and 5×5 convolution kernels to process the feature map to obtain the image structure information at the local, basic, and wide-domain scales. To dynamically fuse the responses of different convolution branches, the module introduces a learnable weight vector and normalizes it through the softmax function to generate the proportional weight. Denote the output feature maps as , , , and the fused feature map is obtained as follows:

[0034]

[0035] To enhance the morphological response while retaining the original structure information, the module adds the input feature map and the fused feature through a residual connection to obtain the final output of this module:

[0036]

[0037] where is a trainable residual adjustment factor used to balance the fusion degree between the original input and the morphology-enhanced feature. The output feature map will be used as the input in the feature reconstruction stage

[0038] Preferably, the CD-FRN model in step 3 linearly reconstructs the features of the query image using a feature map reconstruction network. Specifically, the intermediate layer feature map output by the feature enhancement stage is input into the reconstruction module, where , , , are the batch size, the number of channels, the height and width of the feature map respectively; subsequently, a flattening operation is performed on the feature map in the spatial dimension to convert it into a two-dimensional matrix , where represents the number of spatial positions, represents the feature dimension and serves as the input for subsequent reconstruction operations.

[0039] After the support set and query set are partitioned, feature vectors of images corresponding to each category are extracted from the feature matrix to construct a support set feature matrix and a query image feature matrix respectively. Specifically, for each category , stack the feature vectors of its corresponding support images to form the support feature matrix ; similarly, the flattened feature map of the query image results in the query feature matrix .

[0040] For linear reconstruction of , introduce the reconstruction matrix , and the specific process is as follows:

[0041] S1, calculate the autocorrelation matrix of the support features :

[0042]

[0043] S2, calculate the adaptive regularization factor based on the condition number of the support set feature matrix , and its expression is:

[0044]

[0045] where represents the condition number of the matrix , is the learnable benchmark regularization parameter; multiply by the identity matrix and add it to the autocorrelation matrix to obtain the regularization matrix :

[0046]

[0047] S3, construct the reconstruction matrix :

[0048]

[0049] S4, through Perform linear reconstruction on the query features to obtain:

[0050]

[0051] Among them, is the query feature representation reconstructed based on the category Reconstructed.

[0052] Finally, use the mean square error between the reconstructed feature and the original query feature as the classification basis, and define the reconstruction score as:

[0053]

[0054] Among them, represents the Frobenius norm, represents the score that the query image belongs to the category . The smaller the reconstruction error, the higher the fitting degree of the category to this query sample. The model takes the category with the largest score as the final prediction result.

[0055] Preferably, the CMA (Covariance-Mean Alignment) loss function in step 4 acts on the aligned source domain feature map and the standardized target domain feature map , and reduces the distribution difference between domains by jointly aligning the first-order statistic (mean) and the second-order statistic (covariance). The specific process is as follows:

[0056] S1. Calculate the mean of the feature map in the batch and spatial dimensions (i.e., ), to obtain the channel mean vectors and , and define the mean alignment loss with the mean square error:

[0057]

[0058]

[0059]

[0060] S2. Flatten the feature map into a two-dimensional matrix and then calculate the covariance matrices and , and define the covariance alignment loss :

[0061]

[0062]

[0063]

[0064] S3. Combine the mean and covariance alignment losses in a weighted manner to construct the CMA loss function. :

[0065]

[0066] Wherein, and are adjustable hyperparameters, with default values of 0.1 and 0.05 respectively, and are used to regulate the relative importance of mean alignment and covariance alignment in the loss. The CMA loss function and the main classification loss function together constitute the joint optimization objective in the training process.

[0067] The second aspect of the present invention provides a method for enhancing the recognition of small-sample plankton images, including the following processes:

[0068] S1. Obtain high-resolution microscopic images of plankton to be recognized in real time;

[0069] S2. Input the microscopic images of plankton to be recognized into the small-sample plankton image enhancement recognition model built by the building method described in the first aspect;

[0070] S3. Output real-time classification results.

[0071] The third aspect of the present invention provides a small-sample plankton image enhancement recognition device, which includes at least one processor and at least one memory, and the processor and the memory are coupled; the memory stores a computer execution program of the small-sample plankton image enhancement recognition model built by the building method described in the first aspect; when the processor executes the computer execution program stored in the memory, the processor can execute a method for enhancing the recognition of small-sample plankton images.

[0072] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer execution program of the small-sample plankton image enhancement recognition model built by the building method described in the first aspect, and when the computer execution program is executed by a processor, the processor can execute a method for enhancing the recognition of small-sample plankton images.

[0073] Compared with the prior art, the present invention has the following beneficial effects:

[0074] In view of the problems of insufficient cross - domain generalization ability and high difficulty in morphological discrimination faced by traditional few - shot learning methods in plankton image recognition, this invention proposes a few - shot plankton image enhanced recognition model CD - FRN. Based on the efficient modeling advantage of the feature reconstruction mechanism, this model further enhances the adaptability to cross - domain distribution differences and improves the modeling accuracy of complex morphological features.

[0075] First of all, the model introduces a cross - domain feature alignment strategy. By standardizing and affine - mapping the feature maps of the source domain and the target domain in the spatial dimension, the distribution differences under different sea areas or acquisition conditions are reduced. On this basis, a channel - spatial gating module is introduced, which fuses channel attention, spatial attention and domain - specific response information, and realizes the adaptive fusion of calibrated features and domain - enhanced features through a learnable gating mechanism, improving the model's perception ability and response ability to cross - domain feature differences. To further enhance the alignment effect, a CMA loss function is designed to constrain the mean and covariance of the source - domain and target - domain features in the channel dimension, improving the distribution consistency and cross - domain stability during the training process.

[0076] In addition, aiming at the problems of subtle morphological differences and complex texture changes in plankton images, the model introduces a morphology - aware module. It adopts a parallel multi - scale convolution structure to extract local - to - wide - area structural information, and combines a learnable fusion coefficient and a residual connection mechanism to achieve dynamic modeling and enhanced expression of multi - scale morphological features, thus improving the model's fine - grained recognition ability.

[0077] In summary, this invention provides a few - shot image modeling method with both cross - domain adaptability and fine - grained recognition ability, which can effectively address practical problems such as complex plankton image acquisition conditions, limited samples, and high similarity between morphological classes, and has significant research value. Brief Description of the Drawings

[0078] Figure 1 It is a schematic flow chart of the few - shot plankton image enhanced recognition method proposed by this invention.

[0079] Figure 2 It is a schematic example diagram of plankton microscopic images.

[0080] Figure 3 It is a schematic working flow chart of the few - shot plankton image enhanced recognition model CD - FRN proposed by this invention.

[0081] Figure 4 It is a schematic structure diagram of the feature extraction network ResNet - 12 adopted by this invention.

[0082] Figure 5 It is a schematic structure diagram of the channel - spatial gating module.

[0083] Figure 6 It is a schematic structural diagram of the morphological perception module.

[0084] Figure 7 It is a schematic flow diagram of feature map reconstruction.

[0085] Figure 8 It is a comparison chart of the recognition results of the CD-FRN model and the benchmark model on test samples.

[0086] Figure 9 It is a schematic diagram of the simple structure of the small-sample plankton image enhancement recognition device. Specific implementation mode

[0087] The invention will be further described below in conjunction with specific embodiments.

[0088] In this embodiment, through a specific experimental process, the method proposed by the present invention is further described, and the overall process is as Figure 1 shown.

[0089] 1. Acquisition and preprocessing of original data

[0090] The image data used in this embodiment is sourced from the publicly released WHOI (Woods Hole Oceanographic Institution) plankton dataset and the Kaggle plankton dataset. The above datasets cover a large number of representative plankton species and are widely used in related tasks such as plankton image recognition and classification. Relevant sample examples are as Figure 2 shown.

[0091] To ensure data quality, the original image samples are first screened to eliminate images with quality defects such as missing, blurred, or severely occluded parts. After completing the preliminary quality verification, the images are standardized to unify the data format and distribution; subsequently, multiple data augmentation operations are performed to enrich the diversity of training samples. Specifically, the augmentation operations include: performing a random cropping operation on the image, with the cropping ratio ranging from 80% to 100% of the original image size; applying a random rotation within the range of ±40 degrees to the image; performing a random horizontal flipping operation on the image with a probability of 50%; to further adapt to the differences in illumination and color caused by different imaging devices and marine environments, the color features of the image are perturbed, where the brightness and contrast vary within the range of ±0.2, and the saturation varies within the range of ±0.3; this embodiment also introduces a random occlusion strategy, generating 1 to 3 rectangular occlusion regions at random positions in each image, with the area of each occlusion region accounting for 1% to 5% of the original image, and the aspect ratio controlled between 0.5 and 2.0, and the occlusion regions are filled with Gaussian random noise with a mean of 0.5 and a standard deviation of 0.1. All the above augmentation strategies are integrated and applied in the training stage of this embodiment, and each operation can also be flexibly combined and adjusted as needed.

[0092] 2. Model construction

[0093] Based on the above dataset, this embodiment designs and constructs a deep learning model, Cross Domain-Feature Reconstruction Network (hereinafter referred to as CD-FRN), and its overall processing flow is as Figure 3 shown. The model uses a lightweight deep convolutional neural network as the feature extractor, preferably using ResNet-12 to achieve better feature extraction ability, and also supports using the simpler-structured Conv-4 network as an alternative. This embodiment uses ResNet-12 as the feature extraction network, and its specific structure is as Figure 4 shown. The intermediate features extracted by the backbone network are then subjected to various feature enhancement processes to strengthen the cross-domain adaptability of the model and the representation ability of fine-grained morphological features, so as to still achieve high-precision recognition of plankton microscopic images under the conditions of limited samples and domain differences in data distribution.

[0094] First, the model adopts a cross-domain feature alignment strategy to reduce the distribution differences between different data domains. Specifically, let the source domain feature map be , where , , , are the batch size, number of channels, height, and width of the feature map respectively, and the same symbols in the following text represent the same meanings. First, calculate the mean and the standard deviation , and calculate the standardized source domain feature map :

[0095]

[0096]

[0097]

[0098] Among them, is a constant term used to avoid numerical instability caused by a zero standard deviation. If the target domain feature is available, its statistics are calculated synchronously, including the spatial dimension mean , standard deviation , and the standardized target domain feature map , and its definition form is the same as that of the source domain:

[0099]

[0100]

[0101]

[0102] The target domain feature map here can be obtained in both the training and validation phases: in the training phase, it comes from the training samples of the target domain and is used to assist in the statistical alignment of the source domain features; in the validation phase, it is extracted from the support set images in the new class task and is used to simulate cross-domain adaptation in the real test scenario.

[0103] Next, map the standardized source domain feature map to the target domain distribution space to achieve affine alignment and obtain the aligned source domain feature map :

[0104]

[0105] The model introduces a learnable power-law calibration mechanism to apply a sign-preserving power-law transformation to the aligned features to obtain the calibrated features :

[0106]

[0107] Among them, is a learnable parameter, with an initial value set to 0.95, and is constrained to the interval (0, 1) through the Sigmoid function to improve training stability. This mechanism balances noise suppression and key feature enhancement by dynamically adjusting the non-linear scaling intensity of the feature amplitude.

[0108] Next, the calibrated features are fed into the channel-spatial gating module, which dynamically adjusts the feature response intensity using the joint attention mechanism in both the channel and spatial dimensions, thereby enhancing cross-domain adaptability. The structure of this module is as shown in Figure 5 . First, the channel attention weights are calculated for the input feature map. Specifically, a global average pooling operation is first performed on the input feature map to generate a channel description vector with a shape of . Subsequently, this vector passes through two 1×1 convolutional layers in sequence for inter-channel information interaction: the first layer compresses the number of channels to 1 / 8 of the original number of channels and realizes non-linear mapping through ReLU activation; the second layer restores the channel dimension to the original size and generates a channel attention weight with a shape of via the Sigmoid activation function. At the same time, a 1×1 convolutional operation is applied to the input feature map and activated through the Sigmoid function to generate a spatial attention map with a shape of . Then, the channel attention weight and the spatial attention map are respectively extended to the same dimension as the input feature Figure 1 through the broadcast mechanism, and element-wise multiplication is performed at the element level to obtain the joint attention weight .

[0109] Next, a linear projection is performed on the input feature map to generate a domain-specific response feature map , and this projection operation is implemented by a 1×1 convolutional layer initialized as the identity matrix. The above joint attention weight is applied to the domain feature map to obtain the modulated feature response , where the multiplication is an element-wise weighted operation.

[0110] Subsequently, a learnable scalar gating parameter is introduced, and a gating factor is generated through the Sigmoid function to control the fusion ratio between the calibrated feature and the domain-enhanced feature. To achieve balanced fusion of features at the initial stage, is set to 0 initially, and the corresponding is 0.5 initially. The final fusion result is expressed as follows:

[0111]

[0112] where, is the fused feature. This structure realizes dynamic balance adjustment of different feature sources through joint attention modeling and the gating mechanism.

[0113] To enhance the model's ability to perceive the structure and morphological details of plankton, a morphological perception module is introduced, and its structure is as shown inFigure 6 As shown. This module consists of three parallel convolutional branches, using 1×1, 3×3, and 5×5 respectively, to introduce multi-scale receptive fields. Each branch performs local, basic, and wide-area scale structural modeling on the input feature map and outputs feature maps corresponding to the respective scales , , . To achieve dynamic fusion of multi-scale features, the model introduces a learnable fusion weight vector , whose initial value is [0.33, 0.34, 0.33], and is normalized through the softmax function to generate proportional weights. The fused feature map is calculated by the following formula:

[0114]

[0115] To enhance the morphological response while retaining the original structural information, the model adds the input feature map and the fused feature through a residual connection to obtain the final output of this module :

[0116]

[0117] where is a trainable residual adjustment factor used to balance the fusion degree between the original input and the morphologically enhanced features.

[0118] After the above enhancement process, the final features are fed into the feature map reconstruction network for class discrimination. The basic principle of this reconstruction process is as Figure 7 shown. For each image, the feature map after being processed by the aforementioned module is represented as , and after being flattened, it obtains a two-dimensional matrix representation , where represents the number of spatial positions, represents the feature dimension.

[0119] In each few-shot classification task, the feature representations of the support images and query images corresponding to each category are extracted from the feature matrix . Specifically, for each category , the flattened feature vectors of its corresponding support images are concatenated in sequence to form a support feature matrix ; the feature map of the query image is also flattened to obtain a query feature matrix . Subsequently, the model constructs a reconstruction mapping weight based on the closed-form ridge regression solution and linearly reconstructs the query features. The calculation process is as follows:

[0120]

[0121]

[0122]

[0123]

[0124]

[0125] Among them, represents the autocorrelation matrix of the support set features, represents the condition number of is a learnable baseline regularization parameter, is a class-related adaptive regularization factor, is the identity matrix, is the regularized matrix, is based on the class reconstructed query feature representation. Finally, the model makes class discrimination based on the reconstruction error, and the reconstruction score is defined as:

[0126]

[0127] Among them, represents the Frobenius norm, represents the score that the query image belongs to the class . Calculate in parallel for all candidate classes, and select the class with the maximum score as the prediction result.

[0128] 3. Model Training

[0129] In this embodiment, the implementation environment of a small-sample plankton image enhancement recognition method is based on the Ubuntu22.04 operating system, the programming language is Python 3.10, the deep learning framework uses Pytorch 2.1.0, and CUDA12.1 is used. The model was trained for 1200 rounds on a system equipped with an NVIDIA GeForce RTX 4090 GPU (24GB of video memory). The relevant settings of each training parameter were also recorded in detail during the experiment, and the specific content is shown in Table 1.

[0130] Table 1 Detailed Information of Experimental Configuration

[0131]

[0132] The learning rate scheduling adopts the ReduceLROnPlateau strategy, monitors the change of the validation loss, and when the loss does not decrease for 20 consecutive rounds, the learning rate is decayed to 0.1 times the current value until the learning rate is not lower than 1e-6.

[0133] In this embodiment, a CMA (Covariance-Mean Alignment) loss function is designed to constrain the mean and covariance matrix on the channel dimension respectively. For the aligned source domain feature map and the target domain normalized feature map , the mean vectors and are calculated, and the mean alignment loss is defined as:

[0134]

[0135]

[0136]

[0137] Then, the feature map is flattened into a two-dimensional matrix with a shape of , the covariance matrices and are calculated, and the covariance alignment loss is defined as:

[0138]

[0139]

[0140]

[0141] Finally, the CMA loss function is weighted and fused in the following form to obtain the CMA loss :

[0142]

[0143] Among them, and are loss weighting coefficients (by default =0.1, =0.05), which are used to balance the importance of mean and covariance alignment.

[0144] To construct a complete training objective function, the model comprehensively adopts a joint optimization strategy of reconstruction classification loss and CMA loss. Among them, the reconstruction classification loss is based on the feature map reconstruction error and is used to measure the discriminative ability of the model for query images; the CMA loss It is used to improve the cross - domain generalization performance of the model. The two together constitute the final training objective:

[0145]

[0146] By minimizing this joint loss function That is, the model simultaneously optimizes the classification accuracy and cross - domain alignment ability, thereby achieving the consistency of feature distribution while learning discriminative features, and further improving the recognition accuracy and robustness of the model under the few - shot condition.

[0147] 4. Experimental Results

[0148] To verify the effectiveness and superiority of this method, comparative experiments and ablation experiments were carried out in this embodiment. First, this method was compared with existing few - shot learning methods to evaluate its performance in the plankton image recognition task; second, multiple groups of ablation experiments were constructed, gradually introducing each core module to quantify its impact on the recognition effect.

[0149] The classification accuracy (Accuracy) is used as the core evaluation index of this experiment to measure the accuracy of the model in classifying target samples. In this embodiment, for the few - shot plankton recognition task, the classification accuracy was evaluated under the 5 - way 1 - shot and 5 - way 5 - shot settings respectively.

[0150] Comparative Experiments:

[0151] This embodiment selected a variety of excellent deep - learning models in the field of few - shot learning as comparison models, including Matching Nets, Prototypical Nets, MAML, etc. The specific comparative experiment results are shown in Table 2.

[0152] Table 2 Comparison of 5 - way classification accuracies of different few - shot learning methods

[0153]

[0154] The experimental results show that the present invention is superior to the existing methods under different settings of the WHOI - Plankton and Kaggle Plankton datasets, fully demonstrating its performance advantages in complex tasks such as sample scarcity, domain shift, and fine - grained structure recognition.

[0155] Ablation Experiments:

[0156] To verify the contribution of each structural module in the present invention to the overall model performance, multiple groups of ablation experiments were designed and carried out. Key functional modules were gradually introduced to evaluate their gain effects in the small-sample plankton image recognition task. The results are shown in Table 3. Among them, the "cross-domain adaptation mechanism" includes a cross-domain feature alignment strategy, a channel spatial gating module, and a CMA loss function. The three are related in the model structure and training process, so it is more valuable to conduct ablation verification as a whole.

[0157] Table 3 Results of Ablation Experiments

[0158]

[0159] The results of the ablation experiments show that each of the above components has contributed to the performance gain of the present invention.

[0160] 5. Result Visualization

[0161] To further verify the recognition performance and visual discrimination ability of the proposed method, in this embodiment, a visual analysis was carried out on the prediction results of several plankton images in the test set. Figure 8 The classification result comparison between the baseline model FRN and the model CD-FRN proposed in the present invention on the same sample is shown. Above each subfigure in the figure, the true category (GroundTruth), the model prediction result (Prediction), and the corresponding prediction confidence (Confidence) are marked. It can be seen from the figure that the CD-FRN model shows stronger recognition ability and higher prediction confidence on plankton images with complex morphological structures, further reflecting its advantages in fine-grained feature modeling and cross-domain generalization.

[0162] As Figure 9As shown, the present invention also provides a small-sample plankton image enhancement and recognition device, which includes at least one processor and at least one memory, and also includes a communication interface and an internal bus; a computer execution program is stored in the memory; a computer execution program of the small-sample plankton image enhancement and recognition model constructed by the above-mentioned construction method is stored in the memory; when the processor executes the computer execution program stored in the memory, the processor can execute a small-sample plankton image enhancement and recognition method. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus. The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a disk, or an optical disc, etc.

[0163] The device can be provided as a terminal, a server or other forms of devices. In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to execute the above method.

[0164] The present invention also provides a computer-readable storage medium, in which a computer execution program of the small-sample plankton image enhancement and recognition model constructed by the above-mentioned construction method is stored. When the computer execution program is executed by a processor, the processor can execute a small-sample plankton image enhancement and recognition method.

[0165] Specifically, a system, a device or an equipment equipped with a readable storage medium can be provided. A software program code for implementing the functions of any one of the above embodiments is stored on the readable storage medium, and the computer or processor of the system, the device or the equipment is made to read and execute the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0166] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

[0167] Although the specific implementation manners of the present invention are described above, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A method for building a small-sample plankton image enhancement and recognition model, characterized in that, It includes the following processes: Step 1: Obtain a microscopic image dataset of plankton; Step 2: Preprocess the original image data and divide it into a base class and a new class, with no class overlap between the two; Step 3: Construct a few-shot image enhancement recognition model CD-FRN based on the feature map reconstruction network FRN; The model first receives the input image and extracts the intermediate layer feature map through a deep convolutional neural network; Subsequently, multiple feature enhancement mechanisms are introduced, including a cross-domain feature alignment strategy for performing distribution standardization and calibration, a channel-space gating module for enhancing domain-specific responses, and a morphology-aware module for modeling multi-scale morphological information; The enhanced feature map is fed into the feature map reconstruction process. The model constructs a reconstruction support matrix based on the support set in each few-shot task, uses closed-form ridge regression to achieve linear fitting of the query features, and uses the reconstruction error as the basis for class discrimination; Step 4: Train and validate the CD-FRN model under the meta-learning framework to obtain the optimal model parameters; The loss function consists of a main classification loss function constructed based on the reconstruction error and a CMA loss function based on channel statistic alignment.

2. A method for building a few-shot plankton image enhancement recognition model according to claim 1, characterized in that: The CD-FRN model first processes the input image through a deep convolutional neural network to extract its intermediate layer feature map; The intermediate layer feature map is processed by the cross-domain feature alignment strategy and then sequentially input into the channel-space gating module and the morphology-aware module for feature enhancement; The cross-domain feature alignment strategy first calculates the statistics of the source domain and target domain feature maps in the spatial dimension, performs standardization and affine alignment processing on the feature maps according to the statistic information, and then performs non-linear calibration on the alignment result through an adjustable power-law transformation; The calibrated feature map is then input into the channel-space gating module, adjusts the feature response through the joint modeling of channel attention and spatial attention, and introduces a gating mechanism to fuse the calibrated feature and the domain-enhanced feature to generate a fused feature map; The fusion result is further input into the morphology-aware module, extracts the structural information through multiple convolutional branches of different scales, and fuses the extracted multi-scale features and the fused features through a residual connection to generate an enhanced feature map; The enhanced feature map is fed into the feature map reconstruction process. The model constructs a reconstruction support matrix based on the support set in each few-shot task, calculates the linear mapping result of the query features using closed-form ridge regression, and performs class discrimination based on the difference between the mapping result and the original query features.

3. A method for building a few-shot plankton image enhancement recognition model according to claim 1, characterized in that: The cross-domain feature alignment strategy realizes domain-invariant feature extraction through dynamic standardization and feature calibration, specifically including: Feature Standardization and Domain Alignment: Standardize the input feature map in the spatial dimension to suppress the visual style differences between different domains; Denote the source domain feature map as , where represents the batch size, represents the number of channels, and represent the height and width of the feature map respectively; First, calculate the channel mean and standard deviation of the feature map in the two spatial dimensions of height and width respectively to obtain the statistics and of the source domain features, and complete the standardization process based on this to obtain the standardized source domain feature map: Among them, and are the mean and standard deviation of the source domain features respectively, is the source domain feature map after standardization; If the target domain features are available, synchronously adopt the same method to process the target domain feature map to calculate its mean value in the spatial dimension and standard deviation , and obtain the normalization result ; Then, the standardized source domain feature map is affine-transformed to the target domain distribution space to obtain the aligned source domain feature map : Feature response calibration: Through a learnable exponential factor Apply a sign-preserving power-law non-linear transformation to the aligned feature map to obtain the calibrated feature map : Among them, has an initial value of 0.95, is updated through backpropagation during training, and is constrained to the interval (0, 1) by the Sigmoid function to control the transformation intensity; The calibrated feature is passed on as an intermediate feature.

4. A method for building a few-shot plankton image enhancement recognition model according to claim 1, characterized in that: The channel-space gating module is used to enhance the model's adaptability between different domains; specifically including: S1, Channel attention calculation; For the input feature map perform global average pooling operation to generate a feature map with a shape of ; Then, use two 1×1 convolutional layers to transform the pooled features: the first convolutional layer compresses the number of channels to 1 / 8 of the original number of channels and performs non-linear mapping through the ReLU activation function; the second convolutional layer restores the channel dimension to the original size and generates channel attention weights with a shape of ; S2, Spatial attention calculation; for the same input feature map Apply a 1×1 convolutional layer to generate a spatial attention map of size and generate weights through the Sigmoid function. This spatial attention map is used to weight each spatial position of the input feature map; S3, Feature weighted fusion: Extend the above channel attention and spatial attention to the same dimension and multiply them element by element to generate joint attention weights ; Meanwhile, input into a 1×1 convolutional layer initialized as an identity matrix to generate domain-specific response feature maps ; By applying the joint attention weights to the feature maps , obtain the modulated feature responses ; To achieve the adaptive fusion between calibration features and the modulation response, a learnable scalar parameter is introduced, and a gating factor is generated through the Sigmoid function as the fusion ratio coefficient; the final output feature is composed of the weighted sum of two parts: one part retains the original information of the calibration feature, and the other part fuses the response feature enhanced by the domain-specific, and the ratio between the two is automatically adjusted by the gating factor.

5. A method for building a small-sample plankton image enhancement and recognition model according to claim 1, characterized in that: The morphological perception module extracts morphological features at different scales through a multi-scale convolutional structure to adapt to the characteristics of subtle inter-class differences and complex structures of plankton; the morphological perception module includes three parallel convolutional branches, which use convolutional kernels of 1×1, 3×3, and 5×5 respectively to process the feature map for obtaining image structure information at local, basic, and wide-field scales; to dynamically fuse the responses of different convolutional branches, the module introduces a learnable weight vector , and normalizes it through the softmax function to generate proportional weights; denote the output feature maps as , , , and obtain the fused feature map : To enhance the morphological response while preserving the original structural information, the module adds the input feature map and the fused features through a residual connection to obtain the final output of the module : Among them, is a trainable residual adjustment factor for balancing the fusion degree between the original input and the morphological enhancement features; the output feature map will be used as the input for the feature reconstruction stage.

6. A method for building a small-sample plankton image enhancement and recognition model according to claim 1, characterized in that: The CMA loss function acts on the aligned source domain feature map and the normalized target domain feature map , and reduces the distribution difference between domains by jointly aligning the first-order statistics and the second-order statistics. Specifically: S1. Calculate the mean value of the feature map in the batch and spatial dimensions to obtain the channel mean vectors of the source domain and the target domain and , and define the mean alignment loss using the mean squared error : Among them, represents the batch size, and are the height and width of the feature map, respectively; S2. Flatten the feature map into a two-dimensional matrix and then calculate the covariance matrices of the source domain and the target domain and , and define the covariance alignment loss : S3. Combine the mean and covariance alignment losses in a weighted manner to construct the CMA loss function : Among them, and are tunable hyperparameters with default values of 0.1 and 0.05 respectively, which are used to regulate the relative importance of mean alignment and covariance alignment in the loss.

7. A method for enhancing and recognizing small-sample plankton images, characterized in that, It includes the following processes: S1, obtaining high-resolution microscopic images of plankton to be recognized in real time; S2, inputting the microscopic images of plankton to be recognized into the small-sample plankton image enhancement and recognition model built by the building method according to any one of claims 1 to 6; S3, outputting real-time classification results.

8. A small-sample plankton image enhancement and recognition device, characterized in that: The device includes at least one processor and at least one memory, and the processor is coupled to the memory; the memory stores a computer execution program of the small-sample plankton image enhancement and recognition model built by the building method according to any one of claims 1 to 6; when the processor executes the computer execution program stored in the memory, the processor executes a small-sample plankton image enhancement and recognition method.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer execution program of the small-sample plankton image enhancement and recognition model built by the building method according to any one of claims 1 to 6, and when the computer execution program is executed by the processor, the processor executes a small-sample plankton image enhancement and recognition method.

Citation Information

Patent Citations

  • Hyperspectral image cross-domain ground feature element extraction method based on block characterization mechanism

    CN116883752A

  • Ground disaster remote sensing detection method and device based on YOLOv8 model and medium

    CN118570663A