Small sample plankton image enhancement recognition method and small sample plankton image model building method
By introducing the CD-FRN model of cross-domain feature alignment, channel space gating and morphological perception modules in small sample plankton image recognition task, the problems of insufficient cross-domain generalization ability and difficulty in morphological distinction in plankton image recognition are solved, and high-precision and robust recognition effects are achieved.
Patent Information
- Application Number
- CN202510457298.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In the plankton image recognition task, the existing small sample learning methods have problems such as insufficient cross-domain generalization capability and high difficulty in morphological distinction, which leads to unstable identification and difficult to meet the needs of practical applications for high precision and robustness.
A small sample plankton image enhancement recognition model CD-FRN is proposed. By introducing a cross-domain feature alignment strategy, a channel space gating module and a morphological perception module, the cross-domain adaptability and fine-grain recognition capability of the model are improved.
It effectively improves the cross-domain adaptability and fine-grained classification capabilities of the model in specific marine ecological scenarios, improves the stability and accuracy of recognition, and can achieve high-precision plankton image recognition under limited samples.
Smart Images

Figure CN119992223A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of underwater image enhancement, and in particular relates to a small sample plankton image enhancement recognition method and a model building method thereof. Background Art
[0002] Plankton is a key biological group in aquatic ecosystems. It not only forms the basis of the trophic chain, but also plays a key role in water quality assessment and marine resource management. Automatic identification and monitoring of plankton will help scientifically assess the ecological status of water bodies, predict environmental evolution trends, and provide data support and decision-making basis for fishery development and ecological protection.
[0003] However, in actual operation, the acquisition of plankton images faces many challenges. On the one hand, the collection process is limited by the high equipment cost, the complex and changeable marine environment, and the high dependence on operating skills, resulting in a limited number of available samples; for species that are rare or in specific sea areas, it is even more difficult to build large-scale data sets. On the other hand, plankton has diverse morphologies and subtle differences between genera, which places higher demands on the representation ability of the model; if the training data is insufficient, traditional deep learning methods are often prone to overfitting, resulting in unstable recognition, and it is difficult to meet the requirements of practical applications for high precision and robustness.
[0004] In recent years, a series of small sample learning methods have emerged in the field of computer vision. These methods train the network under the meta-learning framework, so that the model still has the ability to generalize quickly and accurately classify when facing new categories and a very small number of samples. Such methods have shown good application prospects in typical data-scarce scenarios such as medical image diagnosis and remote sensing image analysis, and have important research value and broad practical significance. Among them, the Feature Map Reconstruction Network (FRN) has attracted much attention because it converts small sample classification tasks into feature reconstruction problems. Compared with traditional metric learning or meta-learning methods, FRN can fully retain the spatial structure information of the image and achieve efficient reconstruction of features through closed linear regression, thereby reducing the number of model parameters while improving inference efficiency. This method has achieved excellent performance in multiple typical small sample learning tasks, reaching or approaching the current state-of-the-art level, showing good versatility and application potential.
[0005] Although FRN performs well in small sample learning tasks, its cross-domain adaptability may still be limited in the task of plankton microscopic image recognition. Plankton images are affected by multiple factors such as morphology, background, and acquisition conditions. The source sea areas are diverse and the visual features vary significantly, which can easily lead to a decrease in the generalization ability of the model when migrating to new scenes. In addition, there are many types of plankton, and the morphological differences between some species are extremely small, which puts higher demands on the fine-grained recognition ability of the model. Therefore, improving the cross-domain adaptability and fine-grained classification ability of the model in specific marine ecological scenarios has important research value and deserves in-depth exploration and application. Summary of the invention
[0006] In view of the above problems, the first aspect of the present invention provides a small sample plankton image enhancement recognition method and a model building method thereof, comprising the following process: Step 1, obtaining a plankton microscopic image dataset; Step 2: preprocess the original image data and divide it into base classes and new classes, with no overlapping categories; Step 3, construct a small sample image enhancement recognition model CD-FRN based on feature map reconstruction network FRN; the model first processes the input image through a deep convolutional neural network to extract its intermediate layer feature map; the intermediate layer feature map is processed by a cross-domain feature alignment strategy, and then sequentially input into the channel space gating module and the morphological perception module for feature enhancement. The cross-domain feature alignment strategy first calculates the statistics of the source domain and target domain feature maps in the spatial dimension, performs standardization and affine alignment processing on the feature maps according to the statistical information, and then nonlinearly calibrates the alignment results through an adjustable power law transformation; the calibrated feature map is then input into the channel space gating module, and the feature response is adjusted through the joint modeling of channel attention and spatial attention, and the gating mechanism is introduced to fuse the calibration features and domain enhancement features to generate a fused feature map; the fusion result is further input into the morphological perception module, and the structural information is extracted through multiple convolution branches of different scales, and the extracted multi-scale features and the fused features are fused through residual connections to generate an enhanced feature map. The enhanced feature map is sent to the feature map reconstruction process. The model constructs a reconstructed support matrix based on the support set in each small sample task, uses closed ridge regression to calculate the linear mapping result of the query feature, and performs category discrimination based on the difference between the mapping result and the original query feature; Step 4, training and verifying the model under the meta-learning framework: in the training phase, several small sample tasks are generated from the base class and divided into a support set and a query set. The model is optimized in an end-to-end manner to minimize the value of the joint loss function, which is composed of a main classification loss function based on the reconstruction error and a CMA loss function based on channel statistical alignment; in the verification phase, tasks are constructed from the new class to evaluate the model performance, and the model parameters with the best verification effect are saved for subsequent applications.
[0007] Preferably, a cross-domain feature alignment strategy is adopted in the CD-FRN model in step 3, and domain-invariant feature extraction is achieved through dynamic standardization and feature calibration. The specific implementation includes the following steps: S1, feature normalization and domain alignment. The input feature map is normalized in the spatial dimension to suppress the visual style differences between different domains. Let the source domain feature map be ,in Indicates the batch size, Indicates the number of channels, and Represent the height and width of the feature map respectively. Calculate the spatial dimension mean and standard deviation of the feature map in the spatial dimension:
[0008]
[0009] The standardized features are:
[0010] in, and are the mean and standard deviation of the source domain features, To avoid division by zero errors, is the standardized source domain feature map; If the target domain features are available, the target domain features are calculated synchronously Statistics:
[0011]
[0012]
[0013] in, and are the mean and standard deviation of the target domain features, respectively. is the standardized target domain feature map; then, the standardized source domain feature map is affine transformed to the target domain distribution space to obtain the aligned source domain feature map :
[0014] S2, feature response calibration. Through a learnable exponential factor Aligned feature map Apply a sign-preserving power-law nonlinear transformation to obtain the calibrated features :
[0015] in, The initial value of is 0.95, which is updated by back propagation during training and constrained to the interval (0, 1) by the Sigmoid function to control the transformation strength. The calibrated features are passed on as intermediate features for subsequent fusion and enhancement operations.
[0016] Preferably, a channel space gating module is used in the CD-FRN model in step 3 to enhance the adaptability of the model between different domains. Specifically, the implementation of the channel space gating module includes the following steps: S1, channel attention calculation. For the input feature map Perform a global average pooling operation to generate a shape of Then, two 1×1 convolutional layers are used to transform the pooled features: the first convolutional layer compresses the number of channels to 1 / 8 of the original number of channels and performs nonlinear mapping through the ReLU activation function; the second convolutional layer restores the channel dimension to its original size and generates a shape of The channel attention weight of S2, spatial attention calculation. For the input feature map Apply a 1×1 convolutional layer to generate a network of size The spatial attention map is used to weight each spatial position of the input feature map, and the weights are generated through the Sigmoid function; S3, feature weighted fusion. Expand the channel attention and spatial attention to the same dimension and perform element-by-element multiplication to obtain the joint attention weight. . Then, Input to a 1×1 convolutional layer initialized to the identity matrix to generate domain-specific response feature maps Apply the joint attention weights to the domain feature map to obtain the modulated feature response In order to achieve the adaptive fusion of calibration features and domain enhancement features, a learnable scalar parameter is introduced , and generate the gating factor through the Sigmoid function , as the fusion weight. The fusion operation is expressed as follows:
[0017] in, is the fused feature. This module realizes the effective modeling and weighted integration of cross-domain features through the joint attention mechanism and gated regulation structure.
[0018] Preferably, a morphological perception module is introduced into the CD-FRN model in step 3, and morphological features at different scales are extracted through a multi-scale convolution structure to adapt to the subtle differences and complex structures between plankton classes. The morphological perception module includes three parallel convolution branches, which use 1×1, 3×3, and 5×5 convolution kernel feature maps respectively. Processing is used to obtain image structure information at local, basic and wide scales; in order to dynamically fuse the responses of different convolution branches, the module introduces a learnable weight vector , and normalized by the softmax function to generate proportional weights. The output feature map is recorded as , , , get the fused feature map :
[0019] In order to enhance the morphological response while retaining the original structural information, the module adds the input feature map and the fusion feature through a residual connection to obtain the final output of the module :
[0020] in, is a trainable residual adjustment factor used to balance the degree of fusion between the original input and the morphological enhancement features. It will be used as input to the feature reconstruction phase.
[0021] Preferably, the CD-FRN model in step 3 uses a feature map reconstruction network to linearly reconstruct the features of the query image. Specifically, the intermediate layer feature map output in the feature enhancement stage is Input to the reconstruction module, where , , , are the batch size, number of channels, feature map height and width respectively; then, the feature map is flattened in the spatial dimension to convert it into a two-dimensional matrix ,in represents the number of spatial positions, Represents the feature dimension, which serves as the input for subsequent reconstruction operations.
[0022] After the support set and query set are divided, from the feature matrix Extract the feature vectors of the images corresponding to each category, and construct the support set feature matrix and the query image feature matrix respectively. , and its corresponding The feature vectors of the support images are stacked to form the support feature matrix ; Similarly, the query feature matrix is obtained by flattening the feature map of the query image .
[0023] For Perform linear reconstruction and introduce reconstruction matrix , the specific process is as follows: S1, calculate the autocorrelation matrix of support features :
[0024] S2, based on the support set feature matrix The condition number of the adaptive regularization factor is calculated , whose expression is:
[0025] in, Representation Matrix The condition number of is a learnable baseline regularization parameter; With the identity matrix After multiplication with the autocorrelation matrix Add together to get the regularization matrix :
[0026] S3, build reconstruction matrix :
[0027] S4, by Linearly reconstruct the query features and obtain:
[0028] in, For category-based Reconstructed query feature representation.
[0029] Finally, the reconstructed features With the original query features The mean square error between is used as the classification basis, and the reconstruction score is defined as:
[0030] in, represents the Frobenius norm, Indicates that the query image belongs to the category The smaller the reconstruction error, the better the category The higher the degree of fit for the query sample, the model will take the category with the largest score as the final prediction result.
[0031] Preferably, the CMA (Covariance-Mean Alignment) loss function in step 4 acts on the aligned source domain feature map and the standardized target domain feature map , by jointly aligning the first-order statistics (mean) and the second-order statistics (covariance) to reduce the distribution differences between domains, specifically including the following process: S1, the feature map in batch and spatial dimensions (i.e. ) to find the mean and obtain the channel mean vector and , and define the mean alignment loss as mean square error :
[0032]
[0033]
[0034] S2, flatten the feature map into a two-dimensional matrix and calculate the covariance matrix and , and define the covariance alignment loss :
[0035]
[0036]
[0037] S3, by combining the mean and covariance alignment losses in a weighted manner, constructing the CMA loss function :
[0038] in, and It is an adjustable hyperparameter with default values of 0.1 and 0.05, respectively, which is used to adjust the relative importance of mean alignment and covariance alignment in the loss. The CMA loss function and the main classification loss function together constitute the joint optimization target in the training process.
[0039] The second aspect of the present invention provides a method for enhancing the recognition of small sample plankton images, comprising the following process: S1, real-time acquisition of high-resolution microscopic images of the plankton to be identified; S2, inputting the microscopic image of the plankton to be identified into the small sample plankton image enhancement recognition model constructed by the construction method described in the first aspect; S3, outputs real-time classification results.
[0040] The third aspect of the present invention provides a small sample plankton image enhancement and identification device, the device includes at least one processor and at least one memory, the processor and the memory are coupled; the memory stores a computer execution program of a small sample plankton image enhancement and identification model constructed by the construction method described in the first aspect; when the processor executes the computer execution program stored in the memory, the processor can execute a small sample plankton image enhancement and identification method.
[0041] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer execution program of a small sample plankton image enhancement recognition model constructed by the construction method described in the first aspect. When the computer execution program is executed by a processor, it can enable the processor to execute a small sample plankton image enhancement recognition method.
[0042] Compared with the prior art, the present invention has the following beneficial effects: In response to the problems of insufficient cross-domain generalization ability and high difficulty in morphological distinction faced by traditional small sample learning methods in plankton image recognition, this paper proposes a small sample plankton image enhanced recognition model CD-FRN. On the basis of continuing the efficient modeling advantages of the feature reconstruction mechanism, the model further enhances the adaptability to cross-domain distribution differences and improves the modeling accuracy of complex morphological features.
[0043] First, the model introduces a cross-domain feature alignment strategy, which reduces the distribution differences under different sea areas or acquisition conditions by standardizing and affine mapping the feature maps of the source domain and the target domain in the spatial dimension. On this basis, a channel space gating module is introduced to fuse channel attention, spatial attention and domain-specific response information, and adaptively fuse calibration features and domain enhancement features through a learnable gating mechanism, thereby improving the model's perception and response capabilities to cross-domain feature differences. To further enhance the alignment effect, a CMA loss function is designed to constrain the mean and covariance of the source domain and target domain features in the channel dimension, thereby improving the distribution consistency and cross-domain stability during training.
[0044] In addition, to address the problems of subtle morphological differences and complex texture changes in plankton images, the model introduces a morphological perception module and uses a parallel multi-scale convolutional structure to extract local to wide-area structural information. It also combines learnable fusion coefficients and residual connection mechanisms to achieve dynamic modeling and enhanced expression of multi-scale morphological features, thereby improving the model's fine-grained recognition capabilities.
[0045] In summary, the present invention provides a small sample image modeling method that has both cross-domain adaptability and fine-grained recognition capabilities. It can effectively deal with practical problems such as complex plankton image acquisition conditions, limited samples, and high similarity between morphological classes, and has significant research value. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a schematic diagram of the process of the small sample plankton image enhancement recognition method proposed in the present invention.
[0047] Figure 2 Schematic diagram of a sample microscopic image of plankton.
[0048] Figure 3 This is a schematic diagram of the workflow of the small sample plankton image enhancement recognition model CD-FRN proposed in the present invention.
[0049] Figure 4 This is a schematic diagram of the structure of the feature extraction network ResNet-12 used in the present invention.
[0050] Figure 5 Schematic diagram of the structure of the channel space gating module.
[0051] Figure 6 Schematic diagram of the structure of the morphology perception module.
[0052] Figure 7 Schematic diagram of the feature map reconstruction process.
[0053] Figure 8 The following is a comparison chart of the recognition results of the CD-FRN model and the benchmark model on the test samples.
[0054] Fig. 9 This is a schematic diagram of the simple structure of a small sample plankton image enhancement identification device. DETAILED DESCRIPTION
[0055] The invention will be further described below in conjunction with specific embodiments.
[0056] This embodiment further illustrates the method proposed by the present invention through a specific experimental process. The overall process is as follows: Figure 1 shown.
[0057] 1. Raw data acquisition and preprocessing The image data used in this example comes from the publicly released WHOI (Woods Hole Oceanographic Institution) plankton dataset and Kaggle plankton dataset. The above datasets cover a large number of representative plankton species and are widely used in tasks such as plankton image recognition and classification. Figure 2 shown.
[0058] To ensure data quality, the original image samples are first screened to remove images with quality defects such as missing, blurred, and severely blocked. After completing the preliminary quality check, the images are standardized to unify the data format and distribution; then a variety of data enhancement operations are performed to enrich the diversity of training samples. Specifically, the enhancement operation includes: performing a random cropping operation on the image, with the cropping ratio ranging from 80% to 100% of the original image size; applying a random rotation within the range of ±40 degrees to the image; performing a random horizontal flip operation with a probability of 50% on the image; in order to further adapt to the illumination and color differences caused by different imaging devices and marine environments, the color features of the image are perturbed, where the brightness and contrast vary within the range of ±0.2, and the saturation varies within the range of ±0.3; this embodiment also introduces a random occlusion strategy, generating 1 to 3 rectangular occlusion areas at random positions in each image, each occlusion area accounts for 1% to 5% of the original image, and the aspect ratio is controlled between 0.5 and 2.0. The occlusion area is filled with Gaussian random noise with a mean of 0.5 and a standard deviation of 0.1. This embodiment integrates and applies all the above-mentioned enhancement strategies in the training stage, and various operations can also be flexibly combined and adjusted as needed.
[0059] 2. Model building Based on the above dataset, this embodiment designs and builds a deep learning model Cross Domain-Feature Reconstruction Network (hereinafter referred to as CD-FRN). The overall processing flow is as follows: Figure 3 The model uses a lightweight deep convolutional neural network as a feature extractor, preferably using ResNet-12 to achieve better feature extraction capabilities, and also supports the use of a simpler Conv-4 network as an alternative. This embodiment uses ResNet-12 as the feature extraction network, and its specific structure is as follows Figure 4 The intermediate features extracted by the backbone network are then processed with multiple feature enhancements to enhance the model's cross-domain adaptability and the ability to represent fine-grained morphological features, thereby achieving high-precision recognition of plankton microscopic images under conditions of limited samples and domain differences in data distribution.
[0060] First, the model adopts a cross-domain feature alignment strategy to reduce the distribution differences between different data domains. Specifically, let the source domain feature map be ,in , , , are batch size, number of channels, feature map height and width, and the same symbols in the following text have the same meaning. First, the mean is calculated in the spatial dimension With standard deviation , and calculate the standardized source domain feature map :
[0061]
[0062]
[0063] in, is a constant term used to avoid numerical instability when the standard deviation is zero. If available, its statistics are calculated synchronously, including the mean of the spatial dimension , Standard Deviation , Standardized target domain feature map , which is defined in the same form as the source domain:
[0064]
[0065]
[0066] The target domain feature map here It can be obtained in both training and validation stages: in the training stage, training samples from the target domain are used to assist in the statistical alignment of source domain features; in the validation stage, support set images in the new class of tasks are extracted to simulate cross-domain adaptation in real test scenarios.
[0067] Next, the standardized source domain feature map is mapped to the target domain distribution space to achieve affine alignment and obtain the aligned source domain feature map :
[0068] The model introduces a learnable power-law calibration mechanism, which applies a sign-preserving power-law transformation to the aligned features to obtain the calibrated features. :
[0069] in, It is a learnable parameter with an initial value of 0.95 and is constrained to the interval (0,1) by the Sigmoid function to improve training stability. This mechanism balances noise suppression and key feature enhancement by dynamically adjusting the nonlinear scaling strength of the feature amplitude.
[0070] Next, the calibrated features are fed into the channel-space gating module, which uses the joint attention mechanism of the channel and spatial dimensions to dynamically adjust the feature response strength, thereby improving cross-domain adaptability. The structure of this module is as follows Figure 5 As shown. First, the channel attention weight is calculated for the input feature map. Specifically, the input feature map Perform a global average pooling operation to generate a shape of The channel description vector is then passed through two 1×1 convolutional layers for channel information interaction: the first layer compresses the number of channels to 1 / 8 of the original number of channels and implements nonlinear mapping through ReLU activation; the second layer restores the channel dimension to its original size and generates a shape of through the Sigmoid activation function. At the same time, the input feature map Apply a 1×1 convolution operation and activate it with a Sigmoid function to generate a shape of Then, the channel attention weights and spatial attention maps are expanded to the input features through the broadcast mechanism. Figure 1 The same dimension is obtained by multiplying them point by point at the element level to obtain the joint attention weight. .
[0071] Next, a linear projection is performed on the input feature map to generate a domain-specific response feature map , the projection operation is implemented by a 1×1 convolutional layer initialized to the identity matrix. Applied to domain feature map , and obtain the characteristic response after modulation , where multiplication is an element-wise weighted operation.
[0072] Then, a learnable scalar gating parameter is introduced , the gating factor is generated by the Sigmoid function , to control the fusion ratio between calibration features and domain enhancement features. To achieve balanced feature fusion in the initial stage, The initial value of is set to 0, corresponding to The initial value is 0.5. The final fusion result is as follows:
[0073] in, is the fused feature. This structure realizes dynamic balance adjustment of different feature sources through joint attention modeling and gating mechanism.
[0074] In order to enhance the model's ability to perceive the details of the plankton structure and morphology, the model introduces a morphology perception module, whose structure is as follows: Figure 6 As shown in Figure 2, the module consists of three parallel convolution branches, using 1×1, 3×3 and 5×5 respectively, to introduce multi-scale receptive fields. Each branch performs a multi-scale receptive field on the input feature map. Perform structural modeling at local, basic and wide-area scales, and output feature maps of corresponding scales respectively , , In order to achieve dynamic fusion of multi-scale features, the model introduces a learnable fusion weight vector , whose initial value is [0.33, 0.34, 0.33], and is normalized by the softmax function to generate proportional weights. The fused feature map Calculated by the following formula:
[0075] In order to enhance the morphological response while retaining the original structural information, the model adds the input feature map to the fusion feature through a residual connection to obtain the final output of the module. :
[0076] in, is a trainable residual adjustment factor used to balance the degree of fusion between the original input and the morphological enhancement features.
[0077] After completing the above enhancement processing, the final features are sent to the feature map reconstruction network for category discrimination. The basic principle of the reconstruction process is as follows Figure 7 For each image, its feature map after processing by the above modules is represented as , after flattening, we get a two-dimensional matrix representation ,in Indicates the number of spatial positions, Represents the feature dimension.
[0078] In each small sample classification task, from the feature matrix Extract the feature representation of the support image and query image corresponding to each category. Specifically, for each category , and its corresponding The flattened feature vectors of the support images are concatenated in sequence to form the support feature matrix ; The feature map of the query image is also flattened to obtain the query feature matrix The model then constructs the reconstruction mapping weights based on the closed-form ridge regression solution , and linearly reconstruct the query features. The calculation process is as follows:
[0079]
[0080]
[0081]
[0082]
[0083] in, represents the autocorrelation matrix of the support set features, express The condition number of is a learnable baseline regularization parameter, is the class-dependent adaptive regularization factor, is the unit matrix, is the regularized matrix, For category-based The query feature representation obtained by reconstruction. Finally, the model makes category discrimination based on the reconstruction error, and the reconstruction score is defined as:
[0084] in, represents the Frobenius norm, Indicates that the query image belongs to the category The score of is calculated in parallel for all candidate categories. , select the category with the largest score as the prediction result.
[0085] 3. Model training The implementation environment of a small sample plankton image enhancement recognition method in this embodiment is based on the Ubuntu 22.04 operating system, the programming language is Python 3.10, the deep learning framework uses Pytorch 2.1.0, and CUDA 12.1 is used. The model was trained for 1200 rounds on a system equipped with an NVIDIA GeForce RTX 4090 GPU (video memory 24GB). The relevant settings of each training parameter are also recorded in detail in the experiment, as shown in Table 1.
[0086] Table 1 Experimental configuration details
[0087] The learning rate scheduling adopts the ReduceLROnPlateau strategy to monitor the changes in the validation loss. When the loss does not decrease for 20 consecutive rounds, the learning rate is decayed to 0.1 times the current value until the learning rate is no less than 1e-6.
[0088] This embodiment designs a CMA (Covariance-Mean Alignment) loss function to constrain the mean and covariance matrix in the channel dimension. Normalized feature map with target domain , calculate the mean vector and , and define the mean alignment loss :
[0089]
[0090]
[0091] Then flatten the feature map into the shape of Calculate the covariance matrix of and , and define the covariance of its loss :
[0092]
[0093]
[0094] Finally, the CMA loss function is weighted fused in the following form to obtain the CMA loss :
[0095] in, and is the loss weighting factor (default =0.1, = 0.05) to balance the importance of mean and covariance alignment.
[0096] In order to construct a complete training objective function, the model adopts a joint optimization strategy of reconstruction classification loss and CMA loss. Based on the feature map reconstruction error, it is used to measure the model's ability to discriminate the query image; CMA loss It is used to improve the cross-domain generalization performance of the model. The two together constitute the final training goal:
[0097] By minimizing the joint loss function , that is, the model simultaneously optimizes the classification accuracy and cross-domain alignment capability, thereby achieving consistency in feature distribution while learning discriminative features, further improving the recognition accuracy and robustness of the model under small sample conditions.
[0098] 4. Experimental Results To verify the effectiveness and superiority of this method, this embodiment conducted comparative experiments and ablation experiments. First, this method was compared with the existing small sample learning method to evaluate its performance in the plankton image recognition task; second, multiple groups of ablation experiments were constructed to gradually introduce each core module and quantify its impact on the recognition effect.
[0099] Classification accuracy is the core evaluation indicator of this experiment, which measures the accuracy of the model in classifying the target sample. In this example, for the small sample plankton recognition task, the classification accuracy is evaluated under 5-way 1-shot and 5-way 5-shot settings.
[0100] Comparative experiment: This embodiment selects a variety of excellent deep learning models in the field of small sample learning as comparison models, including Matching Nets, Prototypical Nets, MAML, etc. The specific comparison experimental results are shown in Table 2.
[0101] Table 2 Comparison of 5-way classification accuracy of different small sample learning methods
[0102] Experimental results show that the proposed method outperforms existing methods under different settings of the WHOI-Plankton and Kaggle Plankton datasets, fully demonstrating its performance advantages in complex tasks such as sample scarcity, inter-domain shift, and fine-grained structure recognition.
[0103] Ablation experiment: In order to verify the contribution of each structural module in the present invention to the overall model performance, multiple groups of ablation experiments were designed and carried out, and key functional modules were gradually introduced to evaluate their gain effects in the small sample plankton image recognition task. The results are shown in Table 3. Among them, the "cross-domain adaptation mechanism" includes the cross-domain feature alignment strategy, the channel space gating module and the CMA loss function. The three are related in the model structure and training process, so it is more valuable to perform ablation verification as a whole.
[0104] Table 3 Ablation experiment results
[0105] The ablation experiment results show that each of the above components contributes to the performance gain of the present invention.
[0106] 5. Visualization of results In order to further verify the recognition performance and visual discrimination ability of the proposed method, this example performs a visual analysis of the prediction results of several plankton images in the test set. Figure 8 The comparison of the classification results of the baseline model FRN and the model CD-FRN proposed in this invention on the same sample is shown. The ground truth, model prediction and corresponding prediction confidence are marked above each sub-image in the figure. As can be seen from the figure, the CD-FRN model shows stronger recognition ability and higher prediction confidence in the plankton images with complex morphological structures, further demonstrating its advantages in fine-grained feature modeling and cross-domain generalization.
[0107] like Fig. 9 As shown, the present invention also provides a small sample plankton image enhancement recognition device, the device includes at least one processor and at least one memory, and also includes a communication interface and an internal bus; the memory stores a computer execution program; the memory stores a computer execution program of a small sample plankton image enhancement recognition model constructed by the construction method as described above; when the processor executes the computer execution program stored in the memory, the processor can execute a small sample plankton image enhancement recognition method. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus. The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a U disk, a mobile hard disk, a read-only memory, a disk or an optical disk.
[0108] The device may be provided as a terminal, a server or other forms of device. In an exemplary embodiment, the electronic device may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0109] The present invention also provides a computer-readable storage medium, which stores a computer execution program of the small sample plankton image enhancement recognition model constructed by the construction method described above. When the computer execution program is executed by a processor, the processor can execute a small sample plankton image enhancement recognition method.
[0110] Specifically, a system, device or equipment equipped with a readable storage medium may be provided, on which a software program code for implementing the functions of any of the above-mentioned embodiments is stored, and a computer or processor of the system, device or equipment reads and executes the instructions stored in the readable storage medium. In this case, the program code read from the readable medium itself can implement the functions of any of the above-mentioned embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.
[0111] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0112] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A method for building a small sample plankton image enhancement recognition model, characterized in that: The process includes: Step 1, obtaining a plankton microscopic image dataset; Step 2: preprocess the original image data and divide it into base classes and new classes, with no overlapping categories; Step 3, construct a small sample image enhancement recognition model CD-FRN based on feature map reconstruction network FRN; the model first receives the input image and extracts the intermediate layer feature map through a deep convolutional neural network; then, introduces multiple feature enhancement mechanisms, including a cross-domain feature alignment strategy for performing distribution normalization and calibration, a channel space gating module for enhancing domain-specific responses, and a morphological perception module for modeling multi-scale morphological information; The enhanced feature map is fed into the feature map reconstruction process. In each small sample task, the model constructs a reconstruction support matrix based on the support set, uses closed ridge regression to achieve linear fitting of the query features, and uses the reconstruction error as the basis for class discrimination. Step 4: Train and verify the CD-FRN model under the meta-learning framework to obtain the optimal model parameters; the loss function consists of the main classification loss function based on the reconstruction error and the CMA loss function based on channel statistical alignment.
2. A method for building a small sample plankton image enhancement recognition model as claimed in claim 1, characterized in that: The CD-FRN model first processes the input image through a deep convolutional neural network to extract its middle layer feature map; the middle layer feature map is processed by a cross-domain feature alignment strategy, and then sequentially input into a channel space gating module and a morphological perception module for feature enhancement; the cross-domain feature alignment strategy first calculates the statistics of the source domain and target domain feature maps in the spatial dimension, performs standardization and affine alignment processing on the feature maps according to the statistical information, and then nonlinearly calibrates the alignment results through an adjustable power law transformation; the calibrated feature map is then input into the channel space gating module, and the feature response is adjusted through the joint modeling of channel attention and spatial attention, and a gating mechanism is introduced to fuse the calibration features and domain enhancement features to generate a fused feature map; The fusion result is further input into the morphological perception module, which extracts structural information through multiple convolution branches of different scales, and fuses the extracted multi-scale features with the fusion features through residual connection to generate an enhanced feature map; The enhanced feature map is sent to the feature map reconstruction process. The model constructs a reconstructed support matrix based on the support set in each small sample task, uses closed ridge regression to calculate the linear mapping result of the query feature, and performs category discrimination based on the difference between the mapping result and the original query feature.
3. A method for building a small sample plankton image enhancement recognition model as claimed in claim 1, characterized in that: The cross-domain feature alignment strategy is to achieve domain-invariant feature extraction through dynamic normalization and feature calibration, specifically including: Feature normalization and domain alignment: The input feature map is normalized in the spatial dimension to suppress the visual style differences between different domains; the source domain feature map is recorded as ,in, Indicates the batch size, Indicates the number of channels, and Respectively represent the height and width of the feature map; first, by calculating the channel mean and standard deviation of the feature map in the two spatial dimensions of height and width, the statistics of the source domain features are obtained and , and based on this, the standardization process is completed to obtain the standardized source domain feature map: in, and are the mean and standard deviation of the source domain features, is the standardized source domain feature map; If the target domain features are available, the target domain feature map is synchronized in the same way. Calculate its mean in the spatial dimension With standard deviation , and obtain the standardized result ; Then, the standardized source domain feature map is affine transformed to the target domain distribution space to obtain the aligned source domain feature map : Feature response calibration: via a learnable exponential factor Aligned feature map Apply a sign-preserving power-law nonlinear transformation to obtain the calibrated feature map : in, The initial value of is 0.95, which is updated by back propagation during training and constrained to the interval (0, 1) by the Sigmoid function to control the transformation strength; The calibrated features are then passed on as intermediate features.
4. A method for building a small sample plankton image enhancement recognition model as claimed in claim 1, characterized in that: The channel space gating module is used to enhance the adaptability of the model between different domains; specifically includes: S1, channel attention calculation; for input feature map Perform a global average pooling operation to generate a shape of Then, two 1×1 convolutional layers are used to transform the pooled features: the first convolutional layer compresses the number of channels to 1 / 8 of the original number of channels and performs nonlinear mapping through the ReLU activation function; the second convolutional layer restores the channel dimension to its original size and generates a shape of The channel attention weight of S2, spatial attention calculation; for the same input feature map Apply a 1×1 convolutional layer to generate a network of size The spatial attention map is used to weight each spatial position of the input feature map, and the weights are generated through the Sigmoid function; S3, feature weighted fusion: expand the above channel attention and spatial attention to the same dimension, and multiply them element by element to generate the joint attention weight At the same time, Input to a 1×1 convolutional layer initialized to the identity matrix to generate domain-specific response feature maps ; By adding the joint attention weight Applied to feature map , get the characteristic response after modulation ; To achieve the calibration feature Adaptive fusion between the modulation response and the scalar parameter , and generate the gating factor through the Sigmoid function , as the fusion ratio coefficient; the final output feature It consists of the weighted addition of two parts: one part retains the original information of the calibration features, and the other part fuses the response features after domain-specific enhancement. The ratio of the two is automatically adjusted by the gating factor.
5. The method for building a small sample plankton image enhancement recognition model according to claim 1, characterized in that: The morphological perception module extracts morphological features at different scales through a multi-scale convolution structure to adapt to the subtle differences and complex structures of plankton. The morphological perception module includes three parallel convolution branches, which use 1×1, 3×3, and 5×5 convolution kernel feature maps respectively. Processing is used to obtain image structure information at local, basic and wide scales; in order to dynamically fuse the responses of different convolution branches, the module introduces a learnable weight vector , and normalized by the softmax function to generate proportional weights; the output feature map is recorded as , , , get the fused feature map : In order to enhance the morphological response while retaining the original structural information, the module adds the input feature map and the fusion feature through a residual connection to obtain the final output of the module : in, is a trainable residual adjustment factor used to balance the degree of fusion between the original input and the morphological enhancement features; the output feature map It will be used as input to the feature reconstruction phase.
6. A method for building a small sample plankton image enhancement recognition model as claimed in claim 1, characterized in that: The CMA loss function acts on the aligned source domain feature map and the standardized target domain feature map , by jointly aligning the first-order statistics and the second-order statistics to reduce the distribution differences between domains, specifically: S1, average the feature map in batch and spatial dimensions to obtain the channel mean vector of the source domain and the target domain and , and define the mean alignment loss as mean square error : in, Indicates the batch size, and are the height and width of the feature map respectively; S2, flatten the feature map into a two-dimensional matrix and calculate the covariance matrix between the source domain and the target domain and , and define the covariance alignment loss : S3, by combining the mean and covariance alignment losses in a weighted manner, constructing the CMA loss function : in, and It is an adjustable hyperparameter with default values of 0.1 and 0.05, respectively, which is used to adjust the relative importance of mean alignment and covariance alignment in the loss.
7. A method for enhancing the recognition of small sample plankton images, characterized in that: The process includes: S1, real-time acquisition of high-resolution microscopic images of the plankton to be identified; S2, inputting the microscopic image of the plankton to be identified into the small sample plankton image enhancement recognition model constructed by the construction method according to any one of claims 1 to 6; S3, outputs real-time classification results.
8. A small sample plankton image enhancement identification device, characterized by: The device includes at least one processor and at least one memory, and the processor and the memory are coupled; the memory stores a computer execution program of a small sample plankton image enhancement recognition model constructed by the construction method described in any one of claims 1 to 6; when the processor executes the computer execution program stored in the memory, the processor executes a small sample plankton image enhancement recognition method.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer execution program of a small sample plankton image enhancement recognition model constructed by the construction method as described in any one of claims 1 to 6. When the computer execution program is executed by a processor, the processor executes a small sample plankton image enhancement recognition method.
Citation Information
Patent Citations
Hyperspectral image cross-domain ground feature element extraction method based on block characterization mechanism
CN116883752A
Ground disaster remote sensing detection method and device based on YOLOv8 model and medium
CN118570663A
Breast image classification evaluation method and system based on deep learning
CN119478561A
Wheat growth period identification method based on multi-modal characteristics
CN119516367A
System for fast and accurate visual domain adaptation
US10839269B1
Cited By
Construction method and system of fine tree species identification model, identification method and system of fine tree species identification model
CN120673265A
Training and application method and device of few-sample learning model, equipment and medium
CN121144845A
Method and device for enhancing perception data, electronic equipment and storage medium
CN121329827A
Method and apparatus for enhancing perception data, electronic device, storage medium
CN121329827B
Phytoplankton chromatography sequence identification method and phytoplankton chromatography sequence model building method
CN121838155A