Rare waterfowl identification method under severe weather condition

By constructing damaged rare water bird assisted data sets and joint learning models, the problem of rare water bird image recognition under severe weather conditions is solved, and the recognition performance improvement against severe weather conditions is achieved.

CN120088557AInactive Publication Date: 2025-06-03NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164738.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In severe weather conditions, the image recognition tasks of rare water birds are affected by problems such as noise, low contrast and image blur, resulting in increased recognition difficulty. The existing technology lacks end-to-end training models to link underlying visual tasks with small sample learning.

Method used

A rare water bird recognition method under severe weather conditions is proposed. By constructing a damaged rare water bird assisted data set, and a joint learning model composed of image restoration network and small sample learner is constructed to optimize the learning model to improve recognition performance.

Benefits of technology

Effectively fight the impact of bad weather conditions on small sample image classification tasks, and improve the performance of rare water bird identification tasks in real scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088557A_ABST
    Figure CN120088557A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a rare waterfowl recognition method under severe weather conditions, which comprises the following steps: firstly, collecting common birds and wetland environment images to obtain an auxiliary image data set for rare waterfowl recognition, and carrying out quality degradation processing on the data set to obtain a damaged auxiliary data set; then, constructing a bottom-layer vision and small sample image classification combined learning model; next, inputting the image in the damaged auxiliary data set into the joint learning model, establishing reconstruction loss between the output of the image restoration sub-network and the clear auxiliary image, and establishing a cross entropy loss function between the output of the small sample learner and the classification label value; and finally, fixing the parameters of the joint learning model, adding an adapter, collecting the image of the rare waterfowl, and finely adjusting the parameters of the adapter to complete the recognition of the rare waterfowl. According to the method, the influence of severe weather conditions on a small sample image classification task is resisted, and the performance of a rare waterfowl recognition task in a real scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method for identifying rare water birds under adverse weather conditions. Background Art

[0002] Wetlands are known as the "kidneys of the earth" and are one of the most productive ecosystems on earth, with powerful ecological service functions. Wetland water birds are an important part of the wetland ecosystem. Protecting wetland water birds is of great significance for maintaining the health of the wetland ecosystem, promoting biodiversity, and driving the sustainable development of human society. The number of rare water birds is limited and their distribution range is wide, resulting in a very limited number of samples available for training. Therefore, accurately automatically identifying rare water birds is a very challenging task.

[0003] In the task of automatically identifying rare water birds, adverse weather conditions are an important challenge. Images taken under adverse weather conditions often have problems such as noise, low contrast, and image blurring. At the same time, the appearance characteristics of rare water birds may change, such as the feather color becoming darker and the shape becoming blurred, which greatly increases the difficulty of identification. Underlying computer vision tasks can perform image restoration on the reduction of image quality caused by factors such as noise, distortion, and motion blur. Therefore, in the task of automatically identifying rare water birds, it is necessary to introduce underlying vision tasks into the few-shot learning model. Currently, excellent methods in underlying vision tasks and few-shot learning tasks are all based on deep neural network models. However, the research on these two tasks is separated from each other, and there is still a lack of an end-to-end training model to connect underlying vision tasks and few-shot learning.

[0004] To solve the above problems, the present invention proposes a method for identifying rare water birds under adverse weather conditions, so as to counter the influence of adverse weather conditions on the few-shot image classification task and improve the performance of the rare water bird identification task in real scenarios. Summary of the Invention

[0005] The purpose of the present invention is to solve the disadvantages existing in the prior art, and to propose a method for identifying rare water birds under adverse weather conditions, which can counter the influence of adverse weather conditions on the few-shot image classification task and improve the performance of the rare water bird identification task in real scenarios.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for identifying rare water birds under adverse weather conditions includes the following steps:

[0008] Step 1: Construct a damaged auxiliary dataset for rare waterbirds: Collect images of common birds and wetland environments to obtain an auxiliary image dataset for rare waterbird recognition, and degrade this dataset to obtain a damaged auxiliary dataset;

[0009] Step 2: Construct a joint learning model for low-level vision and few-shot image classification: This model consists of an image restoration sub-network and a few-shot learner;

[0010] Step 3: Optimize the joint learning of low-level vision and few-shot image classification: Establish a reconstruction loss between the output of the image restoration sub-network and the clear auxiliary images, and establish a cross-entropy loss function between the output of the few-shot learner and the classification label values. Use the weighted sum of the reconstruction loss and the cross-entropy loss function as the total loss function to optimize the joint learning model;

[0011] Step 4: Fine-tune the joint learning model for low-level vision and few-shot image classification: Fix the parameters of the joint learning model and add an adapter, collect rare waterbird images to fine-tune the adapter parameters, and complete the recognition of rare waterbirds.

[0012] Preferably, in Step 1, the specific method is as follows:

[0013] Step 11: Collect and organize 200 images of each category in common birds and wetland plants to construct an auxiliary dataset;

[0014] Step 12: For any image in the auxiliary dataset, represented as I(x), perform fogging degradation processing on it according to the following formula:

[0015] H(x) = I(x)T(x) + A(1 - T(x))

[0016] T(x) = exp {-βd(x)}

[0017] where H(x) represents the image of I(x) after fogging, A represents the global atmospheric light coefficient, T(x) represents the medium transmission map, β represents the atmospheric scattering coefficient, and d(x) represents the distance from the camera to the object;

[0018] Step 13: For any image in the auxiliary dataset, represented as I(x), perform rain degradation processing on it according to the following formula:

[0019]

[0020] where R(x) represents the image of I(x) after rain, and S l (x) represents the l-th layer rain streak layer image;

[0021] Step 14: For any image in the auxiliary dataset, represented as I(x), perform snowing degradation processing on it according to the following formula:

[0022] S(x) = I(x)(1 - Z(x)) + C(x)Z(x)

[0023] where S(x) represents the image after snowing of I(x), Z(x) represents the snowflake image, and C(x) represents the mask image;

[0024] Step 15: The foggy image H(x), the rainy image R(x), and the snowing image S(x) constitute the damaged auxiliary dataset.

[0025] Preferably, in Step 2, the specific method is as follows:

[0026] Step 21: The image restoration sub-network adopts a U-shaped structure, mainly composed of two components: an encoder and a decoder. Generally speaking, the encoder and the decoder are divided into 4 stages, and the encoder and the decoder in the same stage are connected by skip layers; the encoder in each stage is composed of a Transformer module and a downsampling layer, and the decoder in each stage is composed of a Transformer layer and an upsampling layer; in order to reduce the computational complexity, a window-type multi-head attention mechanism is adopted in the Transformer module, which divides the feature map into several small windows in space, and the calculation of multi-head attention is limited within the windows; assuming that the feature before the input of each Transformer module is represented as Z, the mathematical representation of the processing of Z by the Transformer module is as follows:

[0027]

[0028] where W Q represents the query value linear representation layer, W K represents the key value linear representation layer, W V represents the value linear representation. A 3×3 convolutional layer is used to perform convolutional operations on the output features of the last stage of the decoder, and the result is superimposed with the input degraded image to obtain the output of the image restoration sub-network;

[0029] Step 22: The few-shot learner is mainly composed of a feature extractor and a linear classifier. The feature extractor is composed of 4 residual blocks, each residual block contains three convolutional layers, and the size of the filter for each layer is 3×3. The first three residual blocks use max pooling operations, while the last residual block uses average pooling operations. The linear classifier selects a Softmax classifier.

[0030] Preferably, in Step 3, the specific method is as follows:

[0031] Step 31: Randomly select a batch of images from the damaged auxiliary dataset. The number of images is N, and any one image is represented as d i (x), and its corresponding clear image is represented as I i (x), and its corresponding class label is represented as O i , and input d i (x) into the image restoration sub-network to obtain the restored image y i (x), then construct the following reconstruction loss function:

[0032]

[0033] Step 32: Input the output image y i (x) of the image restoration sub-network into the few-shot learner. The feature extractor represents it as the deep feature representation F i , and the classifier then converts F i into the output logical value T i , and based on T i and O i calculate the following cross-entropy loss function:

[0034]

[0035] where M is the total number of image categories in the auxiliary dataset, and O ij represents the j-th component of O i , and T ij represents the j-th component of T i ;

[0036] Step 33: Establish a joint loss function, and its mathematical expression is as follows:

[0037] Loss J = Loss C + αLoss R

[0038] where α represents an adjustable parameter. Based on the above formula, use the gradient descent algorithm to optimize the parameters of the joint learning model for low-level vision and few-shot image classification.

[0039] Preferably, in step 4, the specific method is as follows:

[0040] Step 41: Remove the Softmax classifier in the joint learning model for low-level vision and few-shot image classification, and fix the model parameters therein. Add an adapter in the joint learning model. The adapter is represented as A(·), and this adapter is composed of a multi-layer perceptron;

[0041] Step 42: Collect and organize the images of rare water birds. Any one image is represented as b i(x), whose corresponding class label is denoted as U i , input b i (x) into the trained joint learning model of low-level vision and few-shot image classification, and then input it into the adapter A(·) to obtain the output logical value V i , based on V i and U i calculate the following cross-entropy loss function:

[0042]

[0043] where Q is the total number of rare waterbird image categories, and U ij represents the j-th component of U i , and V ij represents the j-th component of V i ;

[0044] Based on the above formula, use the gradient descent algorithm to fine-tune and optimize the parameters of the adapter, and jointly use the trained joint learning model of low-level vision and few-shot image classification without the Softmax classifier and the adapter to complete the recognition of rare waterbirds under adverse weather conditions.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. The present invention can establish an auxiliary dataset for rare waterbirds, providing important data for the research on cross-domain rare waterbird recognition methods.

[0047] 2. The present invention can counter the influence of adverse weather conditions on the few-shot image classification task and improve the performance of the rare waterbird recognition task in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings, so that those skilled in the art can better understand the advantages and features of the present invention, and thus more clearly define the protection scope of the present invention. The embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Refer to Figure 1 , a method for recognizing rare waterbirds under adverse weather conditions, comprising the following steps:

[0051] Step 1: Construct a damaged auxiliary dataset for rare waterbirds: Collect images of common birds and wetland environments to obtain an auxiliary image dataset for rare waterbird recognition, and degrade this dataset to obtain a damaged auxiliary dataset;

[0052] Step 2: Construct a joint learning model for low-level vision and few-shot image classification: This model consists of an image restoration sub-network and a few-shot learner;

[0053] Step 3: Optimize the joint learning of low-level vision and few-shot image classification: Establish a reconstruction loss between the output of the image restoration sub-network and the clear auxiliary images, and establish a cross-entropy loss function between the output of the few-shot learner and the classification label values. Use the weighted sum of the reconstruction loss and the cross-entropy loss function as the total loss function to optimize the joint learning model;

[0054] Step 4: Fine-tune the joint learning model for low-level vision and few-shot image classification: Fix the parameters of the joint learning model and add an adapter, collect rare waterbird images to fine-tune the adapter parameters, and complete the recognition of rare waterbirds.

[0055] Specifically, in Step 1, the specific method is as follows:

[0056] Step 11: Collect and organize 200 images of each category in common birds and wetland plants to construct an auxiliary dataset;

[0057] Step 12: For any image in the auxiliary dataset, denoted as I(x), perform a fogging degradation process on it according to the following formula:

[0058] H(x) = I(x)T(x) + A(1 - T(x))

[0059] T(x) = exp {-βd(x)}

[0060] where H(x) represents the image of I(x) after fogging, A represents the global atmospheric light coefficient, T(x) represents the medium transmission map, β represents the atmospheric scattering coefficient, and d(x) represents the distance from the camera to the object;

[0061] Step 13: For any image in the auxiliary dataset, denoted as I(x), perform a rain degradation process on it according to the following formula:

[0062]

[0063] where R(x) represents the image of I(x) after rain, S l (x) represents the l-th layer of rain streak image;

[0064] Step 14: For any image in the auxiliary dataset, represented as I(x), perform snowing degradation processing on it according to the following formula:

[0065] S(x) = I(x)(1 - Z(x)) + C(x)Z(x)

[0066] Where S(x) represents the image after snowing of I(x), Z(x) represents the snowflake image, and C(x) represents the mask image;

[0067] Step 15: The fogged image H(x), the rained image R(x), and the snowed image S(x) form the damaged auxiliary dataset.

[0068] Specifically, in Step 2, the specific method is as follows:

[0069] Step 21: The image restoration sub-network adopts a U-shaped structure, mainly composed of two components: an encoder and a decoder. Generally speaking, the encoder and the decoder are divided into 4 stages, and the encoder and the decoder in the same stage are connected by skip layers; the encoder in each stage is composed of a Transformer module and a downsampling layer, and the decoder in each stage is composed of a Transformer layer and an upsampling layer; in order to reduce the computational complexity, a window-type multi-head attention mechanism is adopted in the Transformer module, which divides the feature map into several small windows in space, and the calculation of multi-head attention is limited within the windows; assuming that the feature before the input of each Transformer module is represented as Z, the mathematical representation of the Transformer module's processing of it is as follows:

[0070]

[0071] Where W Q represents the query value linear representation layer, W K represents the key-value linear representation layer, W V represents the value linear representation. Use a 3×3 convolutional layer to perform convolutional operations on the output features of the last stage of the decoder, and superimpose it with the input degraded image to obtain the output of the image restoration sub-network;

[0072] Step 22: The few-shot learner is mainly composed of a feature extractor and a linear classifier. The feature extractor is composed of 4 residual blocks, each residual block contains three convolutional layers, and the size of the filter for each layer is 3×3. The first three residual blocks use max pooling operations, while the last residual block uses average pooling operations. The linear classifier selects a Softmax classifier.

[0073] Specifically, in Step 3, the specific method is as follows:

[0074] Step 31: Randomly select a batch of images from the damaged auxiliary dataset. The number of images is N, and any one of the images is represented as d i (x), and its corresponding clear image is represented as I i (x), and its corresponding class label is represented as O i , and input d i (x) into the image restoration sub-network to obtain the restored image y i (x), then construct the following reconstruction loss function:

[0075]

[0076] Step 32: Input the output image y i (x) of the image restoration sub-network into the few-shot learner. The feature extractor represents it as the deep feature representation F i , and the classifier then converts F i into the output logical value T i , and based on T i and O i calculate the following cross-entropy loss function:

[0077]

[0078] where M is the total number of image classes in the auxiliary dataset, and O ij represents the j-th component of O i , and T ij represents the j-th component of T i ;

[0079] Step 33: Establish a joint loss function, and its mathematical expression is as follows:

[0080] Loss J = Loss C + αLoss R where α represents an adjustable parameter. Based on the above formula, use the gradient descent algorithm to optimize the parameters of the joint learning model for low-level vision and few-shot image classification.

[0081] Specifically, in Step 4, the specific method is as follows:

[0082] Step 41: Remove the Softmax classifier in the joint learning model for low-level vision and few-shot image classification, and fix the model parameters therein. Add an adapter in the joint learning model. The adapter is represented as A(·), and this adapter is composed of a multi-layer perceptron;

[0083] Step 42: Collect and organize the images of rare water birds. Any one of the images is represented as b i (x), and its corresponding class label is represented as Ui , input b i (x) into the trained joint learning model of low-level vision and few-shot image classification, and then input it into the adapter A(·) to obtain the output logical value V i , based on V i and U i calculate the following cross-entropy loss function:

[0084]

[0085] where Q is the total number of rare waterbird image categories, and U ij represents the j-th component of U i , and V ij represents the j-th component of V i ;

[0086] Based on the above formula, use the gradient descent algorithm to fine-tune and optimize the parameters of the adapter, and jointly use the trained joint learning model of low-level vision and few-shot image classification without the Softmax classifier and the adapter to complete the recognition of rare waterbirds under adverse weather conditions.

[0087] Example:

[0088] As Figure 1 shown, the specific steps are as follows:

[0089] Step 1: Collect and organize 200 images of each category in common birds and wetland plants to construct an auxiliary dataset. For any image in the auxiliary dataset, perform fogging degradation processing, rain degradation processing, and snow degradation processing on it respectively to obtain fogged image H(x), rained image R(x), and snowed image S(x). These fogged images, rained images, and snowed images together form a damaged auxiliary dataset.

[0090] Step 2: Construct a joint learning model for underlying vision and few-shot image classification. Among them, the image restoration sub-network adopts a U-shaped structure, mainly composed of two components: an encoder and a decoder, which is divided into 4 stages. The encoder and decoder in the same stage are connected by skip layers. The encoder in each stage is composed of a Transformer module and a downsampling layer, and the decoder in each stage is composed of a Transformer layer and an upsampling layer. A window-based multi-head attention mechanism is adopted in the Transformer module, which divides the feature map into several small windows spatially, and the calculation of multi-head attention is restricted within the windows. A 3×3 convolutional layer is used to perform convolution on the output features of the last stage of the decoder, and it is superimposed with the input degraded image to obtain the output of the image restoration sub-network. The few-shot learner is mainly composed of a feature extractor and a linear classifier. The feature extractor is composed of 4 residual blocks, each residual block contains three convolutional layers, and the size of the filter in each layer is 3×3. The first three residual blocks use max pooling operations, while the last residual block uses average pooling operations. The linear classifier selects a Softmax classifier.

[0091] Step 3: Randomly select a batch of images from the damaged auxiliary dataset, the number of images is N, and any one image is represented as d i (x), its corresponding clear image is represented as I i (x), and its corresponding class label is represented as O i . Input d i (x) into the image restoration sub-network to obtain the restored image y i (x), then construct the following reconstruction loss function Loss R . Input the output image y i (x) of the image restoration sub-network into the few-shot learner, and the feature extractor represents it as a deep feature representation F i , and the classifier then converts F i into the output logical value T i . Based on T i and O i , calculate the following cross-entropy loss function Loss C . Obtain the weighted sum of Loss R and Loss C to get the joint loss function. Use the gradient descent algorithm to optimize the parameters of the joint learning model for underlying vision and few-shot image classification.

[0092] Step 4: Remove the Softmax classifier in the joint learning model for underlying vision and few-shot image classification, and fix the model parameters therein. Add an adapter to the joint learning model, and this adapter is composed of a multi-layer perceptron. Collect and organize images of rare water birds, and any one image is represented as bi (x), whose corresponding class label is represented as U i . Input b i (x) into the trained joint learning model of low-level vision and few-shot image classification and the adapter to obtain the output logical value V i . Based on V i and U i Calculate the cross-entropy loss function, and use the gradient descent algorithm to fine-tune and optimize the parameters of the adapter. Combine the trained joint learning model of low-level vision and few-shot image classification without the Softmax classifier with the adapter to complete the recognition of rare water birds under adverse weather conditions.

[0093] In summary, the present invention can counter the influence of adverse weather conditions on the few-shot image classification task and improve the performance of the rare water bird recognition task in real scenarios.

[0094] The descriptions and practices disclosed in the present invention are easy to think about and understand for those of ordinary skill in the art. Without departing from the principle of the present invention, several improvements and retouches can be made. Therefore, the modifications or improvements made without deviating from the spirit of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A method for identifying rare water birds under adverse weather conditions, characterized in that: The steps include: Step 1: Construct a damaged auxiliary dataset of rare water birds: collect images of common birds and wetland environments to obtain an auxiliary image dataset for rare water bird identification, and degrade this dataset to obtain a damaged auxiliary dataset; Step 2: Build a joint learning model for underlying vision and small sample image classification: This model consists of an image complex atomic network and a small sample learner; Step 3: Optimize the joint learning model of underlying vision and small sample image classification: Establish a reconstruction loss between the output of the image complex atom network and the clear auxiliary image, and establish a cross entropy loss function between the output of the small sample learner and the classification label value. Use the weighted sum of the reconstruction loss and the cross entropy loss function as the total loss function to optimize the joint learning model. Step 4: Fine-tune the joint learning model of underlying vision and small sample image classification: Fix the parameters of the joint learning model and add an adapter, collect rare water bird images, fine-tune the adapter parameters, and complete the identification of rare water birds.

2. The method for identifying rare water birds under severe weather conditions according to claim 1, characterized in that: In step 1, the specific method is as follows: Step 11: Collect and organize 200 images of each category of common birds and wetland plants to build an auxiliary dataset; Step 12: For any image in the auxiliary dataset, denoted as I(x), perform a haze degradation process on it according to the following formula: H(x)=I(x)T(x)+A(1-T(x)) T(x)=exp {-βd(x)} Among them, H(x) represents the image of I(x) after being fogged, A represents the global atmospheric light coefficient, T(x) represents the medium transmission diagram, β represents the atmospheric scattering coefficient, and d(x) represents the distance from the camera to the object; Step 13: For any image in the auxiliary dataset, denoted as I(x), perform rain degradation processing on it according to the following formula: Among them, R(x) represents the image of I(x) after being rained, S l (x) represents the image of the rain streak layer at the first layer; Step 14: For any image in the auxiliary dataset, denoted as I(x), the snow degradation is performed according to the following formula: S(x)=I(x)(1-Z(x))+C(x)Z(x) Among them, S(x) represents the image of I(x) after being snowed, Z(x) represents the snowflake image, and C(x) represents the mask image; Step 15: The fog image H(x), the rain image R(x) and the snow image S(x) constitute the damaged auxiliary dataset.

3. The method for identifying rare water birds under adverse weather conditions according to claim 1, characterized in that: In step 2, the specific method is as follows: Step 21: The image complex atomic network adopts a U-shaped structure, which consists of two parts: the encoder and the decoder. The encoder and the decoder are divided into 4 stages. The encoder and the decoder in the same stage are connected by a jump layer; The encoder of each stage is composed of a Transformer module and a downsampling layer, and the decoder of each stage is composed of a Transformer layer and an upsampling layer. In order to reduce the computational complexity, a window-type multi-head attention mechanism is used in the Transformer module, which divides the feature map into several small windows in space, and the calculation of the multi-head attention is limited to the window. Assuming that the feature before the input of each Transformer module is represented as Z, the mathematical representation of its processing by the Transformer module is as follows: Among them, W Q represents the query value linear representation layer, W K represents the key-value linear representation layer, W V To represent the value linear representation, a 3×3 convolutional layer is used to convolve the output features of the final stage of the decoder and to superimpose the output of the image complex subnetwork with the input degraded image; Step 22: The small sample learner consists of a feature extractor and a linear classifier, where the feature extractor consists of 4 residual blocks, each residual block contains three convolutional layers, the size of each filter layer is 3×3, the first three residual blocks use the maximum pooling operation, and the last residual block uses the average pooling operation. The linear classifier selects the Softmax classifier.

4. The method for identifying rare water birds under adverse weather conditions according to claim 1, characterized in that: In step 3, the specific method is as follows: Step 31: Randomly select a batch of images from the damaged auxiliary dataset, the number of images is N, and any image is denoted as d i (x), and its corresponding clear image is denoted as I i (x), its corresponding category label is represented as O i , d i (x) is input into the image complex atom network to obtain the restored image y i (x), then construct the following reconstruction loss function: Step 32: Restore the image to the output image y of the sub-network i (x) is input into the small sample learner, and the feature extractor converts it into a deep feature representation F i , the classifier converts F i Converted into output logic value T i , based on T i and O i The following cross entropy loss function is calculated: Where M is the total number of image categories in the auxiliary dataset, O ij Indicates O i The jth component of ij Indicates T i The jth component of Step 33: Establish a joint loss function, whose mathematical expression is as follows: Loss J =Loss C +αLoss R Among them, α represents an adjustable parameter. Based on the above formula, the gradient descent algorithm is used to optimize the parameters of the underlying vision and small sample image classification joint learning model.

5. The method for identifying rare water birds under adverse weather conditions according to claim 1, characterized in that: In step 4, the specific method is as follows: Step 41: remove the Softmax classifier in the joint learning model of the underlying vision and small sample image classification, fix the model parameters therein, and add an adapter to the joint learning model. The adapter is denoted as A(·), and the adapter is composed of a multi-layer perceptron. Step 42: Collect and organize images of rare water birds, where any image is represented by b i (x), its corresponding category label is denoted as U i , b i (x) is input into the trained underlying vision and small sample image classification joint learning model, and then input into the adapter A(·) to obtain the output logical value V i , based on V i and U i The following cross entropy loss function is calculated: Among them, Q is the total number of rare water bird image categories, U ij Indicates U i The jth component of ij Indicates V i The jth component of Based on the above formula, the gradient descent algorithm is used to fine-tune and optimize the parameters of the adapter. The trained underlying vision and small sample image classification joint learning model without the Softmax classifier is used in conjunction with the adapter to complete the identification of rare water birds under severe weather conditions.