Cross-scene small sample animal key point detection method based on deformable Mama
By introducing deformable Mamba and bidirectional SSM in animal key point detection, combined with contrast learning and feature modulator, the problem of insufficient generalization ability under small samples and cross-scene conditions is solved, and higher detection accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202510346403.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The existing animal key point detection model lacks generalization ability under small samples and cross-scene conditions, making it difficult to identify new species outside the training samples and key points in different scenarios.
A cross-scene small sample animal key point detection method based on deformable Mamba is adopted, and a deformable Mamba encoder and bidirectional state space model (SSM) that adaptively adjusts the size of local receptive fields, combined with contrast learning and feature modulator, the features of animal images are extracted and the key point location is predicted.
It significantly improves the accuracy and generalization ability of key point detection, and can accurately identify unknown species and key points in different scenarios under small sample conditions, reduces GPU memory requirements and improves inference speed.
Smart Images

Figure CN120220190A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cross-scenario few-shot animal key point detection method based on deformable Mamba, belonging to the field of computer vision. Background Art
[0002] Animal key point detection is an important research topic in the field of computer vision. Its core lies in providing concise semantic and structural information by detecting key points, enabling this technology to be widely applied in multiple directions such as human pose estimation, animal pose estimation, and behavior analysis. However, due to the wide variety of animal species and the difficulty in comprehensively collecting data samples, most existing key point detection models can only recognize the species involved during training and lack effective recognition ability for new species outside the training samples. This limitation is particularly prominent in few-shot cross-scenario animal detection. Therefore, there is an urgent need for new methods to improve the generalization performance of the model in unknown species and different scenarios.
[0003] Currently, there are mainly two methods for key point localization: one is to directly regress the key point coordinates, and the other is to decode the coordinates after generating a heat map. Since the heat map regression method utilizes the intuitiveness of spatial mapping and usually can achieve higher detection accuracy, it provides a clear technical basis for subsequent optimization of the few-shot problem and also reveals the advantages and disadvantages of different methods in fine-grained localization.
[0004] In the case of scarce data, few-shot learning can achieve good model performance using limited labeled data. Common few-shot learning methods mainly include three categories: model fine-tuning, data augmentation, and transfer learning. The method based on model fine-tuning first pre-trains the model using a large-scale dataset and then fine-tunes it on a small amount of data. Although this method can quickly adapt to new tasks, it is prone to overfitting problems when the data is insufficient. The data augmentation method alleviates overfitting by expanding the training data, but noise may also be introduced during the augmentation process, which has an adverse effect on the model performance. The method based on transfer learning covers three strategies: metric learning, meta-learning, and graph neural networks, each exploring the internal relationship between data instances from different perspectives. Although certain progress has been made in their respective fields, there are still many challenges in the few-shot key point detection task.
[0005] Regarding the small sample problem in key point detection, existing methods can generally be divided into two categories: one is small sample key point detection for specific categories, which uses a large amount of unlabeled data and a small number of labeled samples to achieve key point localization. This method is suitable for specific key point types. However, when dealing with different unknown species, due to differences in the number and distribution of key points, its generalization ability is limited. The other is a general key point detection method inspired by small sample learning, which adopts an episode training and evaluation strategy. Given support key point cues, the model can flexibly detect any key points in the query image, thus having stronger generalization ability. However, it also poses higher requirements for task design and modeling of the relationship between samples.
[0006] In recent years, the State Space Model (SSM) has received extensive attention because it is good at capturing long-range dependencies and supports parallel training. The Mamba model proposed in 2023 achieved efficient training and inference by introducing time-varying parameters and combining hardware-aware algorithms. Subsequently, in 2024, Vision Mamba successfully transferred Mamba from language modeling to the field of visual modeling. It uses a bidirectional SSM to capture global visual context and combines position embeddings to achieve precise position information perception. It shows superior performance compared to Transformer-based models (such as DeiT) in the ImageNet classification task. At the same time, it also has lower GPU memory requirements and faster inference speed in high-resolution image processing. The further developed deformable Mamba significantly improves the accuracy of key point detection by adaptively adjusting the size of the local receptive field and using Vision Mamba (Visual Mamba) to construct global long-range dependencies. Summary of the Invention
[0007] The purpose of the present invention is to provide a cross-scene small sample animal key point detection method based on deformable Mamba, aiming to solve the technical problems of low key point detection accuracy caused by limited training sample quantity in current detection methods and the inability to identify key points of other species except known species in training samples.
[0008] To achieve the above purpose, the present invention is implemented by the following technical solutions: A cross-scene small sample animal key point detection method based on deformable Mamba, the specific steps are as follows:
[0009] Step1: Use an animal dataset to label the animals and animal key point data in each picture;
[0010] Step2: Divide the labeled animal dataset into a support image set and a query image set, and generate saliency maps for the divided support image set and query image set through a saliency image generator;
[0011] Step 3: Design a cross-scenario small-sample animal key point detection network structure, where the cross-scenario means detecting animal key points in indoor and outdoor scenarios;
[0012] Step 4: Use the divided labeled animal dataset and its saliency map as inputs to train the cross-scenario small-sample animal key point detection network, and finally generate a cross-scenario small-sample animal key point detector.
[0013] The saliency map is specifically:
[0014] An image where the pixel value of each pixel point in the image is only 0 or 1, the pixel value of the foreground area is 1, and the pixel value of the background area is 0.
[0015] The specific content of Step 3 is:
[0016] Step 3.1: Input the divided support image set, query image set and their saliency maps into the deformable Mamba encoder to generate support feature maps and query feature maps; use the marked key points to extract the feature representation of each key point through Gaussian pooling;
[0017] Step 3.2: Use a feature modulator to associate the extracted key point feature representation with the query feature map, and use contrastive learning to extract the attention features related to the key point representation in the query feature map; among them, the related attention features refer to the local area in the query feature map where the semantic similarity with the key point feature is greater than a preset threshold and there is an interaction, and the local area is determined by the feature modulator and contrastive learning;
[0018] Step 3.3: After feature association, project each attention feature into a key point descriptor through a descriptor extractor to achieve dimensionality reduction, and the key point descriptor refers to a feature vector containing the local area information of the key point;
[0019] Step 3.4: Based on the dimension-reduced features, predict the position and uncertainty of the key points through a multi-scale localization network.
[0020] The specific operation of using a feature modulator to associate the extracted key point feature representation with the query feature map is:
[0021] Perform contrast between positive key points, where positive key points refer to the true key points of the animal, and obtain the contrast loss between positive key points through cosine similarity;
[0022] Perform contrast between positive key points and negative key points, where negative key points refer to the noise recognized as non-true key points, and obtain the loss of contrast between positive and negative key points through cosine similarity;
[0023] Based on the contrast loss between positive key points and the loss of contrast between positive and negative key points, the total contrast learning loss is obtained.
[0024] Specifically, Step 3.4 is as follows:
[0025] Step 3.4.1: Use a grid classifier to calculate the probability of each grid in the image. Among them, the grids closer to the key points have higher probabilities, and calculate the classification loss of the key points based on cross-entropy;
[0026] Step 3.4.2: Use an offset regressor to extract the offset vectors of the key points;
[0027] Step 3.4.3: Use a covariance branch to learn the covariance of each key point. Its output is a potential covariance field to model the uncertainty of single or multiple key points;
[0028] Step 3.4.4: Combine the offset vectors and the covariance information of the key points, and calculate the regression loss of the key point offsets through a multivariate Gaussian distribution;
[0029] Step 3.4.5: The total loss L between the predicted output and the true output of the cross-scene few-shot animal key point detection network is obtained as follows:
[0030] L = λ1L os-nll + λ2L cls + λ3L hm + λ4L CL
[0031] where L os-nll is the regression loss of the key point offsets; L cls is the classification loss of the key points, L hm is the heatmap regression loss, which is obtained by calculating the difference between the predicted heatmap and the true Gaussian distribution through mean square error or cross-entropy loss, L CL is the total contrast learning loss, and λ i represents the weights corresponding to each loss function, i = 1, 2, 3, 4.
[0032] Specifically, the training of the cross-scene few-shot animal key point detection network is as follows:
[0033] Take the support image and the query image as an episode and then as the input of the network. Each episode contains multiple support images and query images, and the animal types in the support images and query images are different;
[0034] The labeled key points that support the image only cover a part of the animal's body. First, the image is processed through a Deformable Vision Mamba module; then, by calculating the similarity of each key point in the support image and the query image, new key points of the new species are inferred.
[0035] The specific process of processing the image through the Deformable Mamba module is as follows:
[0036] First, the receptive field size of the local area is adaptively modified according to the image content so that the model can specifically process the animal when detecting key points, and then the Vision Mamba module is used to construct the long-range dependence between global features, so as to realize the adaptive local and global features according to the image content.
[0037] The beneficial effects of the present invention compared with the prior art are:
[0038] (1) The present invention extracts the features of the image by introducing Deformable Vision Mamba. First, the receptive field size of the local area is adaptively modified according to the image content so that the model can better process the animal when detecting key points, and then the Vision Mamba module is used to construct the long-range dependence between global features, which greatly improves the detection accuracy in key point detection; and it uses bidirectional SSM for data-related global visual context modeling and position embedding for position-aware visual recognition, can process pictures in linear time, can effectively capture long-range relationships, and is optimized for different inputs, so that the detection network can more accurately locate the key points and the mutual relationships between key points;
[0039] (2) The present invention solves the problem of weak generalization ability caused by few samples in small-sample key point detection through contrast learning and similarity calculation, reduces the GPU memory required during training, and improves the accuracy of small-sample key point detection. First, through contrast learning, animals different from the original animal species learn similar key points, so that the key point detection of the original species can be generalized to other unknown species; then, calculate the similarity of each region of the new species image to infer new key points similar to the existing key points; through the above two methods, the key point detection of the original species can be generalized to the key point detection of different unknown species. Description of the Drawings
[0040] Figure 1 is a schematic flowchart of the method for detecting key points of small-sample animals of the present invention;
[0041] Figure 2 is a structural diagram of the small-sample animal key point detection network based on Deformable Mamba of the present invention;
[0042] Figure 3 It is a schematic diagram of the RGB image of animals and the generated saliency map in the indoor and outdoor scenarios of the present invention;
[0043] Figure 4 It is a structural diagram of the multi-scale localization network of the present invention for predicting key points;
[0044] Figure 5 It is a detailed structural diagram of the deformable Mamba encoder of the present invention. Detailed implementation manners
[0045] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.
[0046] Example 1: As Figure 1 shown, a cross-scene few-shot animal key point detection method based on deformable Mamba includes the following steps:
[0047] Step1: Use an animal dataset to label the animals and animal key point data in each picture.
[0048] Specifically, the animal dataset has seven categories, namely cat, dog, sheep, cow, horse, elephant, and tiger;
[0049] Furthermore, the labeled data is the animal position information and the key point information on the animal body, such as the left eye, right eye, nose, left hand, etc.
[0050] Step2: Divide the labeled animal dataset into a support image set and a query image set, and generate saliency maps for the divided support image set and query image set through a saliency image generator.
[0051] Specifically, the present invention generates the saliency map of each picture through the SCRN model. The saliency map means that the pixel value of each pixel point in the image can only be 0 or 1, and the pixel value of the foreground area (the area containing the target) is 1, and the pixel value of the background area (the area not containing the target) is 0.
[0052] Step3: Design a cross-scene few-shot animal key point detection network structure. The cross-scene refers to detecting animal key points in indoor and outdoor scenarios, such as Figure 2 shown as the few-shot animal key point detection network structure diagram, and as Figure 3 shown as the RGB image of animals and their saliency images in indoor and outdoor scenarios.
[0053] Step 3.1: Input the divided support image set, query image set, and their saliency maps into the deformable Mamba encoder to generate support feature maps and query feature maps; use the marked key points to extract the feature representations of each key point through Gaussian pooling;
[0054] Step 3.2: Use the feature modulator to associate the extracted key point feature representations with the query feature maps, and use contrastive learning to extract the attention features related to the key point representations in the query feature maps; among them, the related attention features refer to the local regions in the query feature maps where the semantic similarity with the key point features is greater than the preset threshold and there is an interaction, and the local regions are determined by the feature modulator and contrastive learning;
[0055] Specifically, the use of the feature modulator to associate the extracted key point feature representations with the query feature maps is as follows:
[0056] Perform contrast between positive key points, where positive key points refer to the true key points of animals, such as the eyes of a cat and the eyes of a dog, the feet of a horse and the nose of a cat, etc. These are all contrasts between positive key points, and the contrast loss between positive key points is J(T s ,T s' ), as shown in the following formula:
[0057]
[0058] Among them, s and s′ respectively represent two different species; T s represents the average key point feature of one species; T s' represents the average key point feature of another species; cos(.,.) represents the cosine similarity between two key points being contrasted, m i and m i ' represent N types of key points of two different species, i = 1…N;
[0059] Perform contrast between positive key points and negative key points, where negative key points refer to the noises identified as non - true key points, which are used to better correctly locate the key points in the image, such as some points around the mouth of a cat and the mouth of a dog or some points on the body, etc. These are the contrasts between positive and negative key points, and the negative key points are represented as T neg , then the contrast loss between positive and negative key points is J(T s ,T neg ), as shown in the following formula:
[0060]
[0061] Among them, represents the j - th type of negative key point, j = 1…N;
[0062] The total loss generated by contrastive learning is L CL , as shown in the following formula:
[0063] L CL = - <Ⅱ, log(softmax(J(T s , T s′ ∪ T neg )) / τ)>
[0064] Among them, Ⅱ represents the identity matrix; <.,.> represents the inner product; softmax represents a function used to map the element values to between 0 and 1, and the sum of the element values is 1; τ represents the temperature parameter, which is used to control the smoothness of the softmax distribution.
[0065] Step3.3: After feature association, each attention feature is projected into the key point descriptor by a descriptor extractor to achieve dimensionality reduction, and the key point descriptor refers to a feature vector containing the local region information of the key point;
[0066] Step3.4: Based on the dimension-reduced features, the position and uncertainty of the key points are predicted through a multi-scale localization network.
[0067] Step3.4.1: Use a grid classifier to calculate the probability of each grid in the image. Among them, the grid closer to the key point has a higher probability, and the classification loss L cls of the generated key point is as shown in the following formula:
[0068] L cls = - Σ i y i log p i
[0069] Among them, p i represents the probability that the prediction target is the i-th category, and y i represents whether the true category is the i-th category;
[0070] Step3.4.2: Use an offset regressor to extract the offset vector of the key point;
[0071] Step3.4.3: Use a covariance branch to learn the covariance of each key point, and its output is a potential covariance field to model the uncertainty of a single or multiple key points;
[0072] Step3.4.4: Combine the offset vector and the covariance information of the key point, and calculate the regression loss L os-nll of the key point offset through a multivariate Gaussian distribution, as shown in the following formula:
[0073]
[0074] Among them, E is the mathematical expectation, T represents the matrix transpose; x is the predicted offset of the key point; v * is the true position of the key point; Ω = Σ -1 represents the inverse matrix of the covariance matrix, ∑ is the covariance matrix; det(..) represents the determinant of the matrix, as Figure 4 shown is the structural diagram of the multi-scale localization network for predicting key points;
[0075] Step3.4.5: The total loss L between the predicted output and the true output of the cross-scenario small-sample animal key point detection network is:
[0076] L = λ1L os-nll + λ2L cls + λ3L hm + λ4L CL
[0077] Among them, L hm is the heatmap regression loss, which is obtained by calculating the difference between the predicted heatmap and the true Gaussian distribution through mean square error or cross-entropy loss, λ i represents the weight corresponding to each loss function, i = 1, 2, 3, 4.
[0078] Step4: Use the divided labeled animal dataset and its saliency map as inputs to train the cross-scenario small-sample animal key point detection network, and finally generate a cross-scenario small-sample animal key point detector.
[0079] Specifically, the training of the cross-scenario small-sample animal key point detection network is as follows: Take the support image and the query image as an episode, and then use them as the input of the network. Each episode contains multiple support images and query images, and the animal types in the support image and the query image are different. Further, select several categories in the animal dataset as the support image set, and the remaining categories as the query image set;
[0080] The labeled key points of the support image only cover a part of the animal's body. First, process the image through the deformable Mamba module; then, by calculating the similarity of each key point in the support image and the query image, infer the new key points of the new species. For example, if the support image is a cat and the labeled key points are the left eye, left leg, left hand, etc., and the query image is a dog, through this network, the left and right eyes, left and right hands, left and right legs, etc. of the dog can be detected. Through this network, the problem of weak generalization ability caused by few samples in small-sample learning can be improved.
[0081] Furthermore, the processing of the image by the deformable Mamba module is as follows: first, adaptively modify the receptive field size of the local region according to the image content so that the model can specifically process animals when detecting key points, and then construct long-range dependencies between global features through the VisionMamba module, greatly improving the detection accuracy in key point detection, thereby realizing the adaptation of local and global features according to the image content, such as Figure 5 The detailed structure diagram of the deformable Mamba encoder is shown as follows.
[0082] Furthermore, the VisionMamba module is an image temporal model. It uses bidirectional SSM for data-related global visual context modeling and position embedding for position-aware visual recognition. It can process images in linear time, effectively capture long-range relationships, and optimize for different inputs, enabling the detection network to more accurately locate key points and the relationships between key points.
[0083] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A cross-scene small sample animal key point detection method based on deformable Mamba, characterized in that: The method comprises the following steps: Step 1: Use an animal dataset to mark the animals and animal key point data in each picture; Step 2: Divide the labeled animal dataset into a support image set and a query image set, and generate salient maps of the divided support image set and query image set through a salient image generator; Step 3: Design a cross-scene small sample animal key point detection network structure, where cross-scene refers to detecting animal key points in indoor and outdoor scenes; Step 4: Use the divided labeled animal dataset and its saliency map as input to train a cross-scene small-sample animal key point detection network, and finally generate a cross-scene small-sample animal key point detector.
2. The method for detecting key points of small samples of animals across scenes based on deformable Mamba according to claim 1 is characterized in that: The saliency map is specifically: An image in which each pixel has a value of 0 or 1, with the foreground pixel value being 1 and the background pixel value being 0.
3. The method for detecting key points of small samples of animals across scenes based on deformable Mamba according to claim 1 is characterized in that: The Step 3 is specifically as follows: Step 3.1: Input the divided support image set and query image set and their saliency maps into the deformable Mamba encoder to generate support feature maps and query feature maps; use the marked key points to extract the feature representation of each key point through Gaussian pooling; Step 3.2: Use a feature modulator to associate the extracted key point feature representation with the query feature map, and use contrastive learning to extract attention features related to the key point representation in the query feature map; wherein the related attention features refer to local areas in the query feature map whose semantic similarity with the key point feature is greater than a preset threshold and interaction, and the local areas are determined by the feature modulator and contrastive learning; Step 3.3: After feature association, each attention feature is projected into a keypoint descriptor through a descriptor extractor to achieve dimensionality reduction. The keypoint descriptor refers to a feature vector containing the local area information of the keypoint; Step 3.4: Based on the reduced-dimensional features, the location and uncertainty of the key points are predicted through a multi-scale positioning network.
4. The method for detecting key points of small samples of animals across scenes based on deformable Mamba according to claim 3 is characterized in that: The method of using the feature modulator to associate the extracted key point feature representation with the query feature map is specifically as follows: Compare the positive key points, where the positive key points refer to the real key points of the animal, and obtain the comparison loss between the positive key points through cosine similarity; Compare positive key points with negative key points, where negative key points refer to noise that is identified as non-real key points, and the loss of the comparison between positive and negative key points is obtained through cosine similarity; Based on the contrast loss between positive key points and the loss of contrast between positive and negative key points, the total loss of contrastive learning is obtained.
5. The method for detecting cross-scene small sample animal key points based on deformable Mamba according to claim 3 is characterized in that: The Step 3.4 is specifically as follows: Step 3.4.1: Use a grid classifier to calculate the probability of each grid in the image, where the closer the grid is to the key point, the higher the probability is. The classification loss of the key point is calculated based on the cross entropy; Step 3.4.2: Use an offset regressor to extract the offset vector of the key point; Step 3.4.3: Use a covariance branch to learn the covariance of each key point, whose output is a potential covariance field to model the uncertainty of single or multiple key points; Step 3.4.4: Combine the offset vector and the covariance information of the key point, and calculate the regression loss of the key point offset through multivariate Gaussian distribution; Step 3.4.5: The total loss L of the predicted output and the true output of the cross-scene small sample animal key point detection network is: L=λ1L os-nll +λ2L cls +λ3L hm +λ4L CL Among them, L os-nll is the key point offset regression loss; L cls is the classification loss of the key points, L hm is the heatmap regression loss, which is obtained by calculating the difference between the predicted heatmap and the true Gaussian distribution through the mean square error or cross entropy loss. CL is the total loss of contrastive learning, λ i Represents the weight corresponding to each loss function, i=1,2,3,4.
6. The method for detecting key points of small samples of animals across scenes based on deformable Mamba according to claim 1 is characterized in that: The training of the cross-scene small sample animal key point detection network is specifically as follows: The support images and query images are taken as an episode and then used as the input of the network. Each episode contains multiple support images and query images, and the animal types in the support images and query images are different. The annotated keypoints of the support image only cover part of the animal's body. The image is first processed by the deformable Mamba module; then new keypoints for the new species are inferred by calculating the similarity of each keypoint in the support image and the query image.
7. The method for detecting small sample animal key points across scenes based on deformable Mamba according to claim 6 is characterized in that: The processing of the image by the deformable Mamba module is specifically as follows: First, the receptive field size of the local area is adaptively modified according to the image content so that the model can process the animals in a targeted manner when detecting key points. Then, the Vision Mamba module is used to build long-distance dependencies between global features, thereby achieving adaptive local and global features according to the image content.