A new method for identifying golden monkey individuals
Through a new self-supervised individual recognition method based on twin networks and deep feature clustering algorithms, the problem of identifying individuals of golden snub-nosed monkeys without identity information is solved, and the accurate identification of label-free golden snub-nosed monkey data is achieved and the recognition ability of the model is improved.
Patent Information
- Application Number
- CN202210786124.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The prior art is difficult to identify different golden monkey individuals without identity information, especially when new individual data lacks label information, deep learning models lack the ability to construct data and understand the concept of object category.
A new self-supervised individual recognition method based on twin networks and deep feature clustering algorithms is adopted to identify the golden monkey data without identity information by introducing the golden monkey facial image pre-training encoder module, the category number estimation module and the labelless data clustering module.
The accurate identification of labelless golden monkey image data is achieved, which improves the accuracy of new individual recognition, and improves the model's recognition ability of new individuals through data enhancement.
Smart Images

Figure CN115331134B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer application technology, relates to image processing and deep learning methods, and specifically relates to a method for identifying new golden monkey individuals. Background Art
[0002] Golden snub-nosed monkeys have a complex social structure, and their behavioral research needs to be based on the ability to distinguish individuals. Among them, facial recognition of golden snub-nosed monkeys is an important prerequisite for the study of golden snub-nosed monkey behavior. However, individual identification based on golden snub-nosed monkey facial features faces many challenges, including: the facial features of different golden snub-nosed monkeys are very similar and the identity of new individuals is difficult to accurately mark and identify. At present, deep neural networks can identify thousands of wild animals, but they need to rely on auxiliary information of wild animals, such as animal labels and descriptions of appearance and body shape. If there is a lack of image-like information, deep learning lacks the ability to construct data and understand the concept of object categories without external supervision. How can the network recognize different golden snub-nosed monkey individuals without identity information? Unlike humans, neural networks do not know what objects they are identifying. The network extracts key information from the image, which can also be called the attributes of things (such as eye information, hair information, etc.). These key information are the basis for it to distinguish objects. Previous recognition algorithms required manual description of these attributes and converted them into machine language to assist the network in recognition. However, due to the characteristics of golden snub-nosed monkey facial images, it is quite difficult to manually describe the attributes of each golden snub-nosed monkey individual, because golden snub-nosed monkey individuals are extremely similar, resulting in little difference between the description content. In addition, the collection of golden snub-nosed monkey data spans a long period of time, and new golden snub-nosed monkey data (hereinafter referred to as "new individuals") will be collected at different stages. These new individuals cannot be directly identified using previously trained models because new individuals are not repeated with already marked individuals.
[0003] Although supervised learning methods have achieved remarkable results in previous studies and have been applied in many fields. However, in supervised learning, each image needs to be labeled so that the network can learn. In addition, the learned network model can only recognize the trained individual data of golden monkeys, and lacks the ability to accurately recognize unseen image data. Another situation is that in practical applications, individual classes in unlabeled data may not have enough training data, which causes the model to be more inclined to classes with more training samples. In this context, it is of great practical significance to use relevant theoretical knowledge in fields such as image processing and deep learning to develop a twin network and deep feature clustering algorithm to solve these problems. Summary of the invention
[0004] Based on the existing primate identification, a large amount of identity information is required for supervised training. The purpose of the present invention is to provide a new individual identification method for golden monkeys to solve the problem that some image data with missing identity information and wild primate data with no identity information in practice ultimately lead to the inability to perform identity identification.
[0005] In order to achieve the above tasks, the present invention adopts the following technical solutions:
[0006] A method for identifying new individuals of golden snub-nosed monkeys is characterized in that the method adopts a self-supervised new individual identification algorithm based on a twin network and deep feature clustering, and introduces a golden snub-nosed monkey facial image pre-training encoder module, a golden snub-nosed monkey category quantity estimation module, and an unlabeled data clustering module. New individuals are identified for golden snub-nosed monkey data without identity information, specifically comprising the following steps:
[0007] Step S1: firstly, use a high-resolution camera to collect videos of golden monkeys;
[0008] Step S2: Use the image quality assessment algorithm MUSIQ to evaluate the collected golden monkey images, and filter out images with severe blur and jitter;
[0009] Step S3: Use labelme software to label the golden monkey face, and use the target detection network yolov5 to perform monkey face detection on the image containing the golden monkey face information; then segment the located monkey face from the image to obtain an image containing only the golden monkey face.
[0010] Step S4: performing data preprocessing on the detected golden monkey face image, cutting the image to a certain size and performing monkey face correction processing;
[0011] Step S5: The high-quality golden monkey face image pre-trained encoder module is used to extract the visual features of the golden monkey individuals, reduce the ambiguity of clustering through the prior knowledge learned from the labeled data, and improve the accuracy of the new individual division;
[0012] Step S6: The facial images of golden monkeys are highly similar, and considering the small number of images of certain classes in the dataset, this module uses the structure of the Siamese network to learn facial visual features from the data;
[0013] Step S7: a category number estimation module is used to estimate the category number of new golden monkey individuals by using the prior knowledge of labeled data to estimate the category number of unlabeled data through the knowledge transfer method;
[0014] Step S8: Cluster analysis module, clustering the unlabeled data through clustering algorithm to achieve the purpose of individual identification;
[0015] Step S9: Add new golden monkey individuals with high recognition accuracy to the original data set for data enhancement.
[0016] According to the present invention, the specific implementation steps of step S2 include:
[0017] Step S21: First, a multi-scale representation of the golden monkey image is obtained, including the original image and a fixed aspect ratio resized (ARP (aspect ratio preserved) resized) variant. Images of different scales are divided into fixed-size image blocks and then fed into a pre-trained image quality assessment model. Since the image blocks are from images of different spatial resolutions, it is necessary to efficiently encode these inputs of multiple aspect ratios and scales into a token sequence to capture pixel, spatial and scale information.
[0018] Step S22: Based on hashing, the position of a certain image block is recorded in the i-th row and j-th column, and is hashed to the corresponding element in the G×G grid; each element in the grid is a D-dimensional embedding vector; that is, there is a learnable matrix The input image size is H, W, and the image is divided into multiple image blocks of size P. For the image block at position (, j), its spatial embedding is defined as (t i ,t j ) position element;
[0019]
[0020] Step S23: reuse the same hash matrix for the used images. HSE cannot distinguish image blocks from different scales, so an additional scale embedding SCE is introduced to help the model distinguish image blocks from different scales.
[0021] Furthermore, the specific implementation steps of step S3 include:
[0022] Step S31: Use the labelme annotation tool to finely annotate the monkey face of the golden monkey image data, so that the target detection network can accurately obtain the image coordinate information of the monkey face.
[0023] Step S32: Feed the finely annotated data for training the detection network into the target detection network YOLO V5, update the network parameters through back propagation until the network converges, and finally complete the pre-training of the target detection network YOLO V5;
[0024] Step S33: Use the pre-trained model to crop the monkey face part of the unlabeled golden monkey image data to prepare for the later estimation of the number of golden monkey categories and self-supervised identification of new individuals.
[0025] Preferably, the specific implementation steps of step S4 include:
[0026] Step S41: The golden monkey facial data obtained in step S3 is cropped to the same size of 128*128 to facilitate subsequent model processing.
[0027] Step S42: Correct the golden monkey face image based on affine transformation, abstract the projection transformation of the golden monkey face during rotation and convert it into a geometric model. Analyze the specific changes of the monkey face image from the geometric model and analyze the projection plane image of the human face under rotation in combination with the affine transformation theory, so as to achieve correction of the monkey face image.
[0028] Preferably, the specific implementation steps of step S6 include:
[0029] Step 61: Input the high-quality golden monkey image data into the encoder for pre-training. The pre-training module uses the ResNet-50 deep neural network as the backbone network.
[0030] Step 62: Obtain the deep feature vectors of the two golden monkey images through the Siamese network encoder, and calculate the cosine distance between the two feature vectors;
[0031] Step 63: Continue to iterate the SS-NIR network and update the network parameters through back propagation until the network converges.
[0032] Further preferably, the specific implementation steps of step S9 include:
[0033] Step S71: using the prior knowledge of labeled data, clustering the labeled data and unlabeled data sets multiple times using K-means, and continuously estimating and updating the number of classes of unlabeled data;
[0034] Step S72: Calculate two quality indexes to check the obtained clustering results, one of which is to check the quality of the labeled data, the number of its categories is known;
[0035] Step S73: The number of categories of unlabeled data is estimated as the one that maximizes the two quality indexes.
[0036] Preferably, the specific implementation steps of step S9 include:
[0037] Step S81: The number of categories of the unlabeled data is estimated by the category quantity estimation module in step S7, and then the new individual data is labeled using a clustering algorithm;
[0038] Step S82: Update the network parameters of ResNet-50 by alternately training ResNet-50 and K-means clustering algorithm;
[0039] Step S83: Classify the images of the same golden monkey individual into one category and assign a pseudo label to each image.
[0040] The self-supervised new individual recognition algorithm based on the twin network and deep feature clustering algorithm adopted by the present invention introduces three modules to estimate the number of categories of unlabeled golden monkey image data, and clusters the golden monkey data without identity information using a clustering algorithm according to the estimated number of categories, and finally obtains the pseudo-label of the unlabeled golden monkey image data set. The frame images collected in the video segment are extracted at equal intervals, input into the image quality assessment network MUSIQ, and low-quality images are screened out; the high-quality golden monkey images are then passed through the target detection network yolo v5 to extract the facial image of the golden monkey; the golden monkey facial image based on the affine transformation is corrected, and the projection transformation of the golden monkey face during the rotation process is abstracted and converted into a geometric model. The specific changes of the monkey face image are parsed from the geometric model and the projection plane image of the human face under rotation is analyzed in combination with the affine transformation theory, so as to realize the correction of the monkey face image. Then the self-supervised new individual recognition algorithm based on the twin network and deep feature clustering algorithm is used to assign labels to the golden monkey image data to realize the identity recognition of the golden monkey.
[0041] Compared with the existing technology, the method for identifying new golden monkey individuals of the present invention has the following technical innovations: three modules are designed to estimate the number of categories of unlabeled golden monkey image data, and a clustering algorithm is used to cluster the golden monkey data without identity information according to the estimated number of categories, and finally a pseudo-label of the unlabeled golden monkey image data set is obtained. Further, new primate individuals with high recognition accuracy are added to the original data set for data enhancement to improve the model's recognition ability for new individuals. The accuracy of this method is better than that of the existing unlabeled clustering algorithm. It can be widely used in the identification of unlabeled primates. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the golden monkey face positioning;
[0043] Figure 2 : This is a framework diagram of a network for self-supervised new individual identification (SS-NIR) based on a twin network and a deep feature clustering algorithm adopted in the present invention; wherein a, b and c represent three modules respectively, which are executed in sequence from a to c;
[0044] Figure 3 It is the encoder pre-training module;
[0045] Figure 4 It is the framework diagram of cluster analysis module;
[0046] Figure 5 This is a diagram showing the clustering effect.
[0047] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. DETAILED DESCRIPTION
[0048] Golden monkey facial recognition is an important prerequisite for the study of golden monkey behavior recognition, but individual identification based on golden monkey facial features faces many challenges, including: the similarity of facial features of different golden monkey individuals and the difficulty in accurately marking and identifying new individuals. To this end, this paper proposes a hierarchical ensemble network (HE-Nets) and a self-supervised new individual recognition algorithm (SS-NIR) based on a twin network and a deep feature clustering algorithm. The main research contents are as follows:
[0049] For the problem of identifying new golden snub-nosed monkey individuals, the image data of new individuals cannot be directly identified using traditional individual identification models. It is necessary to re-label and update network parameters, which is time-consuming and labor-intensive. The present invention proposes SS-NIR to solve the problem of identifying new golden snub-nosed monkey individuals. Due to the lack of label information, the new individuals are identified simply by relying on the network's own learning ability, and the effect is not ideal. To this end, SS-NIR uses the prior knowledge of the labeled data set to learn and find which features in the labeled data can form a cluster well, so as to reduce the ambiguity of unlabeled data clustering. The present invention first uses the twin network structure to train a deep neural network as a feature encoder to learn the feature distribution of labeled data. Next, the number of new golden snub-nosed monkey individuals is estimated using the prior knowledge of the labeled data set. Experiments show that the error between the category estimate and the true value is less than 3. After adding the estimated number of individuals, the new golden snub-nosed monkey individual data is clustered using the method of alternating training of ResNet-50 and K-means to complete the identity identification of the new golden snub-nosed monkey individuals. Finally, the individuals with high accuracy in the clustering results are added to the original golden snub-nosed monkey data set.
[0050] See also Figures 1 to 5This embodiment provides a method for identifying new individuals of golden snub-nosed monkeys. The method adopts a self-supervised new individual identification algorithm based on a twin network and a deep feature clustering algorithm, and designs a method for estimating the number of categories of unlabeled golden snub-nosed monkey image data. The method introduces a golden snub-nosed monkey facial image pre-training encoder module, a golden snub-nosed monkey category number estimation module, and an unlabeled data clustering module. First, the number of categories of unlabeled golden snub-nosed monkey image data is estimated, and then the golden snub-nosed monkey data without identity information is clustered using a clustering algorithm based on the estimated number of categories. Specifically, the method includes the following steps:
[0051] Step S1: firstly, use a high-resolution camera to collect images of golden monkeys;
[0052] Step S2: Use the image quality assessment algorithm MUSIQ to evaluate the collected golden monkey images, and filter out images with severe blur and jitter;
[0053] Step S3: label the golden monkey faces using labelme software, and perform monkey face detection on the image containing the golden monkey face information using the target detection network yolov5; then segment the located monkey faces from the image to obtain an image containing only the golden monkey faces.
[0054] Step S4: performing data preprocessing on the detected golden monkey face image, cutting the image to a certain size and performing monkey face correction processing;
[0055] Step S5: The high-quality golden monkey face image pre-trained encoder module is used to extract the visual features of the golden monkey individuals. The ambiguity of clustering is reduced by learning the prior knowledge from the labeled data, thereby improving the accuracy of the new individual division.
[0056] Step S6: The facial images of golden monkeys are highly similar, and considering the small number of images of certain classes in the dataset, this module uses the structure of the Siamese network to learn facial visual features from the data;
[0057] Step S7: a category number estimation module is used to estimate the category number of new golden monkey individuals by using the prior knowledge of labeled data to estimate the category number of unlabeled data through the knowledge transfer method;
[0058] Step S8: Cluster analysis module, clustering the unlabeled data through clustering algorithm to achieve the purpose of individual identification;
[0059] Step S9: Add new golden monkey individuals with high recognition accuracy to the original data set for data enhancement.
[0060] It should be noted that: in the following embodiments, deep neural networks can identify thousands of wild animals, but they need to rely on auxiliary information of wild animals, such as animal labels and descriptions of appearance and body shape. If there is a lack of information similar to the image, deep learning lacks the ability to construct data and understand the concept of object categories without external supervision. How to make the network recognize different golden monkey individuals without identity information? Unlike humans, the neural network does not know what object is being recognized. The network extracts key information from the image, which can also be called the attributes of things (such as eye information, hair information, etc.). These key information are the basis for it to distinguish objects. Previous recognition algorithms required manual description of these attributes and converted them into machine language to assist the network in recognition. However, due to the characteristics of golden monkey facial images, it is quite difficult to manually describe the attributes of each golden monkey individual, because there are great similarities between golden monkey individuals, resulting in less difference between the description content.
[0061] Through in-depth research, it was found that the network can accurately identify labeled golden monkey individuals using supervised learning methods. Although the new individuals lack identity information, the key attributes between the new individuals and the original individuals are common. The feature output of the neural network will contain information about the key attributes. The clustering algorithm is a common method for classification through feature representation. The knowledge transfer method can be used to use the prior knowledge of labeled golden monkey data to reduce the ambiguity of the clustering algorithm and better cluster the same individuals in the new individuals.
[0062] The SS-NIR algorithm consists of three modules: encoder pre-training module, golden monkey category number estimation module and unlabeled data clustering module ( Figure 3 ).
[0063] First, the encoder pre-training module uses labeled golden monkey facial image data to train the twin network structure to extract the visual features of the golden monkey facial image. The network will randomly enhance two parts of the image and extract features from the two enhanced points to maximize the similarity between the features of the two enhanced points and enhance the robustness of the network. Next, through the knowledge transfer method, the trained network model is used as the feature encoder of the golden monkey category quantity estimation module to extract the feature representation of the unlabeled image data. The module estimates the number of categories of the unlabeled image data. Finally, after obtaining the number of categories of the unlabeled image data, the K-Means clustering algorithm is used for clustering and the new individual category labels are assigned. The feasibility and effectiveness of the algorithm are verified through experimental data.
[0064] Based on the above facts, the detailed implementation process of the self-supervised new individual recognition algorithm based on the twin network and deep feature clustering algorithm of this embodiment is described as follows:
[0065] 1. Video Processing
[0066] The collected video is processed, and the processing flow is as follows:
[0067] The video of golden monkeys is collected and divided into frames at equal intervals to obtain a video frame sequence. In this embodiment, 30 frames per second are taken for the video. Excessive sampling will lead to data redundancy, which is easy to cause the model to overfit. In addition, in the continuously shot golden monkey video data, the difference between adjacent continuous video frames is not obvious. The incomplete invalid frames are eliminated, and the remaining valid frames are used as the input of the pre-trained converged target detection network yolo v5 for monkey face positioning.
[0068] like Figure 2 As shown in the figure, it is obvious that the golden monkey facial image located by the traditional target detection method is not accurate enough, and some important areas of the golden monkey's face have been cropped. However, the golden monkey facial image located by the pre-trained target detection network yolo v5 is quite accurate, including all the distinguishing information of the golden monkey's face. The more accurate golden monkey facial data obtained in the pre-processing stage helps to improve the final recognition of the golden monkey's identity information.
[0069] All the golden monkey facial data obtained above are normalized in size, and then the normalized golden monkey facial images are corrected based on affine transformation. The projection transformation of the golden monkey face during rotation is abstracted and converted into a geometric model. The specific changes of the monkey face image are analyzed from the geometric model, and the projection plane image of the human face under rotation is analyzed in combination with the affine transformation theory, so as to realize the correction of the monkey face image. The preprocessing finally obtains a set of golden monkey facial images corresponding to a set of golden monkey videos.
[0070] 2. Pre-trained encoder
[0071] The encoder pre-training module uses the ResNet-50 deep neural network as the backbone network. The ResNet-50 network has 49 convolutional layers. Due to the high similarity between golden monkey individuals, and the greater the network depth, the DCNN will overfit due to the small amount of individual image data. In order to solve the above problems, the SS-NIR algorithm performs gradient descent iterative training on the DCNN based on the structure of the twin network.
[0072] In the encoder pre-training module, the characteristics of individual golden monkey facial images and the advantages of the twin network structure are combined to solve the problem of small amount of data for some individual images. The final framework is as follows Figure 4 shown.
[0073] The encoder pre-training module is divided into two parts. The upper part consists of ResNet-50 and a multi-layer perceptron (H), which is denoted as part a. The lower part consists of ResNet-50, which is denoted as part b. The ResNet-50 in a and b is used as an encoder. The two encoders have the same structure and share network parameters. Finally, the feature representation extracted from part a and part b is used to update the network parameters through the cosine similarity loss. The detailed process is as follows:
[0074] First, two regions are randomly selected from image x and expanded into views x1 and x2 as the input of the encoder pre-training module. x1 is passed through part a to obtain the feature representation p1, and x2 is passed through part b to obtain the feature representation η2.
[0075] As follows:
[0076] p1=H(f θ (x1))
[0077] η2=f θ (x2)
[0078] In the above formula, p1 is the output of H, and the loss function can be defined as:
[0079]
[0080] Among them, θ is the parameter of the encoder, which is a learnable parameter. D is the cosine distance. ||·||2 is the L2 norm. The purpose of optimizing the model is achieved by minimizing the following formula, so the loss function can be defined as:
[0081] L'=min θ,η L(θ,η)
[0082] The training of the encoder pre-training module requires solving the parameters θ and η, otherwise the model cannot be trained. θ is a variable, and if η is also a variable, it cannot be solved. Therefore, by stopping the reverse gradient calculation of part b and making it a constant, θ can be solved. Since there are differences between the two enhancements of an image, this embodiment eliminates the differences by exchanging views x1 and x2, that is, x2 extracts features from part a and x1 extracts features from part b. For the learning of each image, the loss function of the network is:
[0083]
[0084] In the above formula, stopgrad(·) stops the reverse gradient calculation.
[0085] The encoder uses the convolutional layer of ResNet-50, and the output is a 2048×1 feature vector, which then passes through three fully connected layers. The specific parameters are shown in Table 4.1. The specific parameters of H are shown in Table 4.2. From the actual effect, in the case of multiple GPUs, the effect is better after adding Batch Normalization.
[0086] Table 4.1: Encoder full-connection layer parameters
[0087]
[0088] Table 4.2: Specific parameters of H
[0089]
[0090] 3. Estimation of the number of golden monkey species
[0091] The main functions of the golden monkey category number estimation module are: in the special diagnosis extraction stage, the facial image data of the golden monkey is input into the twin network structure to train the feature encoder, and the labeled data and unlabeled data sets are clustered multiple times using K-means through the prior knowledge of the labeled data, and the number of classes of the unlabeled data is continuously estimated and updated. Then, the clustering results are checked by calculating two quality indexes, one of which is to check the quality of the labeled data, and its number of categories is known. The other quality index is to measure the clustering effect of the unlabeled data. The number of categories of the unlabeled data will be estimated as the one that maximizes the two quality indexes.
[0092] The specific process is as follows:
[0093] Assuming a labeled dataset Where x is an image, y is the category label corresponding to x, and the total number of labeled data is N. At the same time, assume an unlabeled dataset Category labels for unlabeled datasets is unknown, where K represents the number of golden monkey individuals contained in the unlabeled data, and K is also unknown.
[0094] First, T l Divide into probe subsets and training subset T t l , using the training subset T t lThe ResNet-50 is trained as the data input of ResNet-50. The trained ResNet-50 network is an existing network that can recognize labeled data. How to use the trained network for new individual recognition is the focus of this paper. Next, fix the convolutional layer parameters of ResNet-50, migrate the convolutional layer parameters of ResNet-50 to the encoder pre-training module, and transfer the unlabeled data T u The data input of the encoder pre-training module is used for fine-tuning. This process does not require label information of unlabeled data and is a self-supervised learning process.
[0095] Next, the probe subset With the unlabeled dataset T u It is also used to estimate the number of categories. Partition into fixed probe sets and validation probe sets Fixed probe set and validation probe sets The ratio is 4:1. Then K-means is used to and Clustering is performed to force a fixed probe set The same images are mapped to the same cluster in the validation probe set The images in are regarded as additional unlabeled data. By repeatedly using the K-means algorithm to continuously estimate and update The total number of categories C, and choose to use and The number of categories with the highest clustering quality is C, C is The number of categories and T u The sum of the number of categories K. For each C, it is measured by two indicators. The first indicator is to measure the fixed probe set The second indicator is to measure the clustering quality of Each indicator is used to determine the optimal number of categories, and the average value V of the two indicators is calculated ^ Finally, select the largest V from all the results ^ The corresponding number of categories C, the best category C is brought into the K-means algorithm Perform the last clustering on Any outlier clusters in the clustering results are processed by deleting clusters whose data around the clusters is less than or equal to 1% of the number of clusters around the largest cluster in the clustering results, and obtaining the best estimate of C, and then subtracting the known The number of categories is finally obtained. The algorithm flow is as follows:
[0096] Algorithm 1: Estimating the number of new individual categories
[0097] Input: Unlabeled dataset T u , fixed probe set and validation probe sets The number of clustering counts is the number of categories in the probe set dataset.
[0098] Output: The number of categories C;
[0099] If K satisfies 0≤K≤K max ,but:
[0100] Step 1: In and Run K-means on .
[0101] Step 2: Calculation ACC and CVI.
[0102] Finish.
[0103] Get the optimal solution:
[0104] Assumptions Yes The value of ACC maximizes, Yes The value of CVI that is maximized is calculated And select V ^ The number of categories C corresponding to the maximum value runs the K-means algorithm again.
[0105] Dealing with outlier clusters:
[0106] Check The resulting clusters are discarded and any clusters whose surrounding data are less than or equal to the maximum number of clusters are deleted (τ = 1%). Output the number of remaining clusters.
[0107] 4. Clustering Algorithms
[0108] The clustering algorithm assigns labels to new individual data. The quality of clustering depends on whether the extracted features are key features. Secondly, clustering using only feature representations cannot achieve good results. Therefore, an estimated value of the number of categories is obtained through the category number estimation of the unlabeled data clustering module. Considering that the encoder pre-training module has been used to train the unlabeled data T u Therefore, this section will use the method of alternating training of ResNet-50 and K-means clustering algorithm to update the network parameters of ResNet-50.
[0109] Supervised learning is to continuously adjust the model parameters through data so that the classifier can achieve accurate recognition. First, it is necessary to extract good visual features. θ In (·), f represents the process of extracting features from the network convolutional layer, and θ represents the corresponding parameters. We use labeled data to continuously update θ so that f θ (·) Generate good visual features. Usually a label y corresponds to a large amount of image data x n , by parameterizing the classifier G w (·) In the feature f θ (x n ) to predict the correct label to optimize the following formula:
[0110]
[0111] where l is the logistic loss. When θ is sampled from a Gaussian distribution, without any learning, f θ (·) will not produce good features. In general, data labels are used to guide the prediction of deep convolutional neural networks. The convolution structure has a strong prior for the input signal. This module requires a signal to guide the recognition of deep convolutional neural networks. For unlabeled data, clustering is a common division method. Images of the same golden monkey individual can be classified into one category and each image can be given a pseudo label.
[0112] Clustering has been widely studied and many methods have been developed for various situations. The unlabeled data clustering module uses the K-means clustering algorithm (the framework structure of the unlabeled data clustering module is as follows Figure 4 ). The K-means algorithm requires a certain number of clusters to be specified in advance. After extracting features through the encoder, the high-dimensional features are converted into low-dimensional features using principal component analysis (PCA), which are then used as input to cluster the image data into K different clusters. More precisely, first, the unlabeled data clustering module uses the features extracted by the encoder to cluster the image x i Assigned to a cluster in the clustering result, and then use back propagation to calculate the gradient through these labels to minimize formula 4.8, so as to train ResNet-50 and re-extract features, and then use K-means to T u Reassign each image in the cluster, calculate the distance between the image and its corresponding cluster, determine whether the distance is the smallest, adjust the center point of the cluster, calculate the distance between each cluster center point and the surrounding distribution points, and minimize the sum of all distances. The formula is as follows:
[0113]
[0114] In the above formula, R d×krepresents the d×k center point matrix, C represents the center point coordinates, y n Represents x n The cluster assignment of y. n ∈{0,1} k Indicates whether it belongs to the kth cluster, 0 means it does not belong, and 1 means it does. Until the network training is completed, K-means is used for the last clustering to obtain the final new individual label.
[0115] 5. Experiment
[0116] In order to verify the effectiveness of the method in this embodiment, the collection of golden monkey datasets and the comparison of experiments are carried out:
[0117] 5.1、Collecting Datasets
[0118] The specific implementation plan for golden monkey facial image data collection is:
[0119] The image data used for model training is obtained by directly photographing golden snub-nosed monkeys that have been habituated and attracted in the wild at a relatively close distance. Videos are recorded with mobile phones, SLRs, or digital cameras at a distance of 2-5m from the golden snub-nosed monkeys. The video frame rate is 60fps and the resolution is 1920×1080. When shooting, pay attention to distinguishing and marking individuals. The video duration of each individual is no less than 90s. After completing the video recording of all target individuals, extract pictures from the video at a certain frame rate interval, and then filter the pictures to obtain usable data images. The dataset is then produced through conventional processes such as labeling and cropping, and finally the model is trained.
[0120] 5.2 Experimental Environment
[0121] The experiment uses Python deep learning framework Pytorch to implement the deep neural encoder network. The device used in the experiment has 32GB of running memory, Ubuntu20.04.3LTS operating system, and three RTX2080Ti graphics cards with 11G video memory. The dependent packages of the experimental environment are: torch, numpy, torchvision, numpy, opencv-contrib-pytorh, tqdm, etc.
[0122] 5.3 Analysis of experimental results
[0123] In order to better evaluate the clustering results, the evaluation indicators selected in this embodiment are accuracy (ACC), normalized mutual information (NMI) and adjusted Rand index (ARI). The specific calculation formula is as follows:
[0124] ACC calculation formula:
[0125]
[0126]
[0127] o' i =map(o i )
[0128] In the above formula, g i (j) is the true label, o'(j) is a mapping function, which takes the true label g i (j) is used as the reference label, and then the label order in o'(j) is rearranged in the same way.
[0129] The NMI calculation formula is:
[0130]
[0131]
[0132] In the above formula, H(x) and H(y) are the entropies of x and y respectively.
[0133] The ARI calculation formula is:
[0134]
[0135]
[0136] In the above formula, a represents the image data with the same true label and belonging to the same cluster in the clustering result. b represents the image data with different true labels and not belonging to the same cluster in the clustering result.
[0137] The results show that, under the premise of assuming that the number of unlabeled data categories is 5, 80 golden monkey individuals are used for training. From the results of 5 experiments, it can be seen that the estimated number of categories and the actual number of categories are both 5, and the highest accuracy can reach 89.6%, and the average accuracy of 5 experiments is 84.28% (Table 5.1). Under the premise of assuming that the number of unlabeled data categories is 10, the number of categories estimated in 5 experiments is 10, 12, 12, 11, and 13 respectively. Among them, only one result is the same as the true value, which has a significant error compared with the experimental results of 5 unlabeled data categories. The ACC is significantly lower than the experimental results of 5 categories. However, we found that although the estimated value is higher than the true value, it does not have a significant impact on the clustering results. The highest accuracy can reach 73.3%, and the average accuracy of 5 experiments is 65.97% (Table 5.2).
[0138] Assuming that the number of unlabeled data categories is 15, the estimated number of categories in the five experimental results are 16, 18, 18, 17, and 18 respectively. It can be seen that the results of three experiments are 18, which is 3 different from the actual number of categories. Compared with the experimental results with 5 and 10 unlabeled data categories, the error between the estimated value and the actual value has increased. The highest accuracy can reach 64.6%, and the average accuracy of the five experiments is 60.63% (Table 5.3).
[0139] Table 5.1: Randomly select 5 individuals as unlabeled data for category estimation
[0140]
[0141] Table 5.2: 10 individuals are randomly selected as unlabeled data for category estimation
[0142]
[0143] Table 5.3: 15 individuals are randomly selected as unlabeled data for category estimation
[0144]
[0145] By using the Axes3D library, the clustering results of the pre-assumed 5, 10, and 15 categories of unlabeled individuals are projected into 3D space, and the clustered images are visualized, as shown in Figure 5 shown.
[0146] exist Figure 5 In (a), the five categories of unlabeled data are clustered according to the estimated number of categories (light and dark colors represent different categories). It can be seen that the unlabeled data are clearly divided into five categories.
[0147] exist Figure 5 (b) and 5(c) show the clustering of 10 and 15 unlabeled data with the estimated number of categories added. It can be clearly seen that by using the SS-NIR network proposed in this application, most of the data is clearly divided, but the effect is obviously not as obvious as that of 5 individuals. On the one hand, the number of individuals in the unlabeled data has increased, and the estimated number of categories has an error with the actual number of categories. K-means adds center points during the clustering process, resulting in many images being divided into the wrong center points. On the other hand, as the number of category individuals increases, a part of individuals with high similarity appears in the middle, and these individuals will be confused with other individuals.
[0148] from Figure 5In the second figure (b), it can be seen that there is little overlap between different individuals. It can be found that the SS-NIR network has sufficient discrimination for different new individuals, which clearly shows that the golden monkey new individual recognition method provided in this embodiment can effectively discover new individuals.
[0149] Ten golden monkey individuals were selected from the golden monkey dataset as unlabeled data. Five groups of experiments were conducted to compare with the unsupervised clustering algorithms LSC, Context Encoders, BiGANs, CFN and Split-Brain Auto. The best results among the five groups of experiments were selected for comparison. The experimental results are shown in Table 5.4.
[0150] Table 5.4: Comparison with unsupervised algorithms
[0151] algorithm ACC NMI ARI LSC 0.4460 0.4045 0.2374 Context Encoders 0.6200 0.7174 0.5212 BiGANs 0.6040 0.6619 0.4752 CFN 0.5486 0.6476 0.4088 Split-Brain Auto 0.6580 0.7003 0.5241 SS-NIR 0.7330 0.7445 0.6082
[0152] 6. Conclusion
[0153] In summary, it is shown that SS-NIR can effectively estimate the approximate number of categories of unlabeled data and cluster them according to the estimated number of categories. The clustering accuracy decreases with the increase of data volume, but the overall result is in line with expectations. Errors in the experiment are normal. We do not add all new individuals to the original data set, but add the first few individuals with high accuracy in the recognition results to the original data set.
[0154] In summary, this application first proposes the problem of new individual identification in the golden snub-nosed monkey identification problem. By analyzing the challenges faced by new individual identification and the characteristics of the golden snub-nosed monkey dataset, combined with some knowledge from other fields of deep learning, a golden snub-nosed monkey new individual identification algorithm based on twin networks and clustering algorithms, namely SS-NIR, is designed. Through the twin network structure, the algorithm uses the ResNet-50 network to extract the features of two enhanced points in the golden snub-nosed monkey image, and trains the ResNet-50 network to learn the feature representation of the labeled golden snub-nosed monkey image by calculating the similarity between the two features. Secondly, the knowledge transfer method is applied to estimate the number of unlabeled golden snub-nosed monkey categories using the common attributes between golden snub-nosed monkey individuals. Then, the unlabeled data is clustered using the cluster analysis method based on the estimated number of unlabeled categories. This method mainly solves the problem that the existing methods cannot identify unlabeled golden snub-nosed monkey data. Finally, in order to verify the effectiveness of the algorithm, the experimental design details are introduced, and a large number of comparative experiments are carried out on the golden snub-nosed monkey dataset. From the experimental results, it can be seen that the error between the estimated number of categories and the actual number of categories is within 3, and SS-NIR has a good performance in identifying new golden monkey individuals.
Claims
1. A method for identifying a new golden monkey individual, characterized in that: This method uses a self-supervised new individual recognition algorithm based on a twin network and deep feature clustering. It introduces a golden monkey facial image pre-training encoder module, a golden monkey category quantity estimation module, and an unlabeled data clustering module to perform new individual recognition on golden monkey data without identity information. The specific steps include: Step S1: firstly, use a high-resolution camera to collect videos of golden monkeys; Step S2: The video is divided into frames at equal intervals. After the frames are divided, the golden monkey images are evaluated using the image quality assessment algorithm MUSIQ to filter out images with severe blur and jitter; Step S3: label the face of the golden monkey using labelme software, and perform face detection on the image containing the golden monkey's face information using the object detection network yolov5; then segment the located face from the image to obtain an image containing only the golden monkey's face; Step S4: pre-processing the detected golden monkey facial image data, cropping them to the same size and correcting the golden monkey facial image; Step S5: The high-quality golden monkey facial image pre-trained encoder module is used to extract the visual features of the golden monkey individuals, reduce the ambiguity of clustering through the prior knowledge learned from the labeled data, and improve the accuracy of the new individual division; Step S6: using the structure of the Siamese network to learn facial visual features from the data; The specific implementation steps include: Step 61: Input the high-quality golden monkey image data into the encoder for pre-training. The pre-training module uses the ResNet-50 deep neural network as the backbone network. Step 62: Obtain the deep feature vectors of the two golden monkey images through the Siamese network encoder, and calculate the cosine distance between the two feature vectors; Step 63: Continuously iterate the network for self-supervising new individual recognition based on the twin network and the deep feature clustering algorithm, and update the network parameters through back propagation until the network converges; Step S7: a category quantity estimation module is used to estimate the category quantity of new golden monkey individuals, and estimate the category quantity of unlabeled data using the prior knowledge of labeled data through the knowledge transfer method; the specific implementation steps include: Step S71: using the prior knowledge of labeled data, clustering the labeled data and unlabeled data sets multiple times using K-means, and continuously estimating and updating the number of classes of unlabeled data; Step S72: Calculate two quality indexes to check the obtained clustering results, one of which is to check the quality of the labeled data, the number of its categories is known; Step S73: The number of categories of unlabeled data is estimated to be the one that maximizes the two quality indexes; Step S8: Cluster analysis module, clustering the unlabeled data through clustering algorithm to achieve the purpose of individual identification; the specific implementation steps include: Step S81: The number of categories of the unlabeled data is estimated by the category quantity estimation module in step S7, and then the new individual data is labeled using a clustering algorithm; Step S82: Update the network parameters of ResNet-50 by alternately training ResNet-50 and K-means clustering algorithm; Step S83: Classify the images of the same golden monkey individual into one category and assign a pseudo label to each image; Step S9: Add new golden monkey individuals with high recognition accuracy to the original data set for data enhancement.
2. The method according to claim 1, characterized in that The specific implementation steps of step S2 include: Step S21: First, a multi-scale representation of the golden monkey image is obtained, including the original image and a variant with a fixed aspect ratio. Images of different scales are divided into image blocks of fixed size, and then fed into a pre-trained image quality assessment model. Since the image blocks are from images of different spatial resolutions, it is necessary to efficiently encode these inputs of multiple aspect ratios and scales into a token sequence to capture pixel, spatial and scale information. Step S22: Based on hash-based two-dimensional space embedding, the position of a certain image block is recorded in the Row, No. Columns, hashed to The corresponding elements in the grid of ; each element of the grid is a -dimensional embedding vector; that is, there is a learnable matrix , the input image size is , and divide the image into multiple sizes For the image block at position The spatial embedding of an image patch is defined as In Elements of position; Step S23: reuse the same hash matrix for the used images. HSE cannot distinguish image blocks from different scales, so an additional scale embedding SCE is introduced to help the model distinguish image blocks from different scales. Step S24: Pre-training is performed on ImageNet. Various data augmentation methods are used to improve performance. Fine-tuning is performed on image quality and aesthetic quality datasets. In the fine-tuning stage, the original size image is kept as input, and the data augmentation only uses horizontal flipping that has no effect on image quality. Step S25: Use the quality assessment algorithm MUSIQ to assess the quality of the collected golden monkey images, set a reasonable threshold, and filter out image data below the threshold.
3. The method according to claim 1, characterized in that The specific implementation steps of step S3 include: Step S31: Use the labelme annotation tool to finely annotate the face of the golden monkey image data, so that the target detection network can accurately obtain the image coordinate information of the monkey face; Step S32: Feed the finely annotated data for training the detection network into the target detection network YOLO V5, update the network parameters through back propagation until the network converges, and finally complete the pre-training of the target detection network YOLO V5; Step S33: Use the pre-trained model to crop the monkey face part of the unlabeled golden monkey image data to prepare for the later estimation of the number of golden monkey categories and self-supervised identification of new individuals.
4. The method according to claim 1, characterized in that The specific implementation steps of step S4 include: Step S41: The golden monkey facial data obtained in step S3 is cropped to the same size of 128*128 to facilitate subsequent model processing; Step S42: Correct the golden monkey face image based on affine transformation, abstract the projection transformation of the golden monkey face during the rotation process and convert it into a geometric model; parse the specific changes of the monkey face image from the geometric model and analyze the projection plane image of the human face under rotation in combination with the affine transformation theory, so as to achieve correction of the monkey face image.