An open-world species classification method based on active learning
By integrating the open vocabulary model and closed-set classification model, and using active learning and visual cues technology, the problem of limited performance and computational intensiveness of open vocabulary model in wildlife monitoring image classification is solved, and efficient species classification and new categories are implemented, which significantly improves classification accuracy and performance.
Patent Information
- Application Number
- CN202411733993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing open vocabulary models have problems with limited performance and computational resource intensiveness in wildlife monitoring image species classification tasks, especially when frequently trained and fine-tuned, which may lead to the model forgetting and impairing zero-sample classification capabilities.
An open-world species classification method based on active learning is proposed, which integrates closed-set classification model and open vocabulary model. It adapts to species classification tasks through prompt engineering and visual cues learning, uses label confidence ranking and tag integration strategy for active learning, and periodically updates the closed-set model to include new categories.
It has achieved efficient classification and included in the new categories in open-world wildlife monitoring images, which has improved the accuracy of open vocabulary species classification of CLIP and EVA-CLIP, saved labeling costs, and achieved significant improvements in closed-set recognition performance.
Smart Images

Figure CN119559443B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of species classification, and in particular to an open world species classification method based on active learning. Background Art
[0002] The scale of wildlife resource surveys and biodiversity monitoring is constantly expanding, and the number of monitoring images is growing. In recent years, my country's wildlife surveys, especially those on terrestrial vertebrates, have made rapid progress in both survey scale (geographical area, animal group coverage) and survey methods. Technologies such as infrared triggered cameras provide an effective means of data collection, and wildlife resource surveys and monitoring conducted by national parks, nature reserves at all levels, and forestry departments provide a rich and diverse source of data.
[0003] The open vocabulary model needs to have the ability to process unknown categories in an open environment. Although the zero-shot classifier of the open vocabulary model based on text features has a flexible number of categories and is well adapted to the scenario of increasing species categories in the open environment of wildlife monitoring, the open vocabulary model has a gap in performance and computing resource requirements with the task of species classification in monitoring images. There is a significant difference between the pre-training data of the open vocabulary model and the monitoring images, so there is a lack of domain knowledge for species classification in monitoring images, resulting in limited performance. Wildlife monitoring continuously generates new data, which provides good conditions for the continuous learning of the model, so the model should be trained periodically to obtain the best performance. However, the computationally intensive open vocabulary model is not suitable for frequent training or fine-tuning, and may cause the model to forget the knowledge obtained from large-scale pre-training, and even damage the zero-shot classification ability. Therefore, how to retain the pre-training knowledge while acquiring the knowledge of new data to achieve the classification and continuous learning of unknown categories in an open environment is a problem to be solved. In view of this, the present invention proposes a method for open world species classification based on active learning for the problem of classifying and continuously learning new categories in an open monitoring environment. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] The purpose of the present invention is to propose an open-world species classification method based on active learning to solve the problems mentioned in the background technology. Faced with a steady stream of new data, the present invention can effectively classify open-world wildlife monitoring images and efficiently incorporate new categories into the scope of supervised learning.
[0006] (II) Technical solution
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention proposes an open-world species classification method based on active learning, which fuses a closed-set classification model and an open vocabulary model, wherein the closed-set classification model is supervised and trained by the current monitoring image dataset to accurately identify known categories, and the open vocabulary model is adapted to the species classification task through cue engineering and visual cue learning;
[0009] During testing, the closed-set model routes low-confidence test samples to the open-vocabulary model for unknown category classification, and integrates the open-vocabulary classification results with the high-confidence closed-set classification results to complete the species classification of the open-world scenario;
[0010] The integrated results are used as weakly supervised labels for the closed-set model, and active learning is used to perform label queries to improve the quality of pseudo labels and discover new categories. As new categories emerge and samples accumulate periodically, the closed-set model is retrained and the number of species classifier categories is expanded. The knowledge of the open vocabulary model and expert labels is periodically injected into the closed model to achieve optimal performance, making full use of labeled data while avoiding computationally intensive open vocabulary model parameter adjustments. The following further describes an open-world species classification method based on active learning proposed in the present invention, which specifically includes the following contents:
[0011] An open-world species classification method based on active learning, comprising the following steps:
[0012] S1. Continuously monitor and obtain wildlife monitoring data in open environments, build an open set recognition model to classify the acquired monitoring data, and obtain classification results for known categories;
[0013] S2, realize unknown category detection based on the maximum logits score, use cue engineering and visual cues to obtain species vocabulary and unknown category samples respectively, build an open vocabulary model, assign open vocabulary labels to unknown category samples, and obtain open vocabulary classification results;
[0014] S3, perform label query on the open vocabulary classification results obtained in S2 to obtain high-quality new category labels, update the species vocabulary through the high-quality new category labels, and use label confidence ranking and label integration strategy to select low-confidence new category labels for active learning;
[0015] S4. Integrate the known category classification results obtained in S1 and the high-quality new category labels obtained in S3 to construct a processed wildlife monitoring dataset and store it in the monitoring image database. Use the data in the monitoring image database to periodically update the open set recognition model, incorporate the new categories into the closed set classification range, and realize species classification of open world wildlife monitoring images.
[0016] Preferably, the method for calculating the label confidence comprises the following steps:
[0017] Convert the closed set classifier from the threshold to the open set classifier, denote the input sample as x, the feature extractor as Φ(·), and the closed set classifier weight matrix for classifying K known categories as W K , then the maximum confidence score P of the model output is:
[0018]
[0019] Among them, T represents transposed convolution;
[0020] Assuming that the model's confidence in the prediction of closed-set samples is higher than the confidence in the prediction of open-set samples, the closed-set classifier is extended by the threshold, as follows:
[0021]
[0022] in, Represents the predicted label of the model output; the K+1th class represents the unknown class; threshold represents the open set recognition threshold;
[0023] When the maximum value in the sample prediction distribution exceeds the threshold, it is identified as a closed class, otherwise it is predicted as an unknown class.
[0024] Preferably, the open vocabulary model is used for monitoring image species classification, and further comprises the following steps:
[0025] Clarify the scope of unknown categories and construct corresponding text input;
[0026] Construct corresponding text input for each category based on species name for classification;
[0027] The text descriptions of the salient visual features of a particular species are summarized by the Large Language Models (LLMs) Llama 2 based on the animal descriptions of the corresponding species pages on Wikipedia;
[0028] Keep the feature description text within 30 words to meet the text input limit of the visual language model and avoid exceeding the token limit of the model, which will cause the text input to be truncated.
[0029] Preferably, adjusting the image input using visual cues comprises the following steps:
[0030] Adjust the strategy based on visual cues, by drawing the MegaDetector object detection box to highlight the foreground cues;
[0031] Blurring the background, removing the background color, and completely removing the background are used to suppress background information and highlight the animal foreground cues; the original image is used as a control to balance the foreground and background of the expanded detection frame cropped image cues; the animal foreground is segmented from the monitoring image as other cues;
[0032] According to the image input size of the model, the crop size of the expanded detection box is 336×336;
[0033] For monitoring images outside the experimental dataset, the above method or SAM model is used to generate the foreground mask required for visual cues.
[0034] Preferably, the method of selecting low-confidence newly added category labels for active learning by using label confidence ranking and label integration strategy specifically includes the following steps:
[0035] The model is calibrated by adjusting the temperature, and then the open vocabulary labels are sorted by confidence, and label queries are performed on low-confidence regions;
[0036] The calibrated probability distribution is recorded as q, the predicted distribution before calibration is recorded as s, and the temperature is recorded as τ. The formula is the process of selective label query on samples using temperature scaling:
[0037]
[0038] in, Indicates the label finally assigned to the sample; labeling indicates the label assigned to the sample by the expert; threshold indicates the frequency and cost of querying expert labels;
[0039] When a sample is assigned the same label by two models, the sample is considered to be correctly classified and does not participate in label query. Otherwise, the sample is assigned an open vocabulary label after active learning.
[0040] The above label integration strategy provides more flexibility for classifying samples of known categories, while ensuring that active learning provides more knowledge of unknown categories.
[0041] (III) Beneficial effects
[0042] Compared with the prior art, the present invention provides an open-world species classification method based on active learning, which has the following beneficial effects:
[0043] In view of the long-term nature of wildlife monitoring and the openness of the monitoring environment, the present invention proposes a species classification method for open-world wildlife monitoring images. In the face of a steady stream of new data, the present invention can effectively classify open-world wildlife monitoring images and efficiently include new categories in the scope of supervised learning. When classifying monitoring image species, the method first classifies known category samples by a supervised learning model, and realizes unknown category detection based on the maximum logits score; then the visual language model adapted to species classification is adjusted by prompt engineering and visual prompts to assign open vocabulary labels to unknown category samples; then the label confidence ranking and label integration strategy are used to select low-confidence new category samples for active learning; finally, with the continuous accumulation of new category samples, the framework periodically updates the closed set classification model and includes the new categories in the closed set classification scope, thereby realizing the species classification of open-world wildlife monitoring images. In addition, with the emergence of new categories and the accumulation of samples, the closed set model is retrained and the number of classifier categories is expanded, so that the knowledge of the open vocabulary model and expert labels is periodically updated to the closed set model, thereby gradually transforming the open-world image classification problem into a closed set classification problem. Compared with the baseline method, this method improves the open vocabulary species classification accuracy of CLIP and EVA-CLIP by 10.9% and 9.03%, respectively, and saves about 50% of the label cost when learning new categories. Based on the LoCo model, the maximum logits score is used for open set recognition. The cross entropy model that also uses MLS is selected as the comparison baseline. The open set recognition performance on the iWildCam36 test set. Compared with the baseline method, this method improves the closed set accuracy by 0.66%, and the AUROC and OSCR scores are improved by 0.85 and 1.03, respectively. The open set performance of the model can benefit from the improvement of closed set accuracy. Textual prompts describing the type of wildlife monitoring images and the classification target can improve performance. The combination of the two improves the accuracy by 4.98%. In contrast, prompts that describe the salient visual features of species in detail fail to achieve the expected effect, which may be limited by the text understanding ability of CLIP. From the perspective of visual cues, the combination of visual cues of extended detection box cropping and various cues templates can improve performance, but most of the other methods that suppress the background to highlight the animal foreground are not better than the corresponding baselines, which may be because these visual cues deviate significantly from the pre-training data of the visual language model. The results of this method on the more powerful visual language model EVA-CLIP show that compared with the baseline, the optimal combination of text templates and visual cues of this method improves the accuracy of CLIP and EVA-CLIP by 10.9% and 9.03%, respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flow chart of an open-world species classification method based on active learning proposed in Example 1 of the present invention;
[0045] Figure 2 This is a schematic diagram of the steps of the "divide and conquer" strategy method proposed in Example 1 of the present invention;
[0046] Figure 3 The open set recognition performance evaluation proposed in Embodiment 1 of the present invention;
[0047] Figure 4 This is a CLIP performance comparison of the cueing engineering combined with visual cues proposed in Example 1 of the present invention;
[0048] Figure 5 This is a performance comparison of the prompt engineering proposed in Example 1 of the present invention combined with different visual prompts on EVA-CLIP;
[0049] Figure 6 The performance comparison between the method proposed in Example 1 of the present invention and the baseline method is shown;
[0050] Figure 7 This is the visual prompt adjustment of the wildlife monitoring image proposed in Example 1 of the present invention. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0052] The present invention proposes an open-world species classification method based on active learning. In the face of a steady stream of new data, the method can effectively classify open-world wildlife monitoring images and efficiently include newly added categories in the scope of supervised learning. When classifying monitoring image species, the present invention first classifies known category samples by a supervised learning model, and realizes unknown category detection based on the maximum logits score; then, the visual language model adapted to species classification by prompt engineering and visual prompts is adjusted to assign open vocabulary labels to unknown category samples; then, the label confidence ranking and label integration strategy are used to select low-confidence newly added category samples for active learning; finally, as the newly added category samples continue to accumulate, the framework periodically updates the closed-set classification model and includes the newly added categories in the closed-set classification scope, thereby realizing the species classification of open-world wildlife monitoring images.
[0053] In order to better understand the above technical solution, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0054] See also Figure 1 , Figure 1 This is a flow chart of an open-world species classification method based on active learning proposed in the present invention, which specifically includes the following contents.
[0055] like Figure 1 As shown, the active learning open-world species classification method includes the following steps:
[0056] First, a periodically trained open-set recognition model and a frozen parameter open-vocabulary model are proposed, the latter of which is only used in the inference process. When processing a batch of surveillance images to be classified, the open-set recognition model is first used to identify samples of unknown categories and classify samples of known categories. The closed-set classification model and the open-vocabulary model are fused, wherein the closed-set classification model is supervised and trained by the current surveillance image dataset to accurately identify known categories, and the open-vocabulary model is adapted to the species classification task through cue engineering and visual cue learning; during testing, the closed-set model routes low-confidence test samples to the open-vocabulary model for unknown category classification, and integrates the open-vocabulary classification results and the high-confidence closed-set classification results to complete the species classification of the open-world scene; the integrated results are used as weakly supervised labels for the closed-set model, and label queries are performed through active learning to improve the quality of pseudo-labels and discover new categories. As new categories appear and samples accumulate periodically, the closed model is retrained and the number of species classifier categories is expanded. The knowledge of the open-vocabulary model and expert labels is periodically injected into the closed model to achieve optimal performance, making full use of labeled data while avoiding computationally intensive open-vocabulary model parameter adjustments.
[0057] Secondly, the unknown category samples are further processed by the open vocabulary model, which constructs text prompts based on the list of wildlife species and classifies them into the “unknown” category. The specific content is shown in Table 1.
[0058] Table 1
[0059]
[0060] An open set recognition based on a closed set classifier is proposed. When processing a batch of surveillance images to be classified, the open set recognition model is first used to identify samples of unknown categories and classify samples of known categories. The confidence is calculated. The closed set classifier can be converted to an open set classifier by the threshold. The input sample is x, the feature extractor is Φ(·), and the weight matrix of the closed set classifier for classifying K known categories is W K , then the maximum confidence score P of the model output is:
[0061]
[0062] Assuming that the model predicts closed-set samples with higher confidence than open-set samples, the closed-set classifier can be extended by the threshold, as follows:
[0063]
[0064] in, Represents the predicted label of the model output, the K+1th class represents the unknown class, and threshold represents the open set recognition threshold. When the maximum value in the sample prediction distribution exceeds the threshold, it is identified as a closed set class, otherwise it is predicted as an unknown class.
[0065] The open vocabulary model is adapted to the species classification of monitoring images. It is necessary to clarify the scope of unknown categories and construct corresponding text inputs. It is necessary to construct corresponding text inputs for each category according to the species name for classification. The text description of the significant visual features of a specific species is summarized by the large language model (LLMs) Llama 2 based on the animal description of the corresponding species page on Wikipedia. In order to meet the text input limit of the visual language model, the feature description text is controlled within 30 words to avoid exceeding the token upper limit of the model and causing the text input to be truncated. The image input is adjusted using visual cues. The visual cue adjustment strategy includes cues that highlight the foreground by drawing the MegaDetector target detection box, cues that suppress background information and highlight the animal foreground by blurring the background, removing the background color, and completely removing the background, and cues that crop the image with the expanded detection box to balance the foreground and background. In addition, the original image is used as a control. According to the image input size of the model, the crop size of the expanded detection box is 336×336. Other cues require the animal foreground to be segmented from the monitoring image. For monitoring images outside the experimental dataset, similar methods or models such as SAM can be used to generate the foreground mask required for visual cues.
[0066] Finally, after obtaining the labels and confidences of samples of unknown categories, label queries are performed on difficult samples to actively learn expert knowledge of new categories, and finally the labels of closed-set classification, high-confidence open vocabulary labels, and expert labels are integrated as the classification results of this part of the data. As new data continues to accumulate, the open-set recognition model with low computational cost is periodically trained, and the open vocabulary model and actively learned new category knowledge are updated to the model parameters. This process gradually expands the category range of closed-set classification, converts new categories into known categories, and achieves a good balance between model performance, computational overhead, and annotation cost. For active learning of new categories, the temperature calibration model is adjusted, and then the open vocabulary labels are sorted by confidence, and label queries are performed on low-confidence areas. The calibrated probability distribution is recorded as q, the predicted distribution before calibration is recorded as s, and the temperature is represented by τ. The formula is the process of selective label query on samples using temperature scaling:
[0067]
[0068] in, Indicates the label finally assigned to the sample, labeling indicates the label assigned to the sample by the expert, and threshold controls the frequency and cost of querying the expert label. When a sample is assigned the same label by two models, it is considered that the sample is correctly classified and does not participate in the label query. Otherwise, the sample is assigned an open vocabulary label after active learning. The label integration strategy provides greater flexibility for the classification of samples of known categories, while ensuring that active learning provides more knowledge of unknown categories.
[0069] See also Figure 2 , Figure 2 Schematic diagram of the steps of the "divide and conquer" strategy method of the present invention, as shown in Figure 2 As shown, the "divide and conquer" strategy includes the following steps:
[0070] At the sample level, the classification cost is paid according to the sample difficulty. Samples of known categories are classified by a fully trained closed-set model, while samples of unknown categories are guided to the next stage for open-vocabulary classification through open-set recognition. Difficult samples that still have high uncertainty in open-vocabulary classification are actively learned to obtain expert labels.
[0071] At the method level, the classification method is determined based on the degree to which the model has learned the species categories. When the number of available labels is insufficient to support supervised training of new categories, active learning is combined to assign open vocabulary labels to samples; when new category labels accumulate to a certain number over time, supervised training is used to convert these newly added unknown categories into known categories that can be processed by the closed-set model, completing the transition from relatively rough open vocabulary zero-shot classification to more precise supervised model classification.
[0072] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0073] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions.
[0074] It should be noted that in the claims, any reference numerals placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In the claims enumerating several means, several of these means may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience of expression only and does not indicate any order. These words may be understood as part of the component name.
[0075] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.
[0076] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments after knowing the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0077] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention should also include these modifications and variations.
Claims
1. An open-world species classification method based on active learning, characterized in that: The following steps are involved: S1. Continuously monitor and obtain wildlife monitoring data in open environments, build an open set recognition model to classify the acquired monitoring data, and obtain classification results for known categories; S2, realize unknown category detection based on the maximum logits score, use cue engineering and visual cues to obtain species vocabulary and unknown category samples respectively, build an open vocabulary model, assign open vocabulary labels to unknown category samples, and obtain open vocabulary classification results; S3, performing label query on the open vocabulary classification results obtained in S2 to obtain high-quality new category labels, updating species vocabulary through high-quality new category labels, and selecting low-confidence new category labels for active learning using label confidence ranking and label integration strategy; the label confidence calculation method comprises the following steps: Convert the closed set classifier from the threshold to the open set classifier, denote the input sample as x, the feature extractor as Φ(·), and the closed set classifier weight matrix for classifying K known categories as W K , then the maximum confidence score P of the model output is: Among them, T represents transposed convolution; Assuming that the model's confidence in the prediction of closed-set samples is higher than the confidence in the prediction of open-set samples, the closed-set classifier is extended by the threshold, as follows: in, Represents the predicted label of the model output; the K+1th class represents the unknown class; threshold represents the open set recognition threshold; When the maximum value in the sample prediction distribution exceeds the threshold, it is identified as a closed set category, otherwise it is predicted as an unknown category; S4. Integrate the known category classification results obtained in S1 and the high-quality new category labels obtained in S3 to construct a processed wildlife monitoring dataset and store it in the monitoring image database. Use the data in the monitoring image database to periodically update the open set recognition model, incorporate the new categories into the closed set classification range, and realize species classification of open world wildlife monitoring images.
2. The open world species classification method based on active learning according to claim 1, characterized in that: The open vocabulary model is used for monitoring image species classification, and further comprises the following steps: Clarify the scope of unknown categories and construct corresponding text input; Construct corresponding text input for each category based on species name for classification; The text descriptions of the species’ salient visual features are summarized by the large language model Llama 2 based on the animal descriptions on the corresponding species pages on Wikipedia; Keep the feature description text within 30 words to meet the text input limit of the visual language model and avoid exceeding the token limit of the model, which will cause the text input to be truncated.
3. The open world species classification method based on active learning according to claim 2, characterized in that: Using visual cues to adjust the image input includes the following steps: Adjust the strategy based on visual cues, by drawing the MegaDetector object detection box to highlight the foreground cues; Blurring the background, removing the background color, and completely removing the background are used to suppress background information and highlight the animal foreground cues; the original image is used as a control to balance the foreground and background of the expanded detection frame cropped image cues; the animal foreground is segmented from the monitoring image as other cues; According to the image input size of the model, the crop size of the expanded detection box is 336×336; For monitoring images outside the experimental dataset, the above method or SAM model is used to generate the foreground mask required for visual cues.
4. The open world species classification method based on active learning according to claim 3, characterized in that: The method of selecting low-confidence newly added category labels for active learning by using label confidence ranking and label integration strategy specifically includes the following steps: The model is calibrated by adjusting the temperature, and then the open vocabulary labels are sorted by confidence, and label queries are performed on low-confidence regions; The calibrated probability distribution is recorded as q, the predicted distribution before calibration is recorded as s, and the temperature is recorded as τ. The formula is the process of selective label query on samples using temperature scaling: in, Indicates the label finally assigned to the sample; labeling indicates the label assigned to the sample by the expert; threshold indicates the frequency and cost of querying expert labels; When a sample is assigned the same label by two models, it is considered that the sample is correctly classified and does not participate in label query. Otherwise, the sample is assigned an open vocabulary label after active learning.