A new class discovery method and apparatus for images

CN122821155APending Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610789563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,当前方法在面对由大量“尾部”未知类别图像构成的复杂数据时,图像分类性能显著下降

Benefits of technology

[0007]本发明的有益效果是:本发明在图像分类器和特征提取网络的训练过程中可以发现更多的新类别,在面对由大量“尾部”未知类别图像构成的复杂数据时,可以提升图像分类性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821155A_ABST
    Figure CN122821155A_ABST
Patent Text Reader

Abstract

The application discloses a new class discovery method and device of an image, extracts image features of an image to be classified through a feature extraction network and feeds the image features into an image classifier; the image classifier calculates feature centers of the image features; and labels of the image to be classified are determined based on the feature centers and class prototypes of preset classes in the image classifier; wherein, a training set is constructed based on known class images and unknown class images, the feature extraction network and the image classifier are jointly trained based on the training set, and the number of new classes is not preset in the joint training; more new classes can be discovered in the training process of the image classifier and the feature extraction network, and the image classification performance can be improved when facing complex data composed of a large number of "tail" unknown class images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image classification technology, and in particular relates to a method and apparatus for discovering new categories of images. Background Technology

[0002] In real-world data, the distribution of categories often follows a long-tail pattern: a few categories have a large number of samples (head categories), while the vast majority of categories have only a small number of samples (tail categories). This phenomenon is particularly pronounced in fields such as biomedicine and fine-grained image recognition. For example, in medical imaging, there is ample image data for common diseases, but samples for rare diseases are extremely scarce; in species classification, image data for endangered or newly discovered species is scarce.

[0003] The task of discovering new categories aims to explore and identify images of unknown categories by leveraging knowledge of images with known categories (typically the data-rich head categories). However, current methods suffer significant performance degradation when faced with complex datasets consisting of a large number of "tail" images with unknown categories. Therefore, developing novel image classification methods that can effectively handle long-tail distributions and accurately discover rare categories has become a key challenge driving this field toward practical applications. Summary of the Invention

[0004] The purpose of this invention is to provide a method and apparatus for discovering new image categories, so that image classifiers and feature extraction networks can discover more image categories during the training process.

[0005] This invention employs the following technical solution: a method for discovering new categories of images, characterized by comprising the following steps: Image features of the image to be classified are extracted by a feature extraction network and then fed into an image classifier. Image classifiers calculate feature centers of image features; The label of the image to be classified is determined based on the feature center and the class prototype of the preset category in the image classifier; The training set is constructed based on known and unknown class images. The feature extraction network and image classifier are jointly trained based on the training set, and the number of new classes is not pre-defined in the joint training.

[0006] Another technical solution of the present invention includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the above-described method when executing the computer program.

[0007] The beneficial effects of this invention are: it can discover more new categories during the training process of image classifiers and feature extraction networks, and can improve image classification performance when faced with complex data consisting of a large number of "tail" images of unknown categories. Attached Figure Description

[0008] Figure 1 This is a framework diagram of the model in an embodiment of the present invention; Figure 2 This is a visual analysis diagram showing the feature distribution of each component in this invention and whether it uses a self-attention mechanism. Detailed Implementation

[0009] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0010] Existing methods for discovering new categories largely rely on flattened clustering strategies (such as K-Means). While these methods are relatively simple in concept and implementation, two inherent drawbacks make them ill-suited for handling complex scenarios with long-tailed distributions: (1) Neglect of hierarchical information. The category system in the real world is a natural hierarchical structure (e.g., "disease" → "inflammatory disease" → "specific rare inflammation"). Flat clustering forces all categories to be placed at the same level, discarding this structured information that can help the model to make progressive reasoning from "macro" to "micro". When it is necessary to distinguish a large number of similar and scarce tail categories (e.g., rare diseases of different subtypes), the model without hierarchical guidance has difficulty capturing subtle but key discriminative features, resulting in insufficient ability to discover tail categories.

[0011] (2) Strong dependence on the number of preset categories and incompatibility with long tails. Traditional methods usually require pre-specifying the total number of categories to be discovered, K. Under long-tail distribution, the number of tail categories is large and difficult to count in advance, and this "strong assumption" is almost impossible to satisfy in practice. An inaccurate value of K will directly cause the image classifier to incorrectly merge multiple unknown rare category images or forcibly classify them into head categories, which seriously restricts the robustness and practicality of the method.

[0012] The core challenge of new category discovery lies in how to automatically and accurately identify and learn unknown categories from unlabeled images using supervised information from known categories. Traditional methods often employ flat clustering strategies (such as K-Means), ignoring the inherent hierarchical structure information between categories (e.g., "vehicles" can be subdivided into "cars" and "buses"). Furthermore, flat clustering strategies heavily rely on a pre-defined number of categories, which is a strong assumption in practical applications, leading to performance degradation when there is high similarity between categories (such as different dog breeds).

[0013] Profoundly inspired by biomedical taxonomy and phylogenetics, biologists construct "trees of life" to depict evolutionary relationships between species, subdividing them layer by layer from coarse-grained phyla and classes to fine-grained genera and species, thus systematically encompassing all species from common to rare. Therefore, a robust framework for discovering new categories must be able to automatically construct hierarchical category systems to naturally adapt to long-tailed data distributions.

[0014] The core idea of ​​this invention is to mimic the "top-down, hierarchical subdivision" paradigm of biological taxonomy. First, the model divides the data into several major "superclasses" based on significant feature differences (e.g., distinguishing different tissue types). Then, within each superclass, further subdivisions are made based on more subtle feature differences (e.g., distinguishing benign lesions from cancers of varying degrees of malignancy). This process is recursively repeated until no further subdivision is possible. This structure eliminates the need to pre-define the specific number of tail categories and provides semantic constraints from their parent classes to data-scarce tail categories through hierarchical relationships, thereby improving recognition stability.

[0015] Therefore, this invention proposes a novel category discovery method based on hierarchical clustering trees. The core of this method is the dynamic generation of a hierarchical clustering tree during the training phase. Its construction process is consistent with biological classification logic, spontaneously revealing the inherent category hierarchy from head to tail in an unlabeled image set. To accurately capture local features and contextual relationships crucial for distinguishing fine-grained tail categories during the construction of the hierarchical clustering tree, this invention introduces spatial attention and self-attention mechanisms, enabling the image classifier to focus on discriminative regions in the image. Finally, during the training phase, this invention designs an iterative optimization mechanism, allowing feature representation learning and hierarchical clustering tree construction to mutually promote and evolve, gradually refining the boundaries of the entire category system.

[0016] In the method of this invention, the following new category discovery task scenarios are defined: The input consists of a set of images of known classes and a set of images of unknown classes. The output is an image classifier (which may contain a feature extraction network or a classification layer) and a hierarchical clustering tree. This image classifier can correctly classify images of unknown classes (which may come from the set of images of unknown classes or not) to the labels corresponding to the leaf nodes of the hierarchical clustering tree. These labels include the labels from the set of images of known classes as well as the new labels generated during training.

[0017] The novel category discovery method based on hierarchical clustering proposed in this invention (AHCT-NCD, Attentional Hierarchical Clustering Tree for NCD) framework is as follows: Figure 1As shown, its core is an iterative process that alternates between joint representation learning and hierarchical clustering tree construction.

[0018] Specifically, this invention discloses a method for discovering new image categories, comprising the following steps: extracting image features of the image to be classified through a feature extraction network and feeding them into an image classifier; the image classifier calculating the feature centers of the image features; determining the label of the image to be classified based on the feature centers and the class prototypes of preset categories in the image classifier; wherein, a training set is constructed based on known class images and unknown class images, and the feature extraction network and the image classifier are jointly trained based on the training set, and the number of new categories is not pre-set in the joint training.

[0019] This invention can discover more new categories during the training process of image classifiers and feature extraction networks, and can improve image classification performance when faced with complex data consisting of a large number of images with unknown "tail" categories.

[0020] The overall framework of AHCT-NCD is divided into training and testing phases. It should be noted that during the training phase, the hierarchical clustering tree does not require pre-setting the number of clusters, i.e., the number of leaf nodes.

[0021] During joint training, a class prototype network generates class prototypes for known class images; a feature extraction network extracts image features from unknown class images and calculates their feature centers; the similarity between the feature centers and each class prototype is calculated; when all similarities are less than or equal to a similarity threshold, the corresponding image features are assigned to a new category, resulting in a new category set including known classes and new categories; all categories in the new category set are traversed, and training ends when all categories meet the termination condition. If a category in the new category set does not meet the termination condition, that category is split into K new categories, where K is an integer greater than or equal to 2.

[0022] The termination condition is |S|< , Furthermore, the number of splits is greater than the split threshold, where S represents the number of image features within that category. Represents the image feature threshold. This represents the set of image features within this category. This represents the i-th image feature within this category. This represents the variance of image features within that category. This represents the variance threshold.

[0023] The image classifier iteratively learns feature representations of known class images in the known class image set and unknown class images in the unknown class image set, and constructs a hierarchical clustering tree to generate pseudo-labels for unknown class images. Inputting known and unknown class images, a feature extraction network (i.e., a backbone network with spatial attention) extracts image features. Then, based on the image features of the known class images, the class prototype of the known class is calculated. A hierarchical clustering tree is constructed using all image features (including those of the known and unknown class images). The leaf nodes in the hierarchical clustering tree correspond one-to-one with labels, which include labels for the known classes in the known class image set and labels for the unknown classes generated during training from the unknown class images.

[0024] The feature extraction network consists of a serially connected backbone network and a spatial attention module. The network aims to learn a general feature space that is highly discriminative for both known and unknown classes. During the feature learning phase, the network focuses on the discriminative parts of objects rather than the background or uninformative regions. Given an input image, a primary feature map is first extracted through a backbone network. C represents the number of channels in the feature map, H represents the height of the feature map, and W represents the width of the feature map. In other words, the backbone network is used to extract the primary feature map of the image.

[0025] The spatial attention module is used to perform global max pooling and global average pooling on the primary feature map to obtain two spatial feature descriptors. The two spatial feature descriptors are then concatenated to generate a spatial attention feature map. Finally, the spatial attention feature map and the primary feature map are residually connected to obtain the image features.

[0026] Specifically, the initial feature map is fed into a Spatial Attention Module (SAM). The SAM first obtains two spatial feature descriptors through cross-channel global max pooling and global average pooling. and The concatenation of these elements is then processed by a 7×7 convolution and a sigmoid activation function to generate a spatial attention map. The final feature map is calculated using the following formula: (4) in, This represents element-wise multiplication, where F is processed by global average pooling (GAP) to obtain the image features. .

[0027] Image features Through two different projection heads: (1) Classification head (i.e., image classifier): a linear layer that outputs the probabilities of known classes, used to calculate the supervised loss (i.e., cross-entropy loss). .

[0028] (2) Projector head A multilayer perceptron (MLP) projects image features into a low-dimensional space. This is used for comparative loss. The standard InfoNCE loss is used. For image pairs within a batch... If both are known classes and have the same label, they are considered positive sample pairs; if both are unlabeled data, the pseudo-label generated in the current iteration is used to determine whether they are positive sample pairs; otherwise, they are considered negative sample pairs.

[0029] In other words, after performing a residual connection between the spatial attention feature map and the primary feature map, the method further includes: passing the image features through a projection head to obtain a low-dimensional spatial feature map, which is used to calculate the contrastive loss in joint training.

[0030] First, a pseudo-label needs to be assigned to each input image of an unknown class. This process is implemented based on a class prototype network, which first calculates the class prototype for each known class, such as... Represent the class prototype of the c-th known class, and then calculate the feature centers of the image features of the image of the unknown class. Cosine similarity to each class prototype ,when When an image of an unknown class is assigned to the c-th known class, the image features are assigned to that class if the cosine similarity between the feature center and all class prototypes is less than or equal to the similarity threshold. If the corresponding image features are assigned to a new unknown class, they will be used as pseudo-labels for the corresponding unknown class images.

[0031] It should be noted that, This represents the image features of the i-th unknown class image obtained through the feature extraction network. When multiple cosine similarities are greater than a similarity threshold... If the cosine similarity is the highest, then the known class corresponding to that highest cosine similarity is selected as the category of the image feature, and used as the label for that category. Thus, the label, or pseudo-label, for each image of an unknown class can be obtained.

[0032] In joint training, the method for calculating the class prototype of each known class is as follows: After each training epoch, the model parameters are frozen, and all known class training data are iterated through, based on... Calculate the class prototype of class c. ;in, This represents the number of image features in class c, where i represents the index of the image feature. This represents the i-th image feature in class c. express The tag, This represents a feature extraction network. These prototypes represent the parameters of the feature extraction network. As anchor points for known classes in the feature space, they are used to guide splitting decisions in subsequent hierarchical clustering.

[0033] To make the computational class prototype more representative of the core features of the class and reduce the influence of noisy samples, the method of this invention uses a cross-attention mechanism. Using the query vector Query(Q), the key vector Key(K) of all image features in class c, and the value vector Value(V) of all image features in class c, cross-attention enhancement is performed to obtain the enhanced class prototype. .

[0034] By using a cross-attention layer, the class prototype "queries" and aggregates the most representative information from all images: (6) This optimized class prototype It is more representative and can more accurately guide the decision of "known class or new class" in the hierarchy tree.

[0035] This method iteratively optimizes the image features of known and unknown class images into a clustering tree, where the leaf nodes correspond to the final class assignments. When deciding "how to split" a node, it no longer relies solely on linear, unsupervised dimensionality reduction methods like PCA, but instead utilizes self-attention to learn the complex relationships between data points within a node, finding the optimal split boundary. Splitting the class into K new classes includes: The input image encoder contains a self-attention mechanism and generates a set of encoded image features. ; This represents the i-th encoded image feature within the category; based on the image features within the category, a K-component Gaussian mixture model is fitted, the number of cluster components is set to the corresponding number of groups, and the current category is split into K new categories according to the belonging probability.

[0036] Specifically, the image feature set within the leaf nodes It is treated as a sequence, input to a lightweight Transformer encoder. The self-attention mechanism can model the relationships between images and output a context-enhanced image feature set. .exist We fit a two-component Gaussian mixture model (GMM) (n_components=2, i.e., splitting into 2 new nodes). Based on the probability given by the GMM that the image features belong to the next level leaf node, we split the current leaf node into two new leaf nodes.

[0037] The process of constructing a hierarchical clustering tree is as follows, where each leaf node is a data structure containing its image index, depth, child node pointers, whether it is a leaf node, and label. A leaf node is processed using a recursive splitting algorithm. Then, a termination condition check is performed. If the termination condition is not met, a split is executed. For the two newly generated child nodes, the `build_tree` function is recursively called.

[0038] Finally, after the tree is constructed, all leaf nodes are traversed to generate pseudo-labels. Images in leaf nodes labeled as belonging to a known class are assigned the known class label; leaf nodes labeled as candidates for a new class are assigned a unique pseudo-label ID (Novel_0, Novel_1, ...). These pseudo-labels will be used for the contrastive loss in the next round of training. calculate.

[0039] The characteristics of the method of the present invention are as follows: (1) General feature learning and prototype initialization. A powerful feature extraction network is learned on known classes, and a "prototype" for each known class is calculated as a representative anchor point for that class in the feature space.

[0040] (2) Top-down hierarchical clustering. The image features of the unknown class images are projected into the feature space, and starting from a root node containing all image feature data, the information provided by the class prototypes of the known classes is used to guide the recursive binary or K-separation to form a hierarchical clustering tree.

[0041] (3) Fine-grained category discovery and verification. Reliability assessment is performed on each leaf node (i.e. the finest-grained cluster) in the hierarchical clustering tree, and those clusters that are sufficiently different from known classes and have high internal consistency are identified as "new categories".

[0042] The total loss function consists of supervised loss and contrastive loss, expressed as: (1) in, It is a supervised loss, used for image sets with known classes; It is a contrastive loss used for image sets of known and unknown classes, which are based on pseudo-labels generated from hierarchical clustering trees. It is a hyperparameter used to balance the contributions of the two parts of the loss.

[0043] (1) Monitoring loss .

[0044] Supervised loss is based on the labels of images of known classes, ensuring that the image classifier is discriminative in extracting features from images of known classes. Cross-entropy loss is typically used. This represents the image features output by the feature extraction network (with spatial attention SAM). The parameters of the feature extraction network are then passed through an image classifier (which could also be a classification layer, such as a fully connected layer) to obtain the labels of the known classes. The supervised loss is defined as: (2) in, It is the number of known class images in the joint training. It is an image classifier (or classification layer function if it is a classification layer) used to map image features to a set of images of known classes. probability distribution on, This represents a feature extraction network. This represents the i-th known class image. These represent the parameters of the feature extraction network. express The label CE stands for Cross-Entropy Loss Function.

[0045] (2) Comparison of losses .

[0046] Contrastive loss aims to enhance the discriminativeness of the feature space through pseudo-labels, bringing image features with the same pseudo-label (from the leaf nodes of a hierarchical clustering tree) closer together and image features with different pseudo-labels further apart. Here, a pseudo-label-based contrastive loss (such as InfoNCE loss) is used. express Image features (i.e.) Contrast loss is defined as: (3) in, This represents the sum of the number of known and unknown class images during joint training. It is related to the nth image A set of image indices with the same label (i.e., positive sample pairs). It's a temperature parameter that controls the sensitivity of the loss function. express Low-dimensional spatial feature map, Show medium image The low-dimensional spatial feature map, i.e. China is different The image, express The low-dimensional spatial feature map, that is, the feature map that is different from all images. The images are generated from leaf nodes of a hierarchical clustering tree. Each leaf node is labeled with either a known class or a new class (a candidate new class before the end of the iteration during training). Therefore, each image is assigned a pseudo-label.

[0047] In each iteration of the training phase, the labels of the known classes are used to compute... Calculate using known class labels and pseudo-labels By minimizing The parameters of the feature extraction network are updated. Since the hierarchical clustering tree is rebuilt in each iteration, the pseudo-labels are updated as the feature space changes, and therefore the contrastive loss also changes dynamically. The contrastive loss depends on the accuracy of the pseudo-labels, so it may be inaccurate in the initial iterations, but through iterative optimization, image features and clustering will improve each other. This loss function design ensures discriminative learning of known classes and discovery of new classes, while providing structured pseudo-label guidance through hierarchical clustering.

[0048] The specific process of the new category discovery method based on hierarchical clustering is shown in Table 1.

[0049] Table 1. Novel Category Discovery Methods Based on Hierarchical Clustering During the testing phase, known-class and unknown-class images are input. Features are extracted through a pre-trained backbone network with spatial attention, and then the leaf nodes are traversed through the final hierarchical clustering tree to determine the category and obtain the label.

[0050] In summary, this invention proposes a novel category discovery method based on hierarchical clustering trees and attention mechanisms. This method constructs a systematic discovery framework through three core modules: First, a spatial attention module is introduced to enhance the discriminative representation of features and focus on key region information; second, a cross-attention mechanism is used to optimize known class prototypes, enabling them to more accurately guide the clustering direction; finally, a self-attention mechanism is added to the construction of the hierarchical clustering tree to achieve more robust recursive clustering by enhancing feature associations, thereby enabling the automatic discovery of potential unknown category structures without pre-setting the number of new classes.

[0051] The entire training process employs an iterative optimization mechanism, enabling feature learning and clustering results to mutually promote and evolve collaboratively, forming a dynamically optimized discovery system. Experimental results on publicly available benchmark datasets demonstrate that the method of this invention achieves significant improvements in the identification of both known and novel classes, validating its effectiveness and superiority. By constructing a hierarchical clustering tree, the method of this invention achieves a breakthrough in the discovery of new categories in long-tailed data distributions.

[0052] To verify the effectiveness of the method of the present invention, the following experimental evaluation was conducted.

[0053] 1) Experimental setup The experimental environment for the AHCT-NCD method is listed in Table 2.

[0054] Table 2 Experimental Hardware and Software Environment 2) Selection of backbone network.

[0055] To evaluate the performance of different backbone networks within the AHCT-NCD framework, this experiment selected representative convolutional neural networks and the Transformer architecture for systematic comparison. All experiments were conducted under the same dataset partitioning, hyperparameter settings, and training strategies to ensure fairness in the comparison.

[0056] Table 3 shows the performance comparison results of different backbone networks on the three datasets. ResNet-50 and ResNet-101 show improved performance compared to ResNet-18 on CIFAR-10 and CIFAR-100, but the improvement is small because the CIFAR dataset is relatively simple, and the benefit of increasing model size is limited. On ImageNet-100, ResNet-101 shows a slight improvement over ResNet-50. ViT-B / 16 and ViT-L / 16 perform better on ImageNet-100 because ViT models are suitable for large-scale data; however, they perform slightly worse than ResNet on the CIFAR dataset because ViT requires a large amount of data to realize its advantages and is prone to overfitting on small datasets.

[0057] Table 3. Accuracy using different backbone networks (in %) ResNet-18, one of the lightest versions in the ResNet series, contains 18 convolutional layers and approximately 11.7 million parameters. It employs a basic residual block structure, with each block containing two 3×3 convolutional layers. Despite its smaller model capacity, ResNet-18 offers significant advantages in computational efficiency, making it suitable for resource-constrained applications. Including it in comparative experiments helps evaluate the applicability of the AHCT-NCD framework to lightweight architectures.

[0058] ResNet-50 is a 50-layer version of the residual network, achieving a good balance between feature extraction capability and computational efficiency by introducing a bottleneck module. This network contains approximately 25.6 million parameters, and its mature architecture and abundant pre-training resources make it widely used as a benchmark network for computer vision tasks.

[0059] As a deeper version of the ResNet series, ResNet-101 further enhances feature representation capabilities by increasing the number of layers. Compared to ResNet-50, ResNet-101 adds more residual blocks in the third and fourth stages, reaching a total of 44.5 million parameters. This network is suitable for evaluating the impact of deep feature extraction on the performance of novel feature discovery.

[0060] ViT-Base is a foundational version of the Visual Transformer, successfully applying a pure Transformer architecture to image classification tasks for the first time. It segments the input image into fixed-size patches and obtains serialized token embeddings through linear mapping. ViT-Base contains 12 Transformer encoder layers with approximately 86.6 million parameters and captures long-range dependencies through a global attention mechanism.

[0061] ViT-Large is a larger-scale version of ViT, containing 24 Transformer encoder layers with 307.4 million parameters. Its deeper architecture and more attention heads give it stronger representation learning capabilities, making it particularly suitable for evaluating the impact of large-scale feature representations on the discovery of new categories.

[0062] 3) Comparative experiment.

[0063] To systematically evaluate the performance of the method of this invention on the new category discovery task and verify its effectiveness, comparative experiments were designed and conducted on three general image recognition datasets. Under the standard new category discovery setting, the system compared the performance of representative state-of-the-art methods in known category recognition, new category discovery, and overall classification accuracy, verifying the advantages of the method of this invention over existing methods. Furthermore, the quantitative results were used to analyze the characteristics and limitations of different methods when handling known and unknown categories.

[0064] Considering computational efficiency and broad comparability, ResNet-18 was used as the backbone network for feature extraction on both CIFAR-10 and CIFAR-100, while ResNet-50 was used on ImageNet-100 to match its higher image complexity and feature learning requirements. The comparative experimental results are shown in Table 4.

[0065] FixMatch is a classic semi-supervised learning method that uses Wide ResNet as its backbone and combines consistency regularization and pseudo-labeling techniques to utilize unlabeled data. This method generates pseudo-labels for weakly augmented unlabeled images, filters high-confidence predictions using a confidence threshold, and then enforces consistency constraints on strongly augmented versions, thereby improving model performance. Although FixMatch primarily targets semi-supervised learning for known categories, its concise and effective design provides an important reference for research on new category discovery.

[0066] DS3L is a safe deep semi-supervised learning method specifically designed for handling unlabeled data containing unknown categories. This method employs a two-layer optimization framework, using two CNN layers on MNIST and a Wide ResNet-28-10 as the backbone on CIFAR-10. Its core idea is to learn a weight function that selectively utilizes unlabeled data, ensuring that the model's performance is no less than that of supervised learning using only labeled data, thus safely handling unlabeled data with unknown categories.

[0067] CGDL proposes a learning framework that combines classification and reconstruction. This method utilizes an encoder-decoder structure to simultaneously optimize both the classification loss and the input reconstruction loss, thereby learning feature representations that are both discriminative and generalizable. When discovering new categories, CGDL determines the category of a test sample by jointly evaluating its classification confidence and reconstruction error: if a sample has low confidence and high reconstruction error, it is classified as an unknown category.

[0068] DTC proposes a transfer learning framework based on Deep Embedded Clustering (DEC) for automatically discovering new categories in unlabeled data. This method first pre-trains a feature extraction network on labeled data with known categories. Then, by introducing representational bottlenecks, temporal ensembles, and consistency constraints, it jointly optimizes feature representation and cluster assignment on unlabeled data. DTC does not rely on a pre-defined number of categories; instead, it dynamically estimates the number of unknown categories by dividing a "probe set" from known categories and combining clustering quality metrics and clustering accuracy. Ultimately, this achieves effective classification and identification of unknown categories.

[0069] RankStats uses ResNet-18 (with a VGG-like network on OmniGlot) as its backbone. It first performs self-supervised pre-training on all data (labeled and unlabeled) to obtain unbiased feature representations. Then, it generates pseudo-labels for image pairs based on the consistency of the Top-k activation components of the feature vectors. Finally, it trains a classification head specifically for unlabeled data using binary cross-entropy loss. The model then achieves efficient discovery and representation learning of new categories by jointly optimizing the classification loss for labeled data and the clustering loss for unlabeled data.

[0070] SimCLR uses the standard ResNet as its backbone network. Instead of relying on a specific architecture or memory library, it learns representations invariant to image augmentation transformations by combining various data augmentations, a basic encoder, a nonlinear projection head, and normalized temperature-scaled cross-entropy loss (NT-Xent). Although SimCLR itself does not directly discover new categories, its learned general visual feature representations can serve as high-quality representations, providing a strong feature foundation for downstream clustering or new category recognition tasks, indirectly supporting the identification and classification of new categories.

[0071] OpenLDN proposes an open-world semi-supervised learning framework based on pairwise similarity loss, aiming to automatically discover and identify new categories from unlabeled data. This method uses ResNet-18 (with ResNet-50 on ImageNet-100) as the backbone for feature extraction, employs an additional similarity prediction network to estimate whether sample pairs belong to the same category, and leverages a two-stage optimization strategy to utilize known category annotation information to guide the discovery of new categories. During training, OpenLDN combines cross-entropy loss, entropy regularization, and pairwise similarity loss to simultaneously achieve classification of known categories and clustering of new categories. After discovering new categories, the method further transforms the open-world problem into a closed-set semi-supervised learning problem through iterative pseudo-labeling, thereby leveraging existing closed-set methods to further improve performance.

[0072] OpenCon employs ResNet as its backbone network for feature extraction (e.g., using ResNet-18 on CIFAR-100 and ResNet-50 on ImageNet-100). Methodologically, its core is a prototype-driven contrastive learning framework. It distinguishes known from unknown samples by calculating the similarity between unlabeled samples and known class prototypes, and assigns pseudo-labels to unknown samples to construct positive and negative pairs for contrastive learning. This framework achieves end-to-end training by jointly optimizing supervised contrastive loss (for labeled data), self-supervised contrastive loss (for all unlabeled data), and its newly proposed prototype-based contrastive loss (for unknown data) to learn feature representations that are highly separable for both known and unknown classes.

[0073] NACH proposes a robust method for semi-supervised learning scenarios where "not all categories have labels," aiming to simultaneously identify known categories and discover unknown categories. This method uses ResNet-18 (CIFAR dataset) or ResNet-50 (ImageNet-100) as the backbone network and first performs contrastive learning pre-training via SimCLR to obtain good feature representations. For new category discovery, NACH designs an unsupervised loss based on pairwise similarity, filtering out erroneous sample pairs that may originate from known and unknown classes through a selection strategy, and using binary cross-entropy loss to bring similar samples closer together, thus achieving automatic clustering of unknown classes. Furthermore, to alleviate the problem of inconsistent learning progress between known and unknown classes, NACH introduces an adaptive threshold and distribution alignment mechanism, dynamically adjusting the selection threshold for pseudo-labels of unknown classes, thereby improving overall classification performance.

[0074] Table 4 Comparison results of general image recognition datasets (unit: %) As shown in Table 4, the method of the present invention (AHPT-NCD) has achieved highly competitive results in terms of average accuracy on known classes and new classes in the three datasets, demonstrating good generalization ability.

[0075] On the CIFAR-10 dataset, AHPT-NCD achieves the highest performance (95.0%, 93.4%, 94.4%) on known classes, new classes, and all classes, significantly outperforming recent methods OpenCon and NACH. This demonstrates that on relatively simple datasets, the attention mechanism and hierarchical clustering tree structure integrated in this invention can effectively uncover the class structure in the data.

[0076] On the more challenging CIFAR-100 dataset, the method of this invention leads with an accuracy of 55.6% across all classes. Furthermore, it surpasses OpenCon (47.8%) and NACH (47.0%) in the accuracy of discovering new classes (48.7%), demonstrating that the method of this invention maintains robust new class discovery capabilities even with an increasing number of classes and finer-grained inter-class differences.

[0077] On the large-scale ImageNet-100 dataset, the method of this invention achieved the highest accuracy for new classes (82.4%) and all classes (85.3%). This result verifies the scalability and effectiveness of the method on large-scale, real-world data. The accuracy for known classes (91.4%) is comparable to state-of-the-art methods (NACH: 91.0%, OpenCon: 90.6%), indicating that while retaining the classification ability for known classes, the method of this invention has superior open-world recognition capabilities for new classes.

[0078] Compared to earlier representative baseline methods, the advantages of the method in this invention are more obvious. SimCLR and FixMatch perform poorly on novel class discovery tasks, highlighting the necessity of designing dedicated novel class discovery methods. The performance of dedicated novel class discovery methods such as DTC and RankStats has been significantly surpassed by recent methods, reflecting the rapid development of this field. Building on this, the method in this invention further improves the performance boundary by introducing an attention mechanism and a hierarchical clustering strategy.

[0079] Compared to the two state-of-the-art methods, OpenCon and NACH, which have the closest performance, the method of this invention achieves higher accuracy in discovering new classes on all three datasets. This is mainly due to the hierarchical prototype tree and the self-attention-enhanced clustering mechanism, which can more finely characterize the internal structure of unknown classes and reduce misclassification. The method of this invention not only leads in discovering new classes but also maintains a leading position in the overall evaluation metric of discovering all classes. This demonstrates that the method of this invention improves the ability to discover new classes without sacrificing the recognition accuracy of known classes, achieving a balanced performance improvement. This is attributed to the overall enhancement of feature discriminative power by the spatial attention and prototype optimization modules.

[0080] In summary, the AHPT-NCD method proposed in this invention achieves state-of-the-art performance on multiple standard datasets. Experimental results validate the effectiveness of the core design idea: improving feature quality and prototype representativeness through an attention mechanism across multiple dimensions, and combining a hierarchical, adaptive clustering framework to discover unknown category structures. Compared with existing methods, this invention significantly improves the accuracy of discovering and discriminating new categories while maintaining the ability to identify known categories, demonstrating stronger practicality and robustness.

[0081] 4) Ablation experiment.

[0082] To verify the effectiveness and contribution of the three attention modules in the method of this invention, ablation experiments as shown in Table 5 were conducted. The specific implementation is as follows: Spatial attention removal: The spatial attention module in the backbone network was removed, and the original feature map was used directly for subsequent operations.

[0083] Remove cross attention: When calculating the prototype of a known class, do not use cross attention for optimization, and directly use the average feature as the prototype.

[0084] Remove self-attention: When constructing a hierarchical clustering tree, do not use self-attention to enhance features, but directly use the original features for GMM clustering.

[0085] Based on the ablation experiment results shown in Table 5, it can be concluded that each attention module contributes to the performance improvement. In the complete model (AHCT-NCD), when spatial attention, cross-attention, and self-attention modules are used simultaneously, the model achieves the best performance in both known and new class recognition. This indicates that the combined effect of these three modules can significantly improve the model's representation and discrimination capabilities. Specific analysis follows: Table 5. Ablation experimental analysis of each component of the attention mechanism on the CIFAR-10 dataset (unit: %) The spatial attention module can enhance feature discrimination. When using only spatial attention (first row), the model's performance is relatively low across all categories, indicating that its individual effect is limited. However, after removing the module (fourth row), the model's performance drops significantly compared to the complete model, decreasing by 2.8% and 2.9% on known and new classes, respectively. This demonstrates that spatial attention can effectively improve the response of key regions in the feature map and enhance feature discrimination between categories.

[0086] Self-attention is crucial for constructing hierarchical clustering trees. Using only self-attention (third row), the model performs well in new class identification (75.0%), significantly higher than configurations using only other single modules. Furthermore, removing self-attention (sixth row) causes the model's performance in new classes to drop sharply to 58.9%, demonstrating that self-attention plays an irreplaceable role in modeling feature relationships during the construction of hierarchical clustering trees, effectively improving the discriminative power of the cluster structure.

[0087] Cross attention can improve the representativeness of the prototype. After removing the cross attention module (fifth row), the model performance declined overall, decreasing by 2.1% and 1.1% on known classes and new classes, respectively. This indicates that optimizing the prototype of known classes through cross attention can enable the prototype to better summarize the distribution of class features, thereby improving the model's generalization ability to new classes.

[0088] In summary, the three modules of spatial attention, cross attention, and self-attention work together at three different levels—feature enhancement, prototype optimization, and hierarchical clustering tree construction—to improve the overall performance of the model in the new category discovery task.

[0089] Spatial attention modules act on local spatial regions of feature maps, and their effects are better visualized using heatmaps such as class activation maps (CAMs) to demonstrate their focus on key regions. Cross-attention modules involve the interaction between known class prototypes and sample features, and the analysis of their dynamic weights is more abstract. The core of the self-attention mechanism lies in modeling the long-term dependencies within the sample feature sequence, and its output directly determines the quality of the feature representation used for hierarchical clustering. This causal chain of "feature relationship modeling → cluster structure formation" is presented most intuitively and systematically through t-SNE dimensionality reduction visualization of feature distribution. This method clearly reveals the global structure of the feature space, which is highly consistent with the goal of verifying the crucial role of self-attention in constructing hierarchical clustering trees.

[0090] Furthermore, quantitative ablation results show that removing the self-attention module leads to the most significant decrease in model performance (new class accuracy plummets from 93.4% to 58.9%). This highlights the pivotal role of self-attention in the method of this invention. By enhancing the consistency within features, it is a prerequisite for forming an accurate and robust hierarchical clustering structure. Therefore, a thorough visual analysis of its working mechanism has irreplaceable explanatory value for understanding how the entire method achieves the core objective of new class discovery.

[0091] Self-attention is a crucial step in the feature preparation stage of constructing a clustering tree. Visualizing the difference in feature distribution with and without a self-attention mechanism directly demonstrates how this mechanism transforms "loose, confused features" into "compact, separable clusters," thus laying the foundation for subsequent hierarchical partitioning. This visualization and quantitative analysis complement each other, together forming a complete chain of evidence for the key role of the self-attention mechanism in the hierarchical new category discovery task. Therefore, to more intuitively and systematically verify the key role of the self-attention mechanism in constructing hierarchical clustering trees, Figure 2 The t-SNE visualization compares the feature distributions with and without the use of the self-attention mechanism.

[0092] like Figure 2 As shown in (a), after removing the self-attention mechanism, the features of different categories are loosely distributed in the low-dimensional space, the inter-class boundaries are blurred, and there are obvious overlapping and confusion areas. The icons in the figure, from top to bottom, are airplane, motorcycle, bird, cat, deer, dog, frog, horse, ship, and truck.

[0093] This phenomenon indicates that the lack of self-attention enhancement results in insufficient discriminative power of features, leading to samples with uncertain classifications during the clustering process, which in turn reduces the model's accuracy in identifying new categories.

[0094] In contrast, Figure 2In (b), after introducing the self-attention mechanism, the features of each category exhibit a compact and cohesive distribution pattern, with significant inter-class spacing and a clearly discernible cluster structure. This indicates that self-attention can effectively model long-range dependencies between samples and enhance the semantic consistency of features within the same category, thus providing a more discriminative representation basis for subsequent hierarchical clustering.

[0095] In summary, the visualization results are consistent with the quantitative analysis conclusions in Table 5, jointly confirming that the self-attention mechanism plays an indispensable role in the construction of hierarchical clustering trees by improving the separability and robustness of feature structures.

[0096] The present invention also discloses a new category discovery device for images, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the method described above.

[0097] The present invention also discloses an embodiment that provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the above-described method embodiments.

[0098] The present invention also provides a computer program product that, when run on a data storage device, enables the data storage device to implement the steps in the above-described method embodiments.

[0099] If the integrated unit module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a storage device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0100] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0101] Those skilled in the art will recognize that the algorithmic steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0102] It should be noted that the data used in the implementation of this invention were all collected or gathered through legal and compliant channels, and the collection and gathering activities fully comply with the requirements of relevant laws, regulations and industry standards; the existing technical methods involved in this invention were also obtained and used through legal and compliant means.

Claims

1. A method for discovering new categories of images, characterized in that, Includes the following steps: Image features of the image to be classified are extracted by a feature extraction network and then fed into an image classifier. The image classifier calculates the feature centers of the image features; The label of the image to be classified is determined based on the feature center and the class prototype of the preset category in the image classifier; In this process, a training set is constructed based on known class images and unknown class images. The feature extraction network and the image classifier are jointly trained based on the training set, and the number of new classes is not pre-set in the joint training.

2. The method for discovering new categories of images as described in claim 1, characterized in that, The joint training includes: Generate the class prototype of the known class corresponding to the known class image based on the class prototype network; Image features of unknown class images are extracted using a feature extraction network, and their feature centers are calculated. Calculate the similarity between the feature center and each class prototype; When all similarities are less than or equal to the similarity threshold, the corresponding image features are assigned to a new category, resulting in a new category set that includes known categories and the new category. Iterate through all categories in the new category set, and end training when all categories meet the termination condition; The termination condition is |S|< , Furthermore, the number of splits is greater than the split threshold, where S represents the number of image features within that category. Represents the image feature threshold. This represents the set of image features within this category. This represents the i-th image feature within this category. This represents the variance of image features within that category. This represents the variance threshold.

3. The method for discovering new categories of images as described in claim 2, characterized in that, If a category in the new category set does not meet the termination condition, the category is split into K new categories, where K is an integer greater than or equal to 2.

4. The method for discovering new categories of images as described in claim 3, characterized in that, Splitting this category into K new categories includes: Will The input image encoder contains a self-attention mechanism and generates a set of encoded image features. ; This represents the i-th encoded image feature within this category; Based on the image features within the category, a K-component Gaussian mixture model is fitted, the number of cluster components is set to the corresponding number of groups, and the current category is split into K new categories according to the belonging probability.

5. The method for discovering new categories of images as described in claim 2, characterized in that, The method for calculating class prototypes in the joint training includes: based on Calculate the class prototype of class c. ;in, This represents the number of image features in class c, where i represents the index of the image feature. This represents the i-th image feature in class c. express The tag, This represents a feature extraction network. These represent the parameters of the feature extraction network; by Using the query vector, all image features in class c as the key vector, and all image features in class c as the value vector, cross-attention enhancement is performed to obtain the enhanced class prototype. .

6. A method for discovering new categories of images as described in any one of claims 2-5, characterized in that, The feature extraction network includes a serially connected backbone network and a spatial attention module; The backbone network is used to extract primary feature maps of the image; The spatial attention module is used to perform global max pooling and global average pooling on the primary feature map to obtain two spatial feature descriptors. The two spatial feature descriptors are then concatenated to generate a spatial attention feature map. Finally, the spatial attention feature map and the primary feature map are residually connected to obtain the image features.

7. The method for discovering new categories of images as described in claim 6, characterized in that, After performing a residual connection between the spatial attention feature map and the primary feature map, the following is also included: Image features are passed through a projection head to obtain a low-dimensional spatial feature map, which is used to calculate the contrastive loss during joint training.

8. The method for discovering new categories of images as described in claim 7, characterized in that, In joint training, the loss function includes both supervised loss and contrastive loss; The supervised loss is calculated based on known image classes, specifically as follows: , in, Indicates monitoring losses, This represents the number of known class images in the joint training, and CE is the cross-entropy loss function. Represents an image classifier. This represents a feature extraction network. This represents the i-th known class image. These represent the parameters of the feature extraction network. express The tag.

9. A method for discovering new categories of images as described in claim 8, characterized in that, The comparison loss is specifically as follows: , in, Indicates comparative loss, This represents the sum of the number of known and unknown class images during joint training. Indicates the relationship with the nth image A set of image indexes with the same label express Low-dimensional spatial feature map, express medium image Low-dimensional spatial feature map, For hyperparameters, express The low-dimensional spatial feature map.

10. A new category discovery apparatus for images, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-9.