Open set image classification method based on feature enhancement and boundary optimization
By adopting feature enhancement and boundary optimization methods in image classification, the problem of difficulty in identifying and classifying unknown categories in the prior art is solved, efficient and accurate open-set image classification is achieved, manual labeling costs are reduced, and the robustness and real-timeness of the model are improved.
Patent Information
- Application Number
- CN202510151563.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
When the prior art faces complex and dynamic real-world data, it is difficult to effectively identify and classify unknown categories, and there are problems such as susceptibility to adversarial samples and limited data adaptability.
An open-set image classification method based on feature enhancement and boundary optimization is proposed. Features are extracted through the pre-trained ResNet50 network, combined with Fourier transform, CORAL distance and weighted mixup methods, edge generation adversarial network and region of interest extraction network are constructed, and nonlinear support vector machines are used to realize the identification and classification of unknown categories.
It improves the efficiency and accuracy of open set image classification, reduces the cost of manually labeling samples, realizes effective identification and classification of unlabeled categories in real-world scenarios, and improves the robustness and real-timeness of the model.
Smart Images

Figure CN120070993A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of image recognition and classification, and in particular to an open-set image classification method based on feature enhancement and boundary optimization. Background Art
[0002] With the rapid development of information technology, image recognition and classification technology has become one of the most active research directions in the field of artificial intelligence. The breakthrough of deep learning technology, especially the development of deep neural networks, has greatly promoted the progress of image recognition and classification technology, making it widely used in many fields such as security monitoring, autonomous driving, and medical diagnosis. However, these application scenarios often face a common challenge. The data distribution in the complex and dynamic real world is often open, that is, the test data may contain categories unknown in the model training stage, and this phenomenon is particularly prominent in image classification tasks.
[0003] Most traditional image classification methods are based on the closed-set assumption, that is, it is assumed that all possible categories are known in both the training and test stages. This method performs well under ideal conditions, but in the face of the complex and dynamic real world, its performance often drops significantly. This is because the closed-set method lacks the ability to recognize unknown categories and is prone to misclassify unknown categories as known categories, resulting in a decrease in the accuracy and robustness of classification.
[0004] To solve this problem, researchers have proposed the concept of open-set image classification. The goal of open-set classification is not only to accurately classify known categories but also to be able to identify unknown categories in the test set. This challenge requires the classifier to accurately distinguish between K known categories and an additional unknown category (i.e., the K + 1th category). The introduction of the open-set recognition framework has improved the performance of the classifier when facing new categories, which is of great significance for improving the value and applicability of image classification technology in practical applications.
[0005] Currently, relatively mature research on open-set recognition includes open-set recognition methods based on discriminant models, open-set recognition methods based on generative models, and open-set recognition methods based on multi-modalities, etc.
[0006] Although these methods have made some progress in dealing with unknown category data, they still face some challenges, such as the need for a large amount of known category data, limited adaptability to complex data, susceptibility to adversarial samples, insufficient representation of unknown categories, limited performance improvement, huge network parameter quantities, and high training costs, etc. Summary of the Invention
[0007] To solve the above technical problems, an open-set image classification method based on feature enhancement and boundary optimization is proposed in an embodiment of the present application, which can improve the efficiency and accuracy of open-set image classification, effectively reduce the cost required for manually labeled samples, and well achieve the effective recognition and classification of unlabeled categories in real-world scenarios.
[0008] To achieve the above object, an open-set image classification method based on feature enhancement and boundary optimization is proposed in an embodiment of the present application. The method includes: using a pre-trained ResNet50 network to extract features from sample images of known categories to obtain the basic features of the sample images of known categories; performing Fourier transform on the basic features of the sample images of known categories, calculating the Fourier phase information of the transformed basic features as the domain-internal invariant features of the category domain of the sample images of known categories; calculating the covariance matrix of the domain-internal invariant features of different category domains, and calculating the CORAL distance based on the covariance matrix to measure the differences between different category domains; performing data augmentation on the sample images of known categories through a weighted mixup fusion method, generating new sample images and their category labels by linearly combining different sample images of known categories and their category labels; constructing an edge generative adversarial network including a discriminator and a generator, and making the sample distribution of the sample images generated by the generator approach the real sample distribution through adversarial training based on the sample images; constructing and training a region of interest extraction network, traversing the entire sample image through a sliding window, and obtaining the top-k candidate regions most likely to contain the target according to the top-k strategy; mapping the top-k candidate regions most likely to contain the target from the original space to a higher-dimensional feature space, and realizing the recognition and classification of open-set targets in the sample images with the help of a non-linear support vector machine.
[0009] To achieve the above object, an embodiment of the present application proposes an open-set image classification system based on feature enhancement and boundary optimization. The system includes: a feature extraction module, configured to use a pre-trained ResNet50 network to extract features from sample images of known categories to obtain basic features of the sample images of known categories; a Fourier phase transformation module, configured to perform Fourier transform on the basic features of the sample images of known categories, calculate the Fourier phase information of the transformed basic features, and use it as the within-domain invariant features of the category domain of the sample images of known categories; a cross-domain covariance calculation module, configured to calculate the covariance matrix of the within-domain invariant features of different category domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different category domains; a weighted mixup transformation module, configured to perform data augmentation on the sample images of known categories through a weighted mixup fusion method, and generate sample images of new categories and their category labels by linearly combining different sample images of known categories and their category labels; an edge generative adversarial training module, configured to construct an edge generative adversarial network including a discriminator and a generator, and based on the sample images, make the sample distribution of the sample images generated by the generator approach the real sample distribution through adversarial training; a candidate box extraction module, configured to construct and train a region of interest extraction network, traverse the entire sample image through a sliding window, and obtain the top-k most likely candidate regions containing the target according to the top-k strategy; a support vector machine classification module, configured to map the top-k most likely candidate regions containing the target from the original space to a higher-dimensional feature space, and with the help of a non-linear support vector machine, realize the recognition and classification of open-set targets in the sample images.
[0010] To achieve the above object, an embodiment of the present application also proposes an electronic device. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute an open-set image classification method based on feature enhancement and boundary optimization as described above.
[0011] To achieve the above object, an embodiment of the present application also proposes a computer-readable storage medium storing a computer program, which when executed by a processor, can implement an open-set image classification method based on feature enhancement and boundary optimization as described above.
[0012] Optionally, using a pre-trained ResNet50 network to extract features from sample images of known categories to obtain basic features of the sample images of known categories includes:
[0013] Prepare a set of sample images of known categories, where each sample image of a known category is labeled with a category label. Denote the total number of categories in the set of sample images of known categories as N c ;
[0014] Load the pre-trained ResNet50 network and fine-tune the last several layers of the ResNet50 network to adapt to the set of sample images of known categories;
[0015] Input the sample images of known categories into the fine-tuned ResNet50 network for feature extraction, and obtain the feature vectors output by the fine-tuned ResNet50 network as the basic features of the sample images of known categories. Among them, the basic features of the sample images of known categories represent the representation of the sample images of known categories in the feature space.
[0016] Optionally, perform a Fourier transform on the basic features of the sample images of known categories, and calculate the Fourier phase information of the transformed basic features as the intra-domain invariant features in the category domain of the sample images of known categories, including:
[0017] Perform a Fourier transform on the basic features of the sample images of known categories through the following formula to obtain the frequency-domain representation of the basic features of the sample images of known categories:
[0018]
[0019] where H and W are the height and width of the basic features of the sample images of known categories respectively, and F (u,v) (x) is the frequency-domain representation of the basic features of the sample images of known categories;
[0020] Calculate the Fourier phase information of the corresponding channels based on the frequency-domain representation of the basic features of the sample images of known categories through the following formula:
[0021] x * =P (u,v) (x)=arctan[I (u,v) (x) / R (u,v) (x)]
[0022] where I (u,v) (x) is the imaginary part of F (u,v) (x), R (u,v) (x) is the real part of F (u,v) (x), and x * represents the calculated Fourier phase information;
[0023] Use to train the feature extractor, The class label representing the sample images of known classes, the FFT loss used when training the feature extractor is represented by the following formula:
[0024]
[0025] where, θ f is the learnable parameter of the feature extractor G f , E (x,y) represents taking the expectation, L cls is the cross-entropy loss, is the frequency-domain representation of the basic features of the sample image of the m-th known class, is the class label of the sample image of the m-th known class, d mmd (S i , S j ) is the MMD distance between the current class and other classes, i≠j, i represents the current class, j represents other classes, L fft is the FFT loss used when training the feature extractor;
[0026] Obtain the within-domain invariant features of each class domain based on the trained feature extractor.
[0027] Optionally, calculate the covariance matrix of the within-domain invariant features of different class domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different class domains, including:
[0028] Calculate the mean μ c of the within-domain invariant features of all class domains, μ c is represented by the formula:
[0029]
[0030] where, f(x m ) is the within-domain invariant feature of the class domain extracted from the sample image of the m-th known class;
[0031] Calculate the covariance matrix of the within-domain invariant features of different class domains through the following formula:
[0032]
[0033] where, C c represents the covariance matrix of the c-th class domain;
[0034] Calculate the CORAL distance between two different class domains through the following formula, and finally obtain the CORAL matrix D CORAL ;
[0035]
[0036] Among them, is the inverse square root of the matrix of C c , and C d represents the covariance matrix of the d-th class domain, D CORAL (c, d) represents the CORAL distance between the c-th class domain and the d-th class domain;
[0037] Based on D CORAL construct a regularization loss, and the regularization loss is expressed by the formula:
[0038] L CORAL = L cls + λD CORAL
[0039] Among them, L cls is the classification loss, and λ is the hyperparameter that weighs the CORAL distance regularization term.
[0040] Optionally, perform data augmentation on the sample images of known classes through the weighted mixup fusion method. By linearly combining the sample images and their class labels of different known classes, generate the sample images and their class labels of new classes, including:
[0041] Let the unknown new class be C + 1. For each known class, maintain a fixed-length queue q cm to store the features corresponding to the sample images of each known class, represents the L-th feature corresponding to the m-th sample image in the c-th class;
[0042] Perform in-class data augmentation on the sample images of each class. Each time, randomly select two sample images of the same class and The label of the generated new sample image is still
[0043] Adopt a cross-class sampling strategy to generate sample images of new classes by combining sample images between different classes. Denote the sample images of the new class as It is expressed by the formula:
[0044]
[0045] Among them, m c is the number of sample images of the c-th class, M is the total number of all sample images, X m is the feature matrix of the m-th sample image, τ is the global correlation coefficient, A cis the covariance matrix of the c-th category, ‖·‖ F denotes calculating the Frobenius norm of the matrix.
[0046] Optionally, construct a marginal generative adversarial network including a discriminator and a generator. Based on the sample images, through adversarial training, make the sample distribution of the sample images generated by the generator close to the real sample distribution, including:
[0047] Construct a marginal generative adversarial network composed of a discriminator D and a generator G, and convert the discriminator D into an N c +1 classifier to achieve the purpose of shrinking the classification decision boundary. The N c +1 class is used to determine whether the sample image is real data or fake data from the generator G;
[0048] Calculate the probability score that the sample image output by the classifier belongs to category j through the following formula:
[0049]
[0050] where c i (x) represents the corresponding value of the i-th node of the classifier;
[0051] Train the discriminator D. The loss function used to train the discriminator D is represented by the following formula:
[0052]
[0053] where D Loss is the loss function used to train the discriminator D, is the loss function corresponding to the classifier correctly classifying the known classes, E z~P(z) {log<P class [y = N c +1|G(z)]>} is the loss function for the classifier to correctly identify whether the sample image comes from real data or the generator G;
[0054] Train the generator G. Fix the parameters of the discriminator D and only update the parameters of the generator G. Let the last hidden layer of the classification network of the discriminator D be c’(x). The loss function used to train the generator G is represented by the following formula:
[0055]
[0056] where G Loss is the loss function used to train the generator G, represents the expected value of the real sample image x from D train , E z~P(z)denotes the expected value of the noise z sampled from the prior distribution P(z), and G(z) represents that the generator G receives the noise z and generates a fake sample image;
[0057] Perform adversarial training, alternately update the parameters of the discriminator D and the generator G until the sample distribution of the sample images generated by the generator G is close enough to the true distribution, and at the same time construct the adversarial loss term L Margin , L Margin = D Loss + G Loss .
[0058] Optionally, construct and train a region of interest extraction network, traverse the entire sample image through a sliding window, and obtain the top k candidate regions most likely to contain the target according to the top-k strategy, including:
[0059] Generate a total of 9 anchor boxes with 3 scales and 3 aspect ratios at each position of the feature map of the sample image to cover different scales and shapes of the target as much as possible;
[0060] Pass each anchor box through a 3×3 convolutional kernel and two 1×1 convolutional kernels in sequence. The first 1×1 convolutional kernel outputs the binary classification score of the anchor box, that is, the probability value that the anchor box may contain the target, and the second 1×1 convolutional kernel outputs the offset of the adjusted anchor box position and size, obtaining a number of candidate regions;
[0061] Perform non-maximum suppression on each candidate region, obtain the candidate region with the highest score for each anchor box, and at the same time remove other candidate regions with too high overlap with it until all candidate regions are judged;
[0062] Arrange all the remaining candidate regions in descending order of score, and according to the top-k strategy, select the top k remaining candidate regions with the highest scores as the candidate regions most likely to contain the target.
[0063] Optionally, map the top k candidate regions most likely to contain the target from the original space to a higher-dimensional feature space, and use a non-linear support vector machine to realize the recognition and classification of open-set targets in the sample image, including:
[0064] Map the feature maps corresponding to the top k candidate regions most likely to contain the target from the original space to a higher-dimensional feature space. The model expression of the hyperplane in the high-dimensional space is:
[0065] g(x) = ω T φ(x) + b
[0066] where ω T and b are the feature vectors of the hyperplane, and φ(x) is the representation of the sample image x after feature mapping;
[0067] Calculate the similarity between sample images in the high-dimensional space using the Gaussian kernel function. The sample image x i and the sample image x j The similarity between them is expressed by the formula as follows:
[0068]
[0069] where σ is the bandwidth parameter of the Gaussian kernel, and k(x i , x j ) is the similarity between the sample image x i and the sample image x j ;
[0070] Construct and train a non-linear SVM classifier. The SVM classifier is expressed by the formula as follows:
[0071]
[0072] s.t. y i [ω T φ(x) + b] ≥ 1, i = 1, 2,..., m
[0073] Its dual problem is expressed as:
[0074]
[0075] It can be calculated through the Gaussian kernel function as:
[0076] κ(x i , x j ) = <φ(x i ), φ(x j )> = φ(x i ) T φ(x j )
[0077] This non-linear classification problem can be expressed as solving:
[0078]
[0079] where α i is the Lagrange multiplier, and CR is the regularization parameter;
[0080] Adjust the parameters of the mapping function, kernel function, and the regularization parameter of SVM according to the results of performance evaluation to optimize the performance of the classifier. The loss function of the entire network is:
[0081] L all = L CR + εL CORAL +(1 - ε)L Margin
[0082] Among them, ε is a trade-off parameter used to trade off the influence of the boundary sample image and the sample image of the new category on the classifier.
[0083] An open-set image classification method based on feature enhancement and boundary optimization proposed in the embodiments of the present application uses the weighted Fourier characteristic transformation of the MMD distance to realize the expansion of domain-invariant characteristics, effectively extracts the inherent characteristics inside and between samples, reduces the within-class variance, increases the between-class variance, and enhances the model's recognition and utilization abilities of invariant characteristics. The improved weighted mixup method is used to quickly and highly quality augment the sample images of known categories, improve the diversity of sample images, simulate the sample images of unknown categories, and enhance the model's performance in a variable category environment. The required computing resources are small and the practicability is very strong. Through the candidate box extraction network and the non-linear support vector machine, the hyperplane between known categories and unknown categories is accurately divided, improving the real-time performance and robustness of the classifier and realizing the accurate division of multiple unknown categories. Finally, it effectively reduces the cost required for manually labeled samples and well realizes the effective recognition and classification of unlabeled categories in real-world scenarios. Description of the Drawings
[0084] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0085] Figure 1 is provided in an embodiment of the present application, which is a flowchart of an open-set image classification method based on feature enhancement and boundary optimization;
[0086] Figure 2 is provided in an embodiment of the present application, which is a schematic structural diagram of the entire network;
[0087] Figure 3 is provided in an embodiment of the present application, which is a result comparison diagram of the open-set image classification of the MS-COCO dataset by the present application and the OWOD model;
[0088] Figure 4 is provided in an embodiment of the present application, which is a result comparison diagram of the open-set image classification of the KITTI dataset by the present application and the OWOD model;
[0089] Figure 5 is provided in another embodiment of the present application, which is a schematic structural diagram of an open-set image classification system based on feature enhancement and boundary optimization;
[0090] Figure 6 It is a schematic structural diagram of an electronic device provided in another embodiment of the present application. Specific implementation manners
[0091] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be elaborated in detail below with reference to the accompanying drawings. In various embodiments of the present application, many technical details are proposed for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following various embodiments, the technical solutions claimed in the present application can still be implemented. The division of the following various embodiments is only for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Various embodiments can be combined and cross-referenced with each other on the premise of not being contradictory.
[0092] In recent years, researchers have proposed the concept of open-set image classification. The goal of open-set classification is not only to accurately classify known classes, but also to be able to identify unknown classes in the test set. This challenge requires the classifier to be able to accurately distinguish between the known K classes and an additional unknown class (i.e., the K+1th class). The introduction of the open-set recognition framework has improved the performance of the classifier when facing new classes, which is of great significance for improving the value and applicability of image classification technology in practical applications.
[0093] Currently, relatively mature research on open-set recognition includes open-set recognition methods based on discriminant models, open-set recognition methods based on generative models, and open-set recognition methods based on multi-modalities, etc.
[0094] For the open-set recognition method based on discriminant models, by constructing the average activation vectors and distance sets of each class, and using probability distributions (such as Weibull distribution) to fit the distributions of these activation vectors, the known classes and unknown classes can be distinguished. The OpenMax algorithm is a representative of this type of method. It utilizes the statistical characteristics of known classes to identify unknown classes by evaluating the probability that a test sample belongs to a known class.
[0095] For the open-set recognition method based on generative models, by generating data to enrich samples, the OSR problem is transformed into a K+1-class classification problem. Generative adversarial networks and autoencoders are used to generate open-set synthetic samples to improve the generalization ability of the classifier in the open space.
[0096] For the open-set recognition method based on multi-modalities, by combining image and semantic information and using multi-modal learning techniques to improve the expression ability of the model, the target detector can also perform recognition and localization for unknown classes.
[0097] Although these methods have made some progress in dealing with unknown-class data, they still face some challenges, such as the need for a large amount of known-class data, limited adaptability to complex data, vulnerability to adversarial samples, insufficient representation of unknown classes, limited performance improvement, huge network parameter quantities, and high training costs, etc.
[0098] Zhu Min et al. used Faster R-CNN as the benchmark network for model training. By using an improved region proposal network, the detection capabilities for both the foreground and background of the image were retained. Proposal bounding boxes with object scores higher than the preset threshold but not belonging to known classes were labeled as unknown classes. Meanwhile, a contrastive clustering loss was introduced to reduce the intra-class distance and increase the distance between different classes. The Weibull distribution was incorporated to model the probability density functions of different classes to distinguish known and unknown class objects. Although this method reduced the costs of manual annotation and network update, a certain amount of annotated data was still required during the training phase to train the model, especially in feature clustering and contrastive clustering loss. Also, although the Weibull distribution and feature clustering were introduced to handle unknown classes, the generalization ability of the model for completely unknown classes still needs to be verified, especially in diverse and complex real-world scenarios.
[0099] Sun Jinyong et al. proposed an open-set recognition method based on prototype contrastive learning, which maintains the classification ability for known-class samples while identifying unknown-class samples. Combining contrastive learning and the class prototype theory, an encoder and a projection network were used to model the open-set recognition problem. A prototype contrastive loss function was designed, and the model parameters were learned through the gradient descent method to minimize the distance between the sample and its corresponding class prototype, while maximizing the distance between the sample and other class prototypes. This method also proposed using the sample generation method OSR-Mix to generate unknown-class samples to effectively supplement unknown-class information during the model training process. The generation of mixed samples in the OSR-Mix method relies on random sampling. Due to randomness, the generated samples may not be diverse enough or of high quality, limiting the improvement amplitude of the recognition effect. Additionally, compared with traditional CNN methods, the training time of the OSR model based on contrastive learning has increased significantly. Therefore, this method is not applicable to some scenarios that require fast training.
[0100] Shangguan et al. proposed a network that decouples the pre-trained class and fine-tunes the class parameters, enabling the model to focus more on learning the features of new classes during training. This method modified the skip connection type between the encoder and decoder and introduced a unified decoder module, effectively alleviating the model's bias towards pre-trained classes (multi-sample base classes). However, the model of this method has a high complexity, including multiple self-attention layers. This complexity may lead to overfitting problems in the case of few samples, especially when the number of samples of new classes is very limited. And there are multiple hyperparameters involved in the paper, such as the weight coefficients in the decoupling module, etc., which need to be carefully adjusted to achieve the best effect, increasing the difficulty of model tuning.
[0101] The inventors of the present application found that the above methods generally have four technical problems.
[0102] First, the invariant features within samples and among different classes are ignored.
[0103] In the design and implementation of the above open-set recognition methods, they often focus on mining and utilizing the differential features between classes to achieve the differentiation and recognition of different classes. This strategy is effective in many cases because it can help the model capture the key information for differentiating different classes. However, this method also has some limitations. Especially when dealing with the common shallow semantic features shared by each class, such as basic features like the contour, texture, and shape of the object, these features are also crucial for open-set recognition. Ignoring these basic features may lead to insufficient generalization ability of the model when facing unknown classes, because unknown classes may have similarities with known classes in these basic features, and the model fails to capture these commonalities.
[0104] Second, the sample augmentation strategy is complex and time-consuming.
[0105] When the above open-set recognition methods perform sample data augmentation and expansion, they often use complex neural networks to fully learn the features of known classes. For example, encoder-decoder networks and prototype networks are used for data augmentation. Currently, most sample augmentation methods use generative adversarial networks to expand samples, but it is challenging during the training process. The adversarial training between the generator and the discriminator may lead to unstable training. Especially in complex model structures, problems such as mode collapse and gradient disappearance may occur, and the training process is very time-consuming and requires high-performance computing hardware.
[0106] Third, insufficient attention is paid to strengthening the decision boundary.
[0107] In the sample classification problem, the decision boundary is the boundary that separates different classes. The decision boundary can be drawn by subdividing the feature space and using a model to predict and classify each subdivision point. When determining the decision boundary by the above methods, most rely on existing classifiers. When the number of samples of some classes in the dataset is much larger than that of other classes, the model may be biased towards these classes, resulting in an overly large decision boundary. At the same time, noise or outliers in the training data can also cause the decision boundary to be overly large. When the decision boundary is overly large, unknown classes that were originally outside the decision boundary are very likely to be misclassified as known classes.
[0108] Fourth, most of the above methods focus on the recognition of unknown classes and cannot further classify the unknown classes in detail. For example, most methods can only label all unknown classes as "Unknown", but there are also different specific classifications within the unknown classes.
[0109] To solve the above technical problems, an embodiment of the present application proposes an open-set image classification method based on feature enhancement and boundary optimization, which is applied to an electronic device. The electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is taken as an example of a server for illustration. The implementation details of an open-set image classification method based on feature enhancement and boundary optimization proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.
[0110] The specific process of an open-set image classification method based on feature enhancement and boundary optimization proposed in this embodiment can be as Figure 1 shown and includes:
[0111] Step 101, use the pre-trained ResNet50 network to extract features from the sample images of known classes to obtain the basic features of the sample images of known classes.
[0112] In a specific implementation, after the server obtains the sample images of known classes, it is necessary to use the pre-trained ResNet50 network to extract features from the sample images of known classes to obtain the basic features of the sample images of known classes. It should be noted that networks such as VGG network, Inception network, and DenseNet can all replace the ResNet50 network to perform the feature extraction task (i.e., as a feature extractor).
[0113] In an example, the server needs to prepare a sample image set of known classes. Each sample image of a known class is labeled with a class label, and the total number of classes in the sample image set of known classes is denoted as N c. Subsequently, the pre-trained ResNet50 network is loaded, and the last several layers of the ResNet50 network are fine-tuned to adapt to the acquired sample image set of known categories. Finally, the sample images of known categories are input into the fine-tuned ResNet50 network for feature extraction, and the feature vectors output by the fine-tuned ResNet50 network are obtained as the basic features of the sample images of known categories. Among them, the basic features of the sample images of known categories represent the representation of the sample images of known categories in the feature space.
[0114] Step 102: Perform Fourier transform on the basic features of the sample images of known categories, and calculate the Fourier phase information of the transformed basic features as the domain-internal invariant features of the category domain of the sample images of known categories.
[0115] In a specific implementation, after the server obtains the basic features of the sample images of known categories, it needs to perform Fourier transform on the basic features of the sample images of known categories and calculate the Fourier phase information of the transformed basic features as the domain-internal invariant features of the category domain of the sample images of known categories.
[0116] In an example, the server first performs Fourier transform on the basic features of the sample images of known categories through the following formula to obtain the frequency-domain representation of the basic features of the sample images of known categories:
[0117]
[0118] where H and W are the height and width of the basic features of the sample images of known categories respectively, and F (u,v) (x) is the frequency-domain representation of the basic features of the sample images of known categories, and (u, v) represents the index.
[0119] Next, the server needs to calculate the Fourier phase information of the corresponding channel based on the frequency-domain representation of the basic features of the sample images of known categories through the following formula:
[0120] x * =P (u,v) (x)=arctan[I (u,v) (x) / R (u,v) (x)]
[0121] where I (u,v) (x) is the imaginary part of F (u,v) (x), R (u,v) (x) is the real part of F (u,v) (x), and x * represents the calculated Fourier phase information.
[0122] After obtaining the Fourier phase information, the server needs to use Train the feature extractor (i.e., the ResNet50 network in step 101). Denote the class labels of the sample images of known classes. The FFT loss used when training the feature extractor is expressed by the following formula:
[0123]
[0124] where θ f is the learnable parameter of the feature extractor G f , E (x,y) denotes taking the expectation, L cls is the cross-entropy loss, is the frequency-domain representation of the basic features of the sample image of the m-th known class, is the class label of the sample image of the m-th known class, d mmd (S i , S j ) is the MMD distance between the current class and other classes, i≠j, i represents the current class, j represents other classes, and L fft is the FFT loss used when training the feature extractor.
[0125] Finally, the server obtains the within-domain invariant features of each class domain based on the trained feature extractor. The trained feature extractor can effectively extract the within-domain invariant features of each class domain, and such a feature extractor can be applied to the classification and recognition tasks in the subsequent steps.
[0126] Step 103: Calculate the covariance matrix of the within-domain invariant features of different class domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different class domains.
[0127] In a specific implementation, after the server obtains the within-domain invariant features of the class domains of the sample images of known classes, it needs to calculate the covariance matrix of the within-domain invariant features of different class domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different class domains.
[0128] In an example, the server first needs to calculate the mean μ c of the within-domain invariant features of all class domains, μ c is expressed by the formula:
[0129]
[0130] where f(x m ) is the within-domain invariant feature of the class domain extracted from the sample image of the m-th known class.
[0131] Next, the server calculates the covariance matrix of the domain-internal invariant features of different category domains through the following formula:
[0132]
[0133] where C c represents the covariance matrix of the c-th category domain.
[0134] After that, the server calculates the CORAL distance between two different category domains through the following formula, and finally obtains the CORAL matrix D between all category domains CORAL ;
[0135]
[0136] where is the inverse matrix square root of C c , C d represents the covariance matrix of the d-th category domain, and D CORAL (c, d) represents the CORAL distance between the c-th category domain and the d-th category domain.
[0137] Finally, the server constructs a regularization loss based on D CORAL The regularization loss is expressed by the formula as:
[0138] L CORAL = L cls + λD CORAL
[0139] where L cls is the classification loss, and λ is a hyperparameter that weighs the CORAL distance regularization term.
[0140] Based on the CORAL distance (CORAL matrix), the server details the statistical differences between the covariance of the domain-internal invariant features of different category domains and applies them to model training.
[0141] It should be noted that other cross-domain covariance calculation and distribution adaptation methods can also measure the differences between different category domains, such as transfer component analysis (TCA), Brown distance covariance (BDC), and joint distribution adaptation (JDA), etc.
[0142] Step 104, perform data augmentation on the sample images of known categories through the weighted mixup fusion method, and generate sample images of new categories and their category labels by linearly combining different sample images of known categories and their category labels.
[0143] In a specific implementation, after the server finishes measuring the differences between different category domains, it needs to enter the new category sample image generation stage. The sample images of known categories are data-augmented through the weighted mixup fusion method, and new category sample images and their category labels are generated by linearly combining different known category sample images and their category labels.
[0144] In one example, let the unknown new category be C+1. For each known category, a fixed-length queue q is maintained cm to store the features corresponding to the sample images of each known category, denotes the L-th feature corresponding to the m-th sample image in the c-th category. Usually, L = 20 is set.
[0145] The server performs in-class data augmentation on the sample images of each category. Each time, two sample images of the same category are randomly selected and The label of the generated new sample image remains
[0146] Next, a cross-category sampling strategy is adopted to generate new category sample images by combining sample images between different categories. Denote the new category sample image as It is expressed by the formula as:
[0147]
[0148] where, m c is the number of sample images in the c-th category, M is the total number of all sample images, X m is the feature matrix of the m-th sample image, τ is the global correlation coefficient, A c is the covariance matrix of the c-th category, ‖·‖ F denotes taking the Frobenius norm of the matrix.
[0149] The generated new category sample images will be sent to the classifier for a new round of training.
[0150] Step 105, construct an edge generative adversarial network containing a discriminator and a generator. Based on the sample images, through adversarial training, make the sample distribution of the sample images generated by the generator close to the real sample distribution.
[0151] In a specific implementation, the server needs to construct an edge generative adversarial network containing a discriminator and a generator. Based on the sample images, through adversarial training, make the sample distribution of the sample images generated by the generator close to the real sample distribution.
[0152] In one example, the server constructs an edge generative adversarial network composed of a discriminator D and a generator G, converts the discriminator D into a classifier with N + 1 classes, achieving the purpose of shrinking the classification decision boundary. The N + 1th class is used to determine whether the sample image is real data or fake data from the generator G. c The server needs to calculate the probability score that the sample image output by the classifier belongs to class j through the following formula: c where c(x) represents the corresponding value of the ith node of the classifier.
[0153] Next, train the discriminator D. The loss function used to train the discriminator D is represented by the following formula:
[0154]
[0155] where D is the loss function used to train the discriminator D, is the loss function corresponding to the classifier correctly classifying the known class, and E{log<P[y = N + 1|G(z)]>} is the loss function for the classifier to correctly distinguish whether the sample image comes from real data or the generator G. i (x) represents the corresponding value of the ith node of the classifier.
[0156] Then, train the generator G. Fix the parameters of the discriminator D and only update the parameters of the generator G. Let the last hidden layer of the classification network of the discriminator D be c’(x). The loss function used to train the generator G is represented by the following formula:
[0157]
[0158] where G is the loss function used to train the generator G, represents the expected value of the real sample image x from D, E represents the expected value of the noise z sampled from the prior distribution P(z), and G(z) represents that the generator G receives the noise z and generates a fake sample image. Loss is the loss function used to train the discriminator D, is the loss function corresponding to the classifier correctly classifying the known class, and E{log<P[y = N + 1|G(z)]>} is the loss function for the classifier to correctly distinguish whether the sample image comes from real data or the generator G. z~P(z) {log<P class [y = N c + 1|G(z)]>} is the loss function for the classifier to correctly distinguish whether the sample image comes from real data or the generator G.
[0159] After that, train the generator G. Fix the parameters of the discriminator D and only update the parameters of the generator G. Let the last hidden layer of the classification network of the discriminator D be c’(x). The loss function used to train the generator G is represented by the following formula:
[0160]
[0161] where G is the loss function used to train the generator G, Loss is the loss function used to train the generator G, represents the expected value of the real sample image x from D, train E represents the expected value of the noise z sampled from the prior distribution P(z), and G(z) represents that the generator G receives the noise z and generates a fake sample image. z~P(z) E represents the expected value of the noise z sampled from the prior distribution P(z), and G(z) represents that the generator G receives the noise z and generates a fake sample image.
[0162] Thus, the server performs adversarial training, alternately updating the parameters of the discriminator D and the generator G until the sample distribution of the sample images generated by the generator G is close enough to the real distribution, and at the same time constructs the adversarial loss term L Nargin , L Margin = DLoss +G Loss 。
[0163] Step 106: Construct and train a region of interest (ROI) extraction network, traverse the entire sample image through a sliding window, and obtain the top-k candidate regions that are most likely to contain the target according to the top-k strategy.
[0164] In a specific implementation, the server also needs to construct and train an ROI extraction network, traverse the entire sample image through a sliding window, and obtain the top-k candidate regions that are most likely to contain the target according to the top-k strategy.
[0165] In one example, the server generates a total of 9 anchor boxes with 3 scales and 3 aspect ratios at each position of the feature map of the sample image to cover different scales and shapes of the target as much as possible. Subsequently, each anchor box is passed through a 3×3 convolutional kernel and two 1×1 convolutional kernels in sequence. The first 1×1 convolutional kernel outputs the binary classification score of the anchor box, that is, the probability value that the anchor box may contain the target. The second 1×1 convolutional kernel outputs the offset of the adjusted position and size of the anchor box, obtaining a number of candidate regions. Next, non-maximum suppression is performed on each candidate region. For each anchor box, the candidate region with the highest score is selected, and other candidate regions with too high overlap with it are removed until all candidate regions are judged. Finally, all the remaining candidate regions are sorted from largest to smallest according to the score. According to the top-k strategy, the top-k remaining candidate regions with the highest scores are selected as the candidate regions that are most likely to contain the target.
[0166] Step 107: Map the top-k candidate regions that are most likely to contain the target from the original space to a higher-dimensional feature space, and use a non-linear support vector machine to realize the recognition and classification of the open-set target in the sample image.
[0167] In a specific implementation, the server maps the top-k candidate regions that are most likely to contain the target from the original space to a higher-dimensional feature space, and uses a non-linear support vector machine to realize the recognition and classification of the open-set target in the sample image.
[0168] In one example, the server maps the feature map corresponding to the top-k candidate regions that are most likely to contain the target from the original space to a higher-dimensional feature space. The model expression of the hyperplane in the high-dimensional space is:
[0169] g(x) = ω T φ(x) + b
[0170] where ω T and b are the feature vectors of the hyperplane, and φ(x) is the representation of the sample image x after feature mapping;
[0171] Next, the Gaussian kernel function is used to calculate the similarity between sample images in the high-dimensional space, and the sample image x i and the sample image x j The similarity between them is expressed by the formula as:
[0172]
[0173] where σ is the bandwidth parameter of the Gaussian kernel, and κ(x i , x j ) is the similarity between the sample image x i and the sample image x j .
[0174] Subsequently, a non-linear SVM classifier is constructed and trained. The SVM classifier is expressed by the formula as:
[0175]
[0176] s.t. y i [ω T φ(x) + b] ≥ 1, i = 1, 2,..., m
[0177] Its dual problem is expressed as:
[0178]
[0179] It can be calculated through the Gaussian kernel function:
[0180] κ(x i , x j ) = <φ(x i ), φ(x j )> = φ(x i ) T φ(x j )
[0181] This non-linear classification problem can be expressed as solving:
[0182]
[0183] where α i is the Lagrange multiplier, and CR is the regularization parameter.
[0184] Finally, the server adjusts the parameters of the mapping function, the kernel function, and the regularization parameter of the SVM according to the results of the performance evaluation to optimize the performance of the classifier. The loss function of the entire network is:
[0185] L all = L CR + εL CORAL +(1 - ε)L Margin
[0186] Among them, ε is a trade-off parameter used to balance the influence of the boundary sample image and the sample image of the new category on the classifier.
[0187] It should be noted that as a trade-off parameter, ε uses Mixup again. ε can balance the influence of the sample boundary and the new sample on the classifier. The larger ε is, the more the classifier focuses on the boundary of the known category, that is, it pays more attention to the correct classification of the known category. When ε is smaller, it means that the classifier pays more attention to the recognition of unknown categories and is applicable to scenarios where the detection requirements for abnormal unknown categories are relatively strict.
[0188] After completing the training of the entire network, the nonlinear SVM classifier can be applied to infer and test the image to be classified. In this step, the server uses the trained SVM model to process the image to be classified, correctly classify the known categories, and identify the unknown categories.
[0189] In one example, the structure of the entire network used in an open-set image classification method based on feature enhancement and boundary optimization proposed in this application can be as Figure 2 shown.
[0190] An open-set image classification method based on feature enhancement and boundary optimization proposed in this embodiment uses the weighted Fourier feature transformation of the MMD distance to realize the expansion of domain-invariant features, effectively extracts the inherent features inside and between samples, reduces the within-class variance, increases the between-class variance, and enhances the model's ability to recognize and utilize invariant features. The improved weighted mixup method is used to quickly and highly quality expand the sample images of known categories, improve the diversity of sample images, simulate the sample images of unknown categories, and enhance the model's performance in a variable category environment. It requires less computing resources and has strong practicability. Through the candidate box extraction network and the nonlinear support vector machine, the hyperplane between the known category and the unknown category is accurately divided, improving the real-time performance and robustness of the classifier, and realizing the accurate division of multiple unknown categories. Finally, it effectively reduces the cost required for manually labeled samples and well realizes the effective recognition and classification of unlabeled categories in real-world scenarios.
[0191] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of this application.
[0192] In one embodiment, the effect of an open-set image classification method based on feature enhancement and boundary optimization proposed in this application (hereinafter referred to as OURS) can be further illustrated by the following simulation experiments.
[0193] 1. Simulation conditions.
[0194] The simulation experiment adopted in this embodiment was carried out using Python software on an operating system with an Intel(R) COWOD(TM) i7 3.60GHZ central processing unit, 1G of memory, and WINDOWS10.
[0195] 2. Simulation content.
[0196] Simulation 1: Use OURS and the OWOD model to identify MS-COCO images. The results are as Figure 3 shown. It can be seen from Figure 3 that OURS has a good recognition effect on MS-COCO images and identifies potential unknown classes.
[0197] Simulation 2: Use OURS and the OWOD model to identify KITTI images. The results are as Figure 4 shown. It can be seen from Figure 4 that OURS can also achieve good recognition results in multi-object and complex scenarios.
[0198] The inference time and number of iterations used in Simulation 1 and Simulation 2 are shown in Table 1.
[0199] Table 1: Inference time and number of iterations used in Simulation 1 and Simulation 2
[0200]
[0201] It can be seen from Table 1 that compared with the OWOD model, OURS can obtain recognition results with only a few iterations. The time used for image recognition is much less than that of the OWOD model, and the target recognition efficiency of the image has been greatly improved.
[0202] Another embodiment of this application proposes an open-set image classification system based on feature enhancement and boundary optimization. The implementation details of an open-set image classification system based on feature enhancement and boundary optimization proposed in this embodiment will be specifically described below. The following content is only the implementation details provided for easy understanding and is not necessary for implementing this example.
[0203] The specific structure of an open-set image classification system based on feature enhancement and boundary optimization proposed in this embodiment can be as Figure 5As shown in the figure, the system includes: a feature extraction module 201, a Fourier phase transformation module 202, a cross-domain covariance calculation module 203, a weighted mixup transformation module 204, an edge generative adversarial training module 205, a candidate box extraction module 206, and a support vector machine classification module 207.
[0204] The feature extraction module 201 is used to extract features from the sample images of known categories using a pre-trained ResNet50 network to obtain the basic features of the sample images of known categories.
[0205] The Fourier phase transformation module 202 is used to perform a Fourier transform on the basic features of the sample images of known categories, calculate the Fourier phase information of the transformed basic features, and use it as the within-domain invariant features in the category domain of the sample images of known categories.
[0206] The cross-domain covariance calculation module 203 is used to calculate the covariance matrix of the within-domain invariant features of different category domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different category domains;
[0207] The weighted mixup transformation module 204 is used to perform data augmentation on the sample images of known categories through a weighted mixup fusion method, and generate new sample images and their category labels by linearly combining different sample images of known categories and their category labels.
[0208] The edge generative adversarial training module 205 is used to construct an edge generative adversarial network including a discriminator and a generator, and based on the sample images, make the sample distribution of the sample images generated by the generator approach the real sample distribution through adversarial training.
[0209] The candidate box extraction module 206 is used to construct and train a region of interest extraction network, traverse the entire sample image through a sliding window, and obtain the top-k most likely target-containing candidate regions according to the top-k strategy.
[0210] The support vector machine classification module 207 is used to map the top-k most likely target-containing candidate regions from the original space to a higher-dimensional feature space, and with the help of a non-linear support vector machine, realize the recognition and classification of the open-set targets in the sample images.
[0211] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or can be implemented as a combination of multiple physical units. In addition, to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0212] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.
[0213] Another embodiment of the present application proposes an electronic device, the specific structure of which is as Figure 6 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein, the memory 302 stores instructions executable by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute a method for open-set image classification based on feature enhancement and boundary optimization described in each of the above method embodiments.
[0214] Among them, the memory and the processor can be connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be further described herein. The bus interface is responsible for providing an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0215] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0216] Another embodiment of the present application proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement a method for open-set image classification based on feature enhancement and boundary optimization described in each of the above method embodiments.
[0217] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disks, or optical discs.
[0218] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.
Claims
1. An open set image classification method based on feature enhancement and boundary optimization, characterized in that: include: Use the pre-trained ResNet50 network to extract features of sample images of known categories and obtain the basic features of sample images of known categories; Performing Fourier transformation on basic features of sample images of known categories, calculating Fourier phase information of the transformed basic features as domain-internal invariant features of the category domain of the sample images of known categories; Calculate the covariance matrix of the domain-invariant features of different category domains, and calculate the CORAL distance based on the covariance matrix to measure the differences between different category domains; The weighted mixup fusion method is used to enhance the data of sample images of known categories, and sample images of new categories and their category labels are generated by linearly combining sample images of different known categories and their category labels. Construct an edge generative adversarial network consisting of a discriminator and a generator. Based on sample images, the sample distribution of the sample images generated by the generator is made close to the real sample distribution through adversarial training. Build and train a region of interest extraction network, traverse the entire sample image through a sliding window, and obtain the top k candidate regions most likely to contain the target according to the top-k strategy; The first k candidate regions that are most likely to contain the target are mapped from the original space to a higher-dimensional feature space, and the recognition and classification of open-set targets in the sample image are achieved with the help of nonlinear support vector machines.
2. The open set image classification method based on feature enhancement and boundary optimization according to claim 1, characterized in that: Use the pre-trained ResNet50 network to extract features of sample images of known categories and obtain the basic features of sample images of known categories, including: Prepare a set of sample images of known categories. Each sample image of a known category is marked with a category label. The total number of categories in the sample image set of known categories is N. c ; Load the pre-trained ResNet50 network and fine-tune the last few layers of the ResNet50 network to adapt to the sample image set of known categories; The sample images of known categories are input into the fine-tuned ResNet50 network for feature extraction, and the feature vectors output by the fine-tuned ResNet50 network are obtained as the basic features of the sample images of known categories; wherein the basic features of the sample images of known categories represent the representation of the sample images of known categories in the feature space.
3. The open set image classification method based on feature enhancement and boundary optimization according to claim 2, characterized in that: Perform Fourier transform on the basic features of the sample image of the known category, and calculate the Fourier phase information of the transformed basic features as the domain-internal invariant features of the category domain of the sample image of the known category, including: The basic features of the sample images of known categories are transformed by Fourier transform to obtain the frequency domain representation of the basic features of the sample images of known categories through the following formula: Among them, H and W are the height and width of the basic features of the sample image of known categories, respectively, and F (u,v) (x) is the frequency domain representation of the basic features of the sample image of a known category; The Fourier phase information of the corresponding channel is calculated based on the frequency domain representation of the basic features of the sample image of a known category using the following formula: x * =P (u,v) (x)=arctan[I (u,v) (x) / R (u,v) (x)] Among them, I (u,v) (x) is F (u,v) The imaginary part of (x), R (u,v) (x) is F (u,v) The real part of (x), x * represents the calculated Fourier phase information; use Train the feature extractor, Representing the category label of a sample image of a known category, the FFT loss used when training the feature extractor is expressed by the following formula: Among them, θ f is the feature extractor G f The learnable parameters, E (x,y) Indicates the expectation, L cls is the cross entropy loss, is the frequency domain representation of the basic features of the sample image of the mth known category, is the category label of the sample image of the mth known category, d mmd (S i ,S j ) is the MMD distance between the current category and other categories, i≠j, i represents the current category, j represents other categories, L fft The FFT loss used when training the feature extractor; Based on the trained feature extractor, the domain-invariant features of each category domain are obtained.
4. The open set image classification method based on feature enhancement and boundary optimization according to claim 3, characterized in that: The covariance matrix of the domain-invariant features of different category domains is calculated, and the CORAL distance is calculated based on the covariance matrix to measure the differences between different category domains, including: Calculate the mean μ of the domain-invariant features of all category domains c , μ c It is expressed by the formula: Among them, f(x m ) is the domain-internal invariant feature of the category domain extracted from the sample image of the mth known category; The covariance matrix of the domain-invariant features of different category domains is calculated by the following formula: Among them, C c represents the covariance matrix of the c-th category domain; The CORAL distance between two different category domains is calculated by the following formula, and finally the CORAL matrix D between all category domains is obtained. CORAL ; in, C c The square root of the inverse moment, C d Denotes the covariance matrix of the d-th category domain, D CORAL (c, d) represents the CORAL distance between the c-th category domain and the d-th category domain; Based on D CORAL Construct the regularization loss, which is expressed by the formula: L CORAL =L cls +λD CORAL Among them, L cls is the classification loss, and λ is a hyperparameter that weighs the CORAL distance regularization term.
5. The open set image classification method based on feature enhancement and boundary optimization according to claim 4, characterized in that: The weighted mixup fusion method is used to enhance the data of sample images of known categories. Sample images of new categories and their category labels are generated by linearly combining sample images of different known categories and their category labels, including: Assume that the unknown new category is C+1. For each known category, maintain a fixed-length queue q cm To store the features corresponding to each known category of sample images, Represents the Lth feature corresponding to the mth sample image in the cth class; Perform data enhancement within each category of sample images, and randomly select two sample images of the same category each time and The label of the generated new sample image is still The cross-category sampling strategy is adopted to generate sample images of new categories by combining sample images of different categories. The sample images of new categories are recorded as It is expressed by the formula: Among them, m c is the number of sample images of the cth category, M is the total number of all sample images, X m is the feature matrix of the mth sample image, τ is the global correlation coefficient, A c is the covariance matrix of the cth category, ‖·‖ F It means to find the Frobenius norm of a matrix.
6. The open set image classification method based on feature enhancement and boundary optimization according to claim 5, characterized in that: Construct an edge generative adversarial network consisting of a discriminator and a generator. Based on sample images, adversarial training is performed to make the sample distribution of the sample images generated by the generator close to the real sample distribution, including: Construct an edge generation adversarial network consisting of a discriminator D and a generator G, and convert the discriminator D into an N c +1 classifier, to achieve the purpose of shrinking the classification decision boundary, the Nth c +1 class is used to determine whether the sample image is real data or fake data from the generator G; The probability score of the sample image output by the classifier belonging to category j is calculated by the following formula: Among them, c i (x) represents the corresponding value of the i-th node of the classifier; Train the discriminator D. The loss function used in training the discriminator D is expressed by the following formula: Among them, D Loss is the loss function used to train the discriminator D, is the corresponding loss function for the classifier to correctly classify known categories, E z~P(z) {log <P class [y=N c +1|G(z)]>} is the loss function for the classifier to correctly identify the sample image from the real data or the generator G; Train the generator G, fix the parameters of the discriminator D, and only update the parameters of the generator G. Let the last hidden layer of the classification network of the discriminator D be c'(x). The loss function used to train the generator G is expressed by the following formula: Among them, G Loss is the loss function used to train the generator G, Indicates that from D train The true value of the sample image x is E z~P(z) Represents the expected value of the noise z sampled from the prior distribution P(z), G(z) represents the generator G receiving the noise z and generating a fake sample image; Perform adversarial training, alternately update the parameters of the discriminator D and the generator G, until the sample distribution of the sample images generated by the generator G is close enough to the true distribution, and construct the adversarial loss term L Margin , L Margin =D Loss +G Loss .
7. The open set image classification method based on feature enhancement and boundary optimization according to claim 6, characterized in that: Build and train the region of interest extraction network, traverse the entire sample image through a sliding window, and obtain the top k candidate regions that are most likely to contain the target according to the top-k strategy, including: Generate 9 anchor boxes of 3 scales and 3 aspect ratios at each position of the feature map of the sample image to cover different scales and shapes of the target as much as possible; Each anchor box is passed through a 3×3 convolution kernel and two 1×1 convolution kernels in sequence. The first 1×1 convolution kernel outputs the binary classification score of the anchor box, that is, the probability value that the anchor box may contain the target. The second 1×1 convolution kernel outputs the offset of the adjusted anchor box position and size to obtain several candidate regions. Perform non-maximum suppression on each candidate region, obtain the candidate region with the highest score for each anchor frame, and remove other candidate regions that overlap too much with it until all candidate regions are judged; All retained candidate regions are arranged from large to small according to their scores, and according to the top-k strategy, the first k retained candidate regions with the highest scores are selected as the candidate regions most likely to contain the target.
8. The open set image classification method based on feature enhancement and boundary optimization according to claim 7, characterized in that: The method maps the first k candidate regions most likely to contain the target from the original space to a higher-dimensional feature space, and uses a nonlinear support vector machine to realize the recognition and classification of open-set targets in the sample image, including: The feature maps corresponding to the top k candidate regions that are most likely to contain the target are mapped from the original space to a higher-dimensional feature space. The model of the high-dimensional space hyperplane is expressed as: g(x)=ω T φ(x)+b Among them, ω T , b is the feature vector of the hyperplane, φ(x) is the representation of the sample image x after feature mapping; Use the Gaussian kernel function to calculate the similarity between sample images in high-dimensional space. i With the sample image x j The similarity between them is expressed by the formula: Among them, σ is the bandwidth parameter of the Gaussian kernel, κ(x i ,x j ) is the sample image x i With the sample image x j The similarity between Construct and train a nonlinear SVM classifier. The SVM classifier is expressed by the formula: s.t.y i [ω T φ(x)+b]≥1,i=1,2,...,m Its dual problem is expressed as: The Gaussian kernel function can be calculated: k(x i ,x j )=<φ(x i ),φ(x j )>=φ(x i ) T φ(x j ) This nonlinear classification problem can be expressed as solving: Among them, α i is the Lagrange multiplier, CR is the regularization parameter; According to the results of performance evaluation, the parameters of the mapping function, kernel function and SVM regularization parameters are adjusted to optimize the performance of the classifier. The loss function of the entire network is: L all =L CR +εL CORAL +(1-e)L Margin Among them, ε is a trade-off parameter, which is used to weigh the impact of boundary sample images and sample images of new categories on the classifier.
9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an open set image classification method based on feature enhancement and boundary optimization as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement an open set image classification method based on feature enhancement and boundary optimization as described in any one of claims 1 to 7.