An open set image recognition method based on contrast and ensemble discrimination

By adopting a method based on comparison ideas and integrated discrimination in image open-set recognition technology, using neural network models and integrated unknown discriminators, the problem that known class samples and unknown class samples are easily misclassified in mixed scenarios is solved, and higher recognition accuracy and efficiency are achieved.

CN115331055BActive Publication Date: 2025-05-23TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210973602.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-05-23
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

Existing image open-set recognition technology is prone to problems similar to known class samples and unknown classes in mixed scenarios, resulting in the known class samples and unknown class samples being misclassified, affecting the accuracy of identification.

Method used

The image open-set recognition method based on contrast ideas and integrated discrimination is adopted, and the neural network model structure and integrated unknown discriminator are used, including unknown detectors based on reconstruction error distribution and feature distribution, to reduce error classification and improve recognition accuracy.

Benefits of technology

It effectively reduces the misclassification between known class samples and unknown class samples in mixed scenarios, improves the accuracy of image open set recognition, and improves the recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115331055B_ABST
    Figure CN115331055B_ABST
Patent Text Reader

Abstract

The present invention provides an image open set recognition method based on contrast thinking and integrated discrimination, including reading an image data set, establishing a neural network model structure, training the neural network model structure, establishing an integrated unknown discriminator, and performing image open set recognition. The present invention discloses an image open set recognition method based on contrast thinking and integrated discrimination, which can effectively reduce the misclassification between known class samples and unknown class samples in mixed scenarios and improve the accuracy of image open set recognition by establishing a neural network model structure, using contrast learning training thinking to perform model training, and finally using an integrated unknown discriminator to perform image open set recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer vision, and in particular relates to an image open set recognition method based on contrast ideas and integrated discrimination. Background Art

[0002] Computer vision is a field of artificial intelligence that enables computers and systems to obtain meaningful information from images, videos or other visual inputs and take actions or provide recommendations based on that information. If artificial intelligence gives computers the ability to think, then computer vision gives them the ability to discover, observe and understand.

[0003] In computer vision, image recognition is a basic and core branch of computer vision, which can accurately identify whether a given image belongs to a certain category. With the rise of the field of artificial intelligence, image recognition technology has been widely used in security inspection, identity verification, commodity circulation and other fields, so image recognition has great application value.

[0004] With the rapid improvement of computer performance and the rapid development of deep learning technology, a large number of excellent deep learning-based models have emerged in the field of image recognition. Convolutional neural networks have powerful information extraction capabilities. Deep learning represented by convolutional neural networks far exceeds traditional image recognition algorithms in terms of recognition accuracy and stability. Although research in the field of image recognition has been widely used, in order to be applied to reality and face the open world, image recognition models must also have the ability to handle unknown samples, that is, samples that have never appeared in the training model process should not be misclassified as a known class, but should be judged as unknown categories. Therefore, a branch of image recognition has emerged in the field of open set recognition that can identify unknown class samples while identifying each known class. In the open set recognition model, some models use generative adversarial networks to construct pseudo-unknown class samples for training, so that the model has the ability to deal with unknown class samples. The effect of the open set recognition model based on generated samples is very dependent on the quality of the generated samples. If the quality of the generated samples is not good, the open set recognition model will not work well.

[0005] Although many models with excellent performance have emerged in the field of open-set image recognition, the existing open-set image recognition technology still has the following problems: for example, when the same identity is in different walking states, these samples need to be mixed into one category for recognition, but in this case, the samples of the same identity will be quite different, and there will be problems that the known class samples are similar to the unknown class, resulting in the known class samples and the unknown class samples being misclassified. The misclassification between the known class and the unknown class will bring errors to the judgment of information, and this error will seriously affect the actual application effect of open-set image recognition. Summary of the invention

[0006] In view of this, the present invention aims to propose an image open set recognition method based on contrast idea and integrated discrimination, which, in conjunction with the use of a neural network model structure and an integrated unknown discriminator, can effectively reduce the misclassification between known class samples and unknown class samples in mixed scenarios and improve the accuracy of image open set recognition.

[0007] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0008] Step 1: Read the image data set: The computer reads the image data set and divides the image data set into a training data set and a test data set;

[0009] Step 2: Establish the neural network model structure: The neural network structure includes an encoder, a decoder, and a classifier. The encoder and decoder are used to encode and decode the input image. The classifier is a fully connected network that uses the features extracted by the encoder to classify known classes.

[0010] Step 3: Training the neural network model structure: Input the labeled training data set in step 1 into the neural network model structure established in step 2 for model training, and record the features of the known class image samples and the reconstruction error;

[0011] Step 4: Establish an integrated unknown discriminator: The integrated unknown discriminator includes an unknown detector based on the reconstruction error distribution and an unknown detector based on the feature distribution; the classifier prediction result in step 2 and the reconstruction error in step 3 are respectively input into the unknown detector based on the reconstruction error distribution, and the features extracted by the encoder in step 2 are input into the unknown detector based on the feature distribution. The unknown detector based on the reconstruction error determines whether the reconstruction error belongs to the category under the guidance of the classifier prediction result. At the same time, the feature-based unknown detector is compared with the feature distribution of all categories to determine whether it belongs to a known category. If any unknown detector judges the sample to be identified as an unknown sample through the quantile threshold, the sample is finally regarded as an unknown category, otherwise it is a known category, and the classifier prediction result is used as the classification result of the sample to be identified;

[0012] Step 5: Perform open set image recognition: Input the image to be recognized into the neural network model structure established in step 2, and the integrated unknown discriminator established in step 4 outputs that the image to be recognized belongs to a known category or an unknown category.

[0013] Compared with the prior art, the image open set recognition method based on contrast and integrated discrimination disclosed in the present invention has the following advantages:

[0014] First, the present invention discloses a neural network model structure of an open set image recognition method based on contrast thinking and integrated discrimination, which includes an encoder, a decoder and a classifier, and uses contrast learning thinking to train the neural network model structure, which can efficiently achieve mutual separation between samples.

[0015] Second, the present invention discloses an integrated unknown discriminator for an image open set recognition method based on the contrast idea and integrated discrimination, including an unknown detector based on a reconstructed error distribution and an unknown detector based on a feature distribution. Through the integrated discrimination of the two unknown detectors, the misclassification between known class samples and unknown class samples can be effectively reduced, thereby improving the accuracy of image open set recognition.

[0016] Third, the present invention discloses an image open set recognition method based on contrast thinking and integrated discrimination, which realizes an end-to-end image open set recognition process through the connection between the neural network model structure and the integrated unknown discriminator, and has the characteristics of high recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0018] In the attached picture:

[0019] Figure 1 A schematic diagram of a neural network model structure of an open set image recognition method based on contrast and integrated discrimination according to an embodiment of the present invention;

[0020] Figure 2 A schematic diagram of an image open set recognition method based on contrast and integrated discrimination and integrating an unknown discriminator according to an embodiment of the present invention. DETAILED DESCRIPTION

[0021] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0022] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first", "second", and the like may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.

[0023] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood by specific circumstances.

[0024] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0025] like Figure 1-2 As shown, an image open set recognition method based on contrast thinking and integrated discrimination includes:

[0026] Step 1: Read the image data set: The computer reads the image data set and divides the image data set into a training data set and a test data set;

[0027] In this embodiment, the image dataset contains a total of 20,494 images of nine test persons in eight gait scenarios, and the ratio of the training dataset to the test dataset is 4:1.

[0028] Step 2: Establish the neural network model structure: Figure 1 As shown in the figure, the neural network model structure includes an encoder, a decoder and a classifier. The encoder and the decoder are used to encode and decode the input image. The classifier is a fully connected network that uses the features extracted by the encoder to classify known classes.

[0029] Step 3: Training the neural network model structure: Input the labeled training data set in step 1 into the neural network model structure established in step 2 for model training, and record the features of the known class image samples and the reconstruction error;

[0030] Furthermore, in step three, the convolutional neural network is trained using the contrastive idea, that is, the training data set and the corresponding labels are sent to the neural network model structure for training. The contrastive learning idea is used to make the samples of the same class in the training data set closer to each other, and the samples of different classes are separated from each other. After sufficient training, the corresponding network parameters are saved, and the reconstruction error and characteristics of the known class samples are recorded.

[0031] Step 4: Establish an integrated unknown discriminator: Figure 2 As shown, the integrated unknown discriminator includes an unknown detector based on the reconstruction error distribution and an unknown detector based on the feature distribution; the classifier prediction result in step 2 and the reconstruction error in step 3 are respectively input into the unknown detector based on the reconstruction error distribution, the feature extracted by the encoder in step 2 is input into the unknown detector based on the feature distribution, the unknown detector based on the reconstruction error determines whether the reconstruction error belongs to the category under the guidance of the classifier prediction result, and at the same time, the feature-based unknown detector is compared with the feature distribution of all categories to determine whether it belongs to a known category, if any unknown detector determines the sample to be identified as an unknown sample through the quantile threshold, then the sample is finally regarded as an unknown category, otherwise it is a known category, and the classifier prediction result is used as the classification result of the sample to be identified;

[0032] Further, in step 4, the recorded reconstruction error and features of the known class samples are used as the prior knowledge of the known class, and these prior knowledge are used to construct an integrated unknown discriminator. The use of the integrated unknown discriminator is to cope with the challenge of misclassification of unknown classes as known classes in mixed scenarios. The main reason why the unknown class samples are misclassified as known classes is that the discrimination strength of the unknown class samples is not enough, and the reason for the insufficient discrimination strength is that the single use of a certain prior knowledge is not enough to fully reject the unknown class samples. For example, the reconstruction error can only capture the global difference of the image, but cannot accurately capture the local difference of the sample. In this case, we believe that it is impossible to achieve the ideal effect by relying solely on an unknown discriminator based on reconstruction error or feature to detect unknown samples. Based on the above analysis, the present invention creatively uses the idea of ​​integrated discrimination, that is, these two prior knowledge can be used at the same time to enhance the model's discrimination ability for unknown class samples. According to the concept of image processing, the reconstruction error of the known class samples can be considered as pixel-level prior knowledge, which shows the difference of the image at the pixel level, and the features of the sample can be considered as semantic-level prior knowledge, which reflects the difference of different images at the semantic level.

[0033] Step 5: Perform open set image recognition: Input the image to be recognized into the neural network model structure established in step 2, and the integrated unknown discriminator established in step 4 outputs that the image to be recognized belongs to a known category or an unknown category.

[0034] In step 3, the loss function L used to train the neural network model structure is 总 as follows:

[0035] L 总 =αL cls +βL kls +γL sc +δL rec +ξL klu

[0036] Among them, α, β, γ, δ, ξ are all weight parameters;

[0037] In this embodiment, α is set to 100, β and δ are set to change linearly in the range of 0-1 with the training rounds; γ and ξ are set to be a constant of 1.

[0038] In this embodiment,

[0039] L cls is the classification loss, and its specific formula is as follows:

[0040]

[0041] Where N is the number of samples, i represents the serial number corresponding to the sample, is the total number of known classes, k is the number of classes; I[·] is the indicator function, if the sample x i Label y i If it is equal to k, its value is 1, otherwise it is 0; i,k Represents sample x i The probability of belonging to k;

[0042] L kls It is the regular term loss of automatic encoding with category information, and its specific formula is as follows:

[0043]

[0044] Among them, L is the number of layers of the encoder, N is the number of samples, and i represents the sequence number corresponding to the sample. is the total number of known classes, k is the number of classes, Represents two distributions as well as The KL divergence between and is a sample x belonging to the kth class i The output of the encoder, is the vector obtained by calculating the label through the fully connected layer, I represents the standard deviation vector, and N(.) represents the normal distribution;

[0045] L sc It is the supervised contrast loss, and its specific formula is as follows:

[0046]

[0047] in, Represents sample x i Features of z i The normalized representation of

[0048] S represents the normalized representation set of all samples;

[0049] Indicates division by x i In addition, with the input i-th sample x i The set of normalized representations of all samples of the same category;

[0050] s q Representative Set Normalized representation of samples in;

[0051] Indicates that except for the input i-th sample x i The set of normalized representations of all samples except

[0052] s a Representative Set Normalized representation of samples in;

[0053] τ is a scale parameter;

[0054] L rec is the reconstruction loss, and the specific formula is as follows:

[0055]

[0056] Where N is the number of samples, i represents the serial number corresponding to the sample, is the image sample x i The reconstructed image of

[0057] L klu is the multi-scale feature alignment loss function, and its specific formula is as follows:

[0058]

[0059] Where L is the number of layers of the encoder, l represents the lth layer of the encoder, is the l-th layer output of the decoder, is the output of the encoder at layer l, Represents two distributions as well as The KI divergence between and is a sample x belonging to the kth class i The output of the encoder, is the vector obtained by calculating the label through the fully connected layer, I represents the standard deviation vector, and N(.) represents the normal distribution;

[0060] The data set used in the present invention is a gait identification data set in a mixed scenario. The method disclosed in the present invention is compared and verified with similar algorithms in multiple prior arts. As shown in the following table, the openness setting represents the openness of the open set identification problem. The larger the openness value, the larger the ratio of the number of unknown classes to the number of known classes in the test set, and the greater the difficulty of open set identification. The data in the table are weighted harmonic means of precision and recall. The higher the value, the better the performance of open set identification and the better the model. It can be seen that this method has made significant progress over the prior art.

[0061] Comparison results on gait identification dataset in mixed scenarios

[0062]

[0063] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An open set image recognition method based on contrast and integrated discrimination. Features: include: Step 1: Read the image data set: The computer reads the image data set and divides the image data set into a training data set and a test data set; Step 2: Establish the neural network model structure: The neural network model structure includes an encoder, a decoder, and a classifier. The encoder and decoder are used to encode and decode the input image. The classifier is a fully connected network that uses the features extracted by the encoder to classify known classes. Step 3: Training the neural network model structure: Input the labeled training data set in step 1 into the neural network model structure established in step 2 for model training, and record the features of the known class image samples and the reconstruction error; Step 4: Establish an integrated unknown discriminator: The integrated unknown discriminator includes an unknown detector based on the reconstruction error distribution and an unknown detector based on the feature distribution; the classifier prediction result in step 2 and the reconstruction error in step 3 are respectively input into the unknown detector based on the reconstruction error distribution, and the features extracted by the encoder in step 2 are input into the unknown detector based on the feature distribution. The unknown detector based on the reconstruction error distribution determines whether the reconstruction error belongs to the category under the guidance of the classifier prediction result. At the same time, the unknown detector based on the feature distribution is compared with the feature distribution of all categories to determine whether it belongs to a known category. If any unknown detector judges the sample to be identified as an unknown sample through the quantile threshold, the sample is finally regarded as an unknown category, otherwise it is a known category, and the classifier prediction result is used as the classification result of the sample to be identified; Step 5: Perform open set image recognition: Input the image to be recognized into the neural network model structure established in step 2, and the integrated unknown discriminator established in step 4 outputs that the image to be recognized belongs to a known category or an unknown category.

2. According to claim 1, an image open set recognition method based on contrast and integrated discrimination, Features: In step 3, the loss function L used to train the neural network model structure is 总 as follows: L 总 =αL cls +βL kls +γL sc +δL rec +ξL klu Among them, α, β, γ, δ, ξ are all weight parameters; Where N is the number of samples, i represents the serial number corresponding to the sample, is the total number of known classes, k is the number of classes; I[·] is the indicator function, if the sample x i Label y i If it is equal to k, its value is 1, otherwise it is 0; i,k Represents sample x i The probability of belonging to k; Among them, L is the number of layers of the encoder, N is the number of samples, and i represents the sequence number corresponding to the sample. is the total number of known classes, k is the number of classes, Represents two distributions as well as The KI divergence between and is a sample x belonging to the kth class i The output of the encoder, is the vector obtained by calculating the label through the fully connected layer, I represents the standard deviation vector, and N(.) represents the normal distribution; in, Represents sample x i Features of z i The normalized representation of S represents the normalized representation set of all samples; Indicates division by x i In addition, with the input i-th sample x i The set of normalized representations of all samples of the same category; s q Representative Set Normalized representation of samples in; Indicates that except for the input i-th sample x i The set of normalized representations of all samples except s a Representatives Normalized representation of samples in; τ is a scale parameter; Where N is the number of samples, i represents the serial number corresponding to the sample, is the image sample x i The reconstructed image of Where L is the number of layers of the encoder, l represents the lth layer of the encoder, is the l-th layer output of the decoder, is the output of the encoder at layer l, Represents two distributions as well as The KL divergence between , N(.) represents the normal distribution, and is a sample x belonging to the kth class i The output of the encoder, is the label vector obtained by calculating the fully connected layer, I represents the standard deviation vector, and N(·) represents the normal distribution.

3. The image open set recognition method based on contrast and integrated discrimination according to claim 1, Features: The ratio of the training data set to the test data set is 4:1.

Citation Information

Patent Citations

  • Open set identification method and device and computer readable storage medium

    CN109784325A

  • Radar interference semi-supervised open set identification system based on generative adversarial network

    CN114241263A