Model training method and device and object recognition method and device
By performing two forward propagation and weight updates on the object recognition model, the problem that the existing model cannot recognize objects of unknown categories is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510165461.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The existing object recognition model only has the function of enclosed set recognition, and cannot effectively identify objects of unknown categories, resulting in recognition errors.
By region division and feature extraction of sample images, two forward propagation is performed using the visual converter model, combining neutral signal vectors and category text features, the total loss is calculated and back-gradient propagation is performed, and the model weight is updated to improve the open set recognition ability.
The recognition performance of the object recognition model for unknown categories and fuzzy categories is improved, the robustness of the model in out-of-distribution scenarios is enhanced, and the probability of identification errors is reduced.
Smart Images

Figure CN120107984A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training and object recognition method and device. Background Art
[0002] The object recognition model trained by traditional methods usually only has closed set recognition function, that is, it can only identify the category of the object in the image to be recognized as the known category set in the training stage. However, in actual application scenarios, the image to be recognized often contains objects of unknown categories. In this case, the object recognition model will forcibly classify the objects of unknown categories into known categories, resulting in object recognition errors.
[0003] The object recognition model with open set recognition function can distinguish objects of known categories from objects of unknown categories, reducing the probability of object recognition errors. In view of this, how to make the object recognition model have open set recognition function has become an important research topic.
[0004] In the related art, the output results of the trained object recognition model are generally processed based on methods such as Maximum Softmax Probability (MSP) to enable the object recognition model to have the function of open set recognition. Specifically, the maximum soft classification probability method directly uses the maximum probability value of the object output by the object recognition model that belongs to a known category to determine whether the object belongs to a known type or an unknown category. Although the above method is relatively simple to implement, it has a low sensitivity to unknown categories, resulting in a low accuracy of object recognition using the object recognition model. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a model training, object recognition method and device to improve the accuracy of object recognition. The specific technical solution is as follows:
[0006] In a first aspect, an embodiment of the present application provides a model training method, the method comprising:
[0007] Divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features;
[0008] Inputting the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtaining a first object category output by the visual converter model;
[0009] Obtaining a category text feature of the first object category, inputting the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtaining a second object category output by the visual converter model;
[0010] Determine a first loss generated by the visual converter model predicting the first object category based on the first object category and the true category label of the object, and determine a second loss generated by the visual converter model predicting the second object category based on the second object category and the true category label;
[0011] Based on the first loss and the second loss, a total loss generated by the visual converter model for category prediction is determined, and reverse gradient propagation is performed based on the total loss to update the weight coefficient of the visual converter model to obtain an object recognition model.
[0012] In a second aspect, an embodiment of the present application provides an object recognition method, the method comprising:
[0013] Obtaining an image to be recognized;
[0014] The image to be identified is input into an object recognition model to obtain the object category of the object in the image to be identified output by the object recognition model, wherein the object recognition model is a model trained according to the method described in the first aspect.
[0015] In a third aspect, an embodiment of the present application provides a model training device, the device comprising:
[0016] The feature acquisition module is used to divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features;
[0017] A first prediction module, configured to input the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtain a first object category output by the visual converter model;
[0018] A second prediction module obtains a category text feature of the first object category, inputs the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtains a second object category output by the visual converter model;
[0019] a loss determination module, configured to determine a first loss generated by the visual converter model predicting the first object category based on the first object category and the true category label of the object, and to determine a second loss generated by the visual converter model predicting the second object category based on the second object category and the true category label;
[0020] A weight updating module is used to determine the total loss generated by the visual converter model for category prediction based on the first loss and the second loss, perform reverse gradient propagation based on the total loss, update the weight coefficient of the visual converter model, and obtain an object recognition model.
[0021] In a fourth aspect, an embodiment of the present application provides an object recognition device, the device comprising:
[0022] An image acquisition module, used for acquiring an image to be recognized;
[0023] An object recognition module is used to input the image to be recognized into an object recognition model to obtain the object category of the object in the image to be recognized output by the object recognition model, wherein the object recognition model is: a model trained according to the method described in the first aspect.
[0024] In a fifth aspect, an embodiment of the present application provides an electronic device, including:
[0025] Memory, used to store computer programs;
[0026] The processor is used to implement the method described in the first aspect or the second aspect when executing the program stored in the memory.
[0027] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect or the second aspect is implemented.
[0028] In a seventh aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the method described in the first aspect or the second aspect.
[0029] As can be seen from the above, when the scheme provided by the embodiment of the present application is used for model training, the neutral signal vector and the image block features are input into the visual converter model to obtain the first object category output by the visual converter model; then, the category text features and the image block features of the first object category are input into the visual converter model to obtain the second object category output by the visual converter model; then, based on the first object category and the true category label of the object, the first loss generated by the visual converter model is determined, and based on the second object category and the true category label, the second loss generated by the visual converter model is determined; finally, based on the first loss and the second loss, the total loss is determined, and reverse gradient propagation is performed based on the total loss to update the weight coefficient of the visual converter model to obtain the object recognition model. It can be seen that the model training scheme provided by the embodiment of the present application involves two forward propagation processes: in the first forward propagation process, the neutral signal vector and the image block features are used to obtain the initial predicted category; in the second forward propagation process, the category prediction is performed again based on the category text features of the initial prediction result, the confidence of the first prediction is reduced, a new prediction result is obtained, and the ability to refine the initial predicted category is improved. Moreover, during the second forward propagation process, the category text features can reflect the semantic information of the initial predicted category, so that the model can learn refined category information, improve the recognition performance of the trained object recognition model for fuzzy categories and unknown categories, and enhance the robustness of the object recognition model in out-of-distribution scenarios.
[0030] Of course, implementing any product or method of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0032] Figure 1 A schematic diagram of the flow chart of the first model training method provided in the embodiment of the present application;
[0033] Figure 2 A schematic diagram of a flow chart of a second model training method provided in an embodiment of the present application;
[0034] Figure 3 A schematic diagram of a model training process provided in an embodiment of the present application;
[0035] Figure 4 A flowchart of an object recognition method provided in an embodiment of the present application;
[0036] Figure 5 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;
[0037] Figure 6 A schematic diagram of the structure of an object recognition device provided in an embodiment of the present application;
[0038] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field based on the present application belong to the scope of protection of the present application.
[0040] First, the execution subject of the solution provided in the embodiment of the present application is described.
[0041] The executor of the solution provided in the embodiment of the present application is: any electronic device with data processing, storage and other functions.
[0042] The application scenarios of the solutions provided in the embodiments of the present application are introduced.
[0043] The application scenario of the solution provided in the embodiment of the present application is: training the neural network model to be trained to obtain a trained object recognition model.
[0044] The object recognition model trained by traditional methods usually only has closed set recognition function, and can only identify the category of the object in the image to be recognized as the known category set in the training stage, that is, it can only recognize the in-distribution (ID) data. However, in actual application scenarios, the image to be recognized often contains objects of unknown categories. In this case, the object recognition model will forcibly classify the objects of unknown categories into known categories, resulting in object recognition errors, that is, the model cannot recognize out-of-distribution (OOD) data.
[0045] An object recognition model with an open set recognition function can recognize OOD data. In the prior art, the output results of the trained object recognition model are generally processed by methods such as maximum soft classification probability to enable the object recognition model to have the function of open set recognition.
[0046] However, the above methods have poor application effects in the Vision Transformer (ViT) model and the Iterative Neural Network (INN). It can be seen that the current research limitations on new models such as ViT and INN limit the practicality and scalability of existing technologies in open set recognition scenarios.
[0047] In view of this, an embodiment of the present application provides a model training method to train an object recognition model that can recognize OOD data.
[0048] The model training scheme provided in the embodiment of the present application is introduced in detail below in conjunction with the accompanying drawings.
[0049] See also Figure 1 , is a flow chart of the first model training method provided in an embodiment of the present application, and the above method includes the following steps S101-S105.
[0050] Step S101: Divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features.
[0051] In this step, the sample image may be divided into a plurality of non-overlapping image blocks, each of which may be referred to as a Patch.
[0052] The embodiment of the present application does not introduce the specific method of dividing multiple non-overlapping image blocks, but is explained below by examples.
[0053] In one implementation, the sample image may be divided into a plurality of non-overlapping image blocks of the same size.
[0054] For sample images in, represents a real number set, H and W represent the height and width of the sample image respectively, and C represents the number of channels of the sample image. If the height and width of the image block are set to P, then x can be divided into There are non-overlapping image patches of size P×P.
[0055] For example, if the H and W of the sample image x are both 224 pixels, and the height and width of the image block are set to 16 pixels, then x is divided into non-overlapping image patches.
[0056] After obtaining each image block, feature extraction can be performed on each divided image block to obtain image block features.
[0057] In one implementation, each image patch may be projected into a d-dimensional vector space to obtain an image patch feature (Patchtoken).
[0058] The above d dimension refers to the embedding dimension, that is, the feature dimension required by the model, which can be determined according to actual conditions.
[0059] Specifically, the image block may be flattened into a block vector first, and then the flattened block vector may be projected to the embedding dimension d through a dimension conversion matrix. Of course, if the dimension of the flattened block vector is d, no projection operation is required.
[0060] In another implementation, feature extraction algorithms such as Histogram of Oriented Gradients (HOG) and Scale-Invariant Feature Transform (SIFT) may be used to extract features of the image block to obtain image block features, which will not be described in detail here.
[0061] Step S102: inputting the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtaining a first object category output by the visual converter model.
[0062] The above-mentioned neutral signal vector refers to a vector that is neither biased towards a specific category nor exhibits obvious characteristics.
[0063] In one case, the neutral signal vector is: a zero vector of the second dimension, and the second dimension is: a feature dimension of the obtained image block feature, that is, the aforementioned embedding dimension.
[0064] The zero vector does not show obvious characteristics and is easy to generate, and can reduce the difficulty of subsequent calculations. Therefore, using the zero vector of the second dimension as a neutral signal vector can improve the efficiency of model training.
[0065] In this step, a pre-generated neutral signal vector may be obtained; or a neutral signal vector may be generated before this step is performed.
[0066] When the neutral signal vector is a zero vector, the neutral signal vector may be generated in the following manner.
[0067] The product of the zero vector and the second-dimensional conversion matrix is calculated to obtain the second-dimensional zero vector as the neutral signal vector. The second-dimensional conversion matrix can be the same as the first-dimensional conversion matrix.
[0068] In this way, based on the first dimensional conversion matrix, the dimensional conversion can be realized quickly and efficiently to obtain a neutral signal vector.
[0069] Specifically, the neutral signal vector and the obtained image block features can be input into the encoder in the visual converter model. The encoder adopts a multi-head self-attention mechanism to capture the relationship between the obtained image block features. Multiple self-attention heads calculate the attention weights in parallel, and then output the attention-weighted features; then, the features output by the encoder are linearly transformed through a feedforward neural network (FFN), and finally the features corresponding to the neutral signal vector in the change result are input into the classification head to obtain the classification result of the classification head, and the first object category is output based on the classification result of the classified meat.
[0070] The classification result of the classification head may be a classification vector, and each element in the vector represents the confidence that the object category is a set category. In this way, the object category with the highest confidence may be determined as the first object category.
[0071] Step S103: obtaining a category text feature of the first object category, inputting the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtaining a second object category output by the visual converter model.
[0072] The first object category is the object category predicted by the visual transformer model, and the first category text of the first object category refers to the text of the object category. For example, if the first object category is a sedan, the first category text is the text of "sedan".
[0073] The above-mentioned category text features are features related to the first category text, and can be obtained in the following way:
[0074] In one implementation, the category description text of the first object category can be obtained based on a preset large language model (LLM), and then the category text feature can be obtained based on the category description text. Figure 2 The illustrated embodiment will not be described in detail here.
[0075] In another implementation, text features of the first category text of the first object category may be extracted as category text features of the first object category.
[0076] In this way, the category text feature can be conveniently obtained based on the text feature of the first category text of the first object category, thereby improving the efficiency of obtaining the category text feature.
[0077] In one case, the first category text may be input into a Contrastive Language-Image Pre-training (CLIP) model to obtain the category text features output by the CLIP model.
[0078] In some cases, features of the first category text may also be extracted based on a word embedding algorithm or a bidirectional encoder representations from transformers (BERT) model as category text features, which is not limited in the embodiments of the present application.
[0079] Step S104: Based on the first object category and the true category label of the object, determine a first loss generated by the visual converter model predicting the first object category, and based on the second object category and the true category label, determine a second loss generated by the visual converter model predicting the second object category.
[0080] Specifically, a loss calculation function such as a cross entropy loss function and a mean square error loss function may be used to calculate the first loss and the second loss.
[0081] Step S105: Based on the first loss and the second loss, determine the total loss generated by the visual converter model for category prediction, perform reverse gradient propagation based on the total loss, update the weight coefficient of the visual converter model, and obtain the object recognition model.
[0082] Specifically, the weight coefficients of the visual converter model can be updated based on the gradient of the total loss for the weight coefficients, and then a gradient descent algorithm such as a stochastic gradient descent (SGD) algorithm and a stochastic gradient descent with momentum (SGD with Momentum) algorithm can be used.
[0083] Among them, each time the weight coefficient of the visual converter model is adjusted according to the two object prediction results of the visual converter model for the sample image, it can be considered that a round of model training is completed.
[0084] In one embodiment of the present application, multiple rounds of training may be performed until a model that meets the requirements is obtained. Specifically, after the weight coefficient is updated, the step S101 may be returned to execute until the set training end condition is met, and the model after the weight coefficient is updated is determined as the object recognition model.
[0085] The above-mentioned training end condition can be model convergence, or the number of training times reaches a set number, etc., which is not limited in the embodiments of the present application.
[0086] As can be seen from the above, when the scheme provided by the embodiment of the present application is used for model training, the neutral signal vector and the image block features are input into the visual converter model to obtain the first object category output by the visual converter model; then, the category text features and the image block features of the first object category are input into the visual converter model to obtain the second object category output by the visual converter model; then, based on the first object category and the true category label of the object, the first loss generated by the visual converter model is determined, and based on the second object category and the true category label, the second loss generated by the visual converter model is determined; finally, based on the first loss and the second loss, the total loss is determined, and based on the total loss, reverse gradient propagation is performed to update the weight coefficient of the visual converter model to obtain an object recognition model.
[0087] It can be seen that the model training scheme provided by the embodiment of the present application involves two forward propagation processes: in the first forward propagation process, the neutral signal vector and the image block features are used to obtain the initial prediction category; in the second forward propagation process, the category prediction is performed again based on the category text features of the initial prediction result to reduce the confidence of the first prediction, obtain a new prediction result, and improve the ability to refine the initial prediction category. In addition, in the second forward propagation process, the category text features can reflect the semantic information of the initial prediction category, so that the model can learn the refined category information, improve the recognition performance of the trained object recognition model for fuzzy categories and unknown categories, and enhance the robustness of the object recognition model in out-of-distribution scenarios.
[0088] In one embodiment of the present application, the following steps A-B may be used to obtain category text features.
[0089] Step A: Obtain initial text features of the first object category.
[0090] In this step, the method of obtaining the initial text features is the same as the method of obtaining the category text features introduced in the aforementioned step S103, the only difference being the feature names, which will not be described in detail here.
[0091] Step B: If the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature, the initial text feature is converted into a text feature of the second dimension as a category text feature of the first object category.
[0092] The second dimension mentioned above is also the embedded dimension of the setting mentioned above.
[0093] In one implementation, the product of the initial text feature and the first dimensional conversion matrix can be calculated to obtain the text feature of the second dimension. In this way, the dimensional conversion of the initial text feature can be quickly and efficiently achieved based on the first dimensional conversion matrix.
[0094] In some possible implementations, the initial text features may also be converted into text features of the second dimension based on a principal component analysis (PCA) algorithm, an autoencoder, etc., which will not be described in detail here.
[0095] It can be seen that when the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature, the dimension of the converted initial text feature can be unified with the dimension of the image block feature through dimensionality conversion, thereby improving the execution efficiency of the subsequent model training process.
[0096] exist Figure 1 Based on the embodiment shown, when obtaining the category text feature of the first object category, the category description text of the first object category can be obtained first, and then the category text feature is obtained based on the category description text. In view of the above situation, the embodiment of the present application provides a second model training method.
[0097] See also Figure 2 , is a flow chart of the second model training method provided in an embodiment of the present application, and the above method includes the following steps S201-S207.
[0098] Step S201: Divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features.
[0099] Step S202: inputting the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtaining a first object category output by the visual converter model.
[0100] The above steps S201-S202 are the same as the above steps Figure 1 In the illustrated embodiment, steps S101 and S102 are the same and will not be described in detail here.
[0101] Step S203: inputting the first category text of the first object category and the first instruction text into a preset large language model to perform a detailed description on the first category text, and obtaining a category description text output by the large language model.
[0102] The first instruction text is a text that instructs the large language model to perform a detailed description of the first category of text.
[0103] The embodiment of the present application does not limit the specific content of the first instruction text, which is introduced below through examples.
[0104] For example, the first instruction text may be "detailed description of the following category text:" etc.
[0105] In this way, the large language model can determine, based on the first instruction text, that the task goal is to provide a detailed description of the first category text, thereby being able to perform corresponding text amplification, feature description and other operations on the first category text to obtain a category description text.
[0106] Step S204: extracting text features of the category description text as category text features of the first object category.
[0107] The method of extracting text features of category description text can be found in the above Figure 1 The method of extracting text features of the first category of text introduced in step S103 in the illustrated embodiment will not be repeated here.
[0108] Step S205: inputting the category text features and the obtained image block features into the visual converter model to predict the category of the object, and obtaining a second object category output by the visual converter model.
[0109] Step S206: Based on the first object category and the true category label of the object, determine a first loss generated by the visual converter model predicting the first object category, and based on the second object category and the true category label, determine a second loss generated by the visual converter model predicting the second object category.
[0110] Step S207: Based on the first loss and the second loss, determine the total loss generated by the visual converter model for category prediction, perform reverse gradient propagation based on the total loss, update the weight coefficient of the visual converter model, and obtain the object recognition model.
[0111] The above steps S206-S207 are the same as the above steps Figure 1 In the illustrated embodiment, steps S104 and S105 are the same and will not be described in detail here.
[0112] In this embodiment, when generating the category text feature of the first object category, the first category text is firstly refined by means of the large language model to obtain the category description text, and then the text features of the category description text are extracted as the category text feature of the first object category. In this way, the category text feature containing more fine-grained semantic information can be obtained based on the refined category description text, so that the subsequent model can learn more comprehensive and detailed category information, further improving the recognition performance of the trained object recognition model for fuzzy categories and unknown categories.
[0113] Based on the above embodiments, Figure 3 , a specific model training process provided in an embodiment of the present application is described.
[0114] See also Figure 3, is a schematic diagram of a model training process provided in an embodiment of the present application. It should be noted that the two visual converters in the figure are the same visual converter. The above model training process can be divided into the following five stages.
[0115] Phase 1: Patch segmentation and token representation
[0116] In this stage, the sample image x is divided into There are P×P non-overlapping image patches, and each patch is projected into a d-dimensional vector space to obtain z 1 ,z 2 ,…,z N There are N image patch vectors (Patch token) in total, which can also be called image patch features.
[0117] Phase 2: First forward propagation
[0118] In this process, the class token in the traditional training process is replaced by the context (CTX) token.
[0119] In this stage, the CTX token is first initialized as a neutral signal: Then, through a linear layer Transformed into the input space, we get That is to say It can be called the neutral signal vector.
[0120] Among them, W lin Represents a learnable weight matrix, which is used to transform The dimension of is converted to the embedding dimension (such as 768 dimensions).
[0121] Next, Input the visual converter model with the image patch vector to obtain the first predicted category output by the visual converter model That is,
[0122] Among them, W cls represents the learnable classification head weight matrix, the Softmax() function is used to normalize the probability distribution, and ViT() represents the internal processing flow of the visual transformer model.
[0123] So far, the first predicted category is obtained The first forward propagation is completed.
[0124] Phase 3: Second forward propagation
[0125] This stage can be divided into the following steps 1-3:
[0126] Step 1: Generate descriptive embeddings
[0127] Based on the first predicted category Generate using large language models The category descriptive text S of the category text is input into the CLIP model to generate an embedding vector including refined semantic features: CLIP =CLIP-TextEncoder(S), where CLIP-TextEncoder() represents the internal processing flow of the CLIP model.
[0128] Step 2: Conditionalize CTX token
[0129] Specifically, this step can be CLIP The dimension of is projected onto the embedding dimension, and we get It can also be called contextual features.
[0130] Right now, Where W lin Used to CLIP The dimension of is converted to the embedding dimension.
[0131] Step 3: Feature Input
[0132] After the conditional The image block features are input into the visual converter model so that the visual converter model performs a refined prediction and outputs the second prediction result, i.e., the second prediction category in,
[0133] So far, the second prediction category is obtained The second forward propagation ends.
[0134] Stage 4: Determine the total loss function
[0135] In this stage, the total loss is calculated based on the first prediction category and the second prediction category. Specifically, the sum of the loss functions of the two forward propagations can be used as the optimization target, that is, the total loss function L=L1+L2.
[0136] in,
[0137] Among them, CrossEntropy() represents the cross entropy loss function, y true Represents the true category of the object in the sample image.
[0138] Stage 5: Backward Gradient Propagation
[0139] In this stage, reverse gradient propagation is performed based on the total loss. Specifically, the gradient of the total loss function with respect to the weight coefficient of the visual converter model can be calculated, and then the gradient descent algorithm is used to update the weight coefficient.
[0140] At this point, the weight coefficient of the visual converter model is adjusted according to the two object prediction results of the visual converter model for the sample image, completing a round of model training.
[0141] The above training process is repeated based on new sample images until the set number of times is reached or the model converges, and an object recognition model can be obtained.
[0142] In the solution of the embodiment of the present application, the model training mechanism using two forward propagations can be called a Zigzag mechanism, and the open set recognition framework for training visual converters based on the Zigzag mechanism can be called Zigzag-ViT.
[0143] It can be seen from the above that in the above training process, for samples within the distribution, the model improves the prediction accuracy by minimizing the loss; for samples outside the distribution, the model enhances the robustness to OOD data by reducing the confidence of the first prediction.
[0144] The following is an introduction to the performance of the object recognition model trained using the solution provided in the embodiments of the present application.
[0145] First, refer to the following Table 1, which shows the performance of the object recognition model trained by the solution provided in the embodiment of the present application under different ID data sets and OOD data sets.
[0146] Table 1
[0147]
[0148] In the above Table 1, AUROC represents the area under the receiver operating characteristic curve, and the larger the value, the better the performance of the model; FPR95 represents the false positive rate at 95% recall, and the smaller the value, the better the performance of the model; IDAcc. represents the in-class accuracy, and the larger the value, the better the performance of the model; the above datasets 1 to 6 are SVHN dataset, iSUN dataset, Places365 dataset, Textures dataset, LSUN-R dataset and LSUN-C dataset respectively. Among them, the three parameters separated by " / " in the 3rd to 5th columns represent: the parameters of the model (baseline model) obtained by traditional training, the parameters of the model trained by one forward propagation (replacing the class token with a neutral signal vector only), and the parameters of the model trained by double forward propagation.
[0149] It can be seen from Table 1 that Zigzag-ViT significantly improves the model's ability to detect OOD samples in the open set recognition task.
[0150] Referring again to Table 2 below, the performance of the object recognition model trained using the solution provided in the embodiment of the present application is shown in combination with different post-processing methods.
[0151] Table 2
[0152]
[0153] The meanings of the same indicators or parameters in Table 2 refer to Table 1. Methods 1 to 6 represent the MSP method, the Outlier Detection using Incremental Networks (ODIN) method, the Reliable and Adaptive Classification for Open Set Recognition (ReAct) method, the Adversarial Self-Hardening for Open Set Recognition-Softmax (ASH-S) method, the SCALE method, and the Energy method, respectively. The parameter values corresponding to the methods in rows 7 to 10 are calculated based on the energy scores.
[0154] As can be seen from Table 2, compared with the existing mainstream OOD detection post-processing methods (such as MSP, Energy, ODIN, ReAct, ASH-S, SCALE, etc.), Zigzag-ViT can continuously improve the performance when combining these methods, especially showing the greatest improvement in the case of double forward propagation.
[0155] For example, in terms of the AUROC indicator, Zigzag-ViT has a significant improvement over the baseline model obtained by traditional training, and at the same time has a significant reduction in the FPR95 indicator, and effectively improves the overall performance of the model while maintaining the same ID Acc.
[0156] Finally, refer to Table 3 below, which shows the performance of the object recognition model trained using the solution provided in the embodiments of the present application under different training methods.
[0157] Table 3
[0158]
[0159] The meanings of the same indicators or parameters in Table 3 refer to Table 1, where the first column represents the training methods based on different classifiers, and Classifier 1 to Classifier 4 are Classifier Ensemble (CE) classifier, Regularized Mixture (RegMixup) classifier, Statistical Uncertainty Reduction for Ensemble (SURE) classifier and Incremental Semi-supervised Hierarchical (ISH) classifier respectively; Method 7 represents the Deep Ensemble method, and the parameter values corresponding to the methods in the last row are calculated based on the energy score.
[0160] As can be seen from Table 3, when Zigzag-ViT is combined with advanced classifiers (such as RegMixup, SURE, etc.), its performance can be further improved, making it more competitive in OOD detection.
[0161] At the same time, Zigzag-ViT shows good generalization ability on different prompt generation strategies, CLIP encoders and model sizes, demonstrating the effectiveness and versatility of its approach.
[0162] In summary, the Zigzag-ViT framework provided in the embodiments of the present application effectively improves the reliability and performance of the visual transformer in the open set recognition task through innovative architecture design and prompt engineering methods, and provides more reliable technical support for practical applications in the field of computer vision.
[0163] Corresponding to the above-mentioned model training method, an embodiment of the present application provides an object recognition method.
[0164] See also Figure 4 , is a flow chart of an object recognition method provided in an embodiment of the present application, the method includes the following steps S401-S402.
[0165] Step S401: Obtain an image to be recognized.
[0166] The image to be recognized may be an image including any object to be recognized.
[0167] Step S402: inputting the image to be identified into the object recognition model, and obtaining the object category of the object in the image to be identified output by the object recognition model.
[0168] Among them, the object recognition model is: a model trained according to the aforementioned model training method.
[0169] Since the aforementioned model training method enables the model to learn detailed category information, improves the recognition performance of the trained object recognition model for fuzzy categories and unknown categories, and enhances the robustness of the object recognition model in out-of-distribution scenarios, therefore, in this embodiment, identifying objects in images based on the object recognition model can reduce the probability of object recognition errors and improve the accuracy of object recognition.
[0170] Among them, the process of object recognition model identifying objects in the application stage is similar to the aforementioned model training process, which is briefly introduced in two stages below.
[0171] Phase 1: Patch segmentation and token representation
[0172] The image to be recognized is divided into non-overlapping patches of the same size, and each patch is projected into a d-dimensional vector space to obtain N image block vectors.
[0173] Phase 2: First forward propagation
[0174] Initialize the CTX token to a neutral signal and will Input linear layer W lin ,get Will Input the visual converter model with the image patch vector to obtain the first predicted category output by the visual converter model
[0175] Phase 3: Second forward propagation
[0176] The CLIP embedding library may pre-store descriptive embeddings corresponding to each prediction category, so that the corresponding descriptive embedding may be retrieved from the CLIP embedding library according to the first prediction category. CLIP .
[0177] Then, e CLIP The dimension of is projected onto the embedding dimension, and we get Right now
[0178] Finally, the conditioned The image block features are input into the visual converter model so that the visual converter model performs a refined prediction and outputs a second prediction category as the final output result.
[0179] Corresponding to the above-mentioned model training method, an embodiment of the present application also provides a model training device.
[0180] See also Figure 5 , is a schematic diagram of the structure of a model training device provided in an embodiment of the present application, wherein the device comprises the following modules:
[0181] The feature acquisition module 501 is used to divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features;
[0182] A first prediction module 502, configured to input the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtain a first object category output by the visual converter model;
[0183] A second prediction module 503 obtains a category text feature of the first object category, inputs the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtains a second object category output by the visual converter model;
[0184] A loss determination module 504 is configured to determine a first loss generated by the visual converter model predicting the first object category based on the first object category and the true category label of the object, and to determine a second loss generated by the visual converter model predicting the second object category based on the second object category and the true category label;
[0185] The weight updating module 505 is used to determine the total loss generated by the visual converter model for category prediction based on the first loss and the second loss, perform reverse gradient propagation based on the total loss, update the weight coefficient of the visual converter model, and obtain an object recognition model.
[0186] As can be seen from the above, when the scheme provided by the embodiment of the present application is used for model training, the neutral signal vector and the image block features are input into the visual converter model to obtain the first object category output by the visual converter model; then, the category text features and the image block features of the first object category are input into the visual converter model to obtain the second object category output by the visual converter model; then, based on the first object category and the true category label of the object, the first loss generated by the visual converter model is determined, and based on the second object category and the true category label, the second loss generated by the visual converter model is determined; finally, based on the first loss and the second loss, the total loss is determined, and based on the total loss, reverse gradient propagation is performed to update the weight coefficient of the visual converter model to obtain an object recognition model.
[0187] It can be seen that the model training scheme provided by the embodiment of the present application involves two forward propagation processes: in the first forward propagation process, the neutral signal vector and the image block features are used to obtain the initial prediction category; in the second forward propagation process, the category prediction is performed again based on the category text features of the initial prediction result, the confidence of the first prediction is reduced, and a new prediction result is obtained, which improves the ability to refine the initial prediction category. In addition, in the second forward propagation process, the category text features can reflect the semantic information of the initial prediction category, so that the model can learn the refined category information, improve the recognition performance of the trained object recognition model for fuzzy categories and unknown categories, and enhance the robustness of the object recognition model in out-of-distribution scenarios.
[0188] In one embodiment of the present application, the second prediction module 503 is specifically used to input the first category text and the first instruction text of the first object category into a preset large language model to perform a detailed description of the first category text, and obtain a category description text output by the large language model, wherein the first instruction text is: text instructing the large language model to perform a detailed description of the first category text; extracting text features of the category description text as category text features of the first object category; inputting the category text features and the obtained image block features into the visual converter model to predict the category of the object, and obtain a second object category output by the visual converter model.
[0189] In this embodiment, when generating the category text feature of the first object category, the first category text is firstly refined by means of the large language model to obtain the category description text, and then the text features of the category description text are extracted as the category text feature of the first object category. In this way, the category text feature containing more fine-grained semantic information can be obtained based on the refined category description text, so that the subsequent model can learn more comprehensive and detailed category information, further improving the recognition performance of the trained object recognition model for fuzzy categories and unknown categories.
[0190] In one embodiment of the present application, the second prediction module 503 is specifically used to extract text features of the first category text of the first object category as category text features of the first object category; input the category text features and the obtained image block features into the visual converter model to predict the category of the object, and obtain the second object category output by the visual converter model.
[0191] In this way, the category text feature can be conveniently obtained based on the text feature of the first category text of the first object category, thereby improving the efficiency of obtaining the category text feature.
[0192] In one embodiment of the present application, the second prediction module 503 includes:
[0193] A first feature acquisition submodule, used to obtain initial text features of the first object category;
[0194] A second feature acquisition submodule is used for converting the initial text feature into a text feature of the second dimension as a category text feature of the first object category if the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature;
[0195] The second prediction submodule is used to input the category text features and the obtained image block features into the visual converter model to predict the category of the object, so as to obtain a second object category output by the visual converter model.
[0196] It can be seen that when the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature, the dimension of the converted initial text feature can be unified with the dimension of the image block feature through dimensionality conversion, thereby improving the execution efficiency of the subsequent model training process.
[0197] In one embodiment of the present application, the second feature acquisition submodule is specifically used to calculate the product of the initial text feature and the first dimensional conversion matrix to obtain the text feature of the second dimension if the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature.
[0198] In this way, the dimensionality conversion of the initial text features can be achieved quickly and efficiently based on the first dimensional conversion matrix.
[0199] In one embodiment of the present application, the neutral signal vector is: a zero vector of a second dimension, and the second dimension is: a feature dimension of the obtained image block feature.
[0200] The zero vector does not show obvious characteristics and is easy to generate, and can reduce the difficulty of subsequent calculations. It can be seen that using the zero vector of the second dimension as a neutral signal vector can improve the efficiency of model training.
[0201] In one embodiment of the present application, the device further includes:
[0202] The neutral signal vector acquisition module is used to calculate the product of the zero vector and the second-dimensional conversion matrix before the first prediction module 502 is triggered, so as to obtain the zero vector of the second dimension as the neutral signal vector.
[0203] In this way, based on the first dimensional conversion matrix, the dimensional conversion can be realized quickly and efficiently to obtain a neutral signal vector.
[0204] Corresponding to the above-mentioned object recognition method, an embodiment of the present application also provides an object recognition device.
[0205] See also Figure 6 , is a schematic diagram of the structure of an object recognition device provided in an embodiment of the present application, wherein the device comprises the following modules:
[0206] An image acquisition module 601 is used to obtain an image to be recognized;
[0207] The object recognition module 602 is used to input the image to be recognized into the object recognition model to obtain the object category of the object in the image to be recognized output by the object recognition model, wherein the object recognition model is: a model trained according to the aforementioned model training method.
[0208] Since the aforementioned model training method enables the model to learn detailed category information, improves the recognition performance of the trained object recognition model for fuzzy categories and unknown categories, and enhances the robustness of the object recognition model in out-of-distribution scenarios, therefore, in this embodiment, identifying objects in images based on the object recognition model can reduce the probability of object recognition errors and improve the accuracy of object recognition.
[0209] The present application also provides an electronic device, such as Figure 7 As shown, including:
[0210] Memory 701, used for storing computer programs;
[0211] The processor 702 is used to implement the aforementioned model training method or object recognition method when executing the program stored in the memory 701.
[0212] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 702, the communication interface, and the memory 701 communicate with each other via the communication bus.
[0213] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0214] The communication interface is used for communication between the above electronic device and other devices.
[0215] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0216] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0217] In another embodiment provided in the present application, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned model training or object recognition method are implemented.
[0218] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any model training or object recognition method in the above embodiments.
[0219] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a solid-state hard disk (SSD), etc.
[0220] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0221] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, electronic device and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0222] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features; Inputting the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtaining a first object category output by the visual converter model; Obtaining a category text feature of the first object category, inputting the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtaining a second object category output by the visual converter model; Determine a first loss generated by the visual converter model predicting the first object category based on the first object category and the true category label of the object, and determine a second loss generated by the visual converter model predicting the second object category based on the second object category and the true category label; Based on the first loss and the second loss, a total loss generated by the visual converter model for category prediction is determined, and reverse gradient propagation is performed based on the total loss to update the weight coefficient of the visual converter model to obtain an object recognition model.
2. The method according to claim 1, characterized in that The obtaining of the category text feature of the first object category includes: Inputting the first category text and the first instruction text of the first object category into a preset large language model to perform a detailed description on the first category text, thereby obtaining a category description text output by the large language model, wherein the first instruction text is: text instructing the large language model to perform a detailed description on the first category text; The text features of the category description text are extracted as category text features of the first object category.
3. The method according to claim 1, characterized in that The obtaining of the category text feature of the first object category includes: A text feature of a first category text of the first object category is extracted as a category text feature of the first object category.
4. The method according to claim 1, characterized in that The obtaining of the category text feature of the first object category includes: Obtaining initial text features of the first object category; If the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature, the initial text feature is converted into a text feature of the second dimension as a category text feature of the first object category.
5. The method according to claim 4, characterized in that The converting the initial text features into text features of the second dimension comprises: The product of the initial text feature and the first dimensional conversion matrix is calculated to obtain the text feature of the second dimension.
6. The method according to any one of claims 1 to 5, characterized in that The neutral signal vector is: a zero vector of the second dimension, and the second dimension is: a feature dimension of the obtained image block feature.
7. The method according to claim 6, characterized in that Before the neutral signal vector and the image block features of the sample image are input into the preset visual converter model, the method further includes: The product of the zero vector and the second-dimensional conversion matrix is calculated to obtain the zero vector of the second dimension as the neutral signal vector.
8. An object recognition method, characterized in that: The method comprises: Obtaining an image to be recognized; The image to be identified is input into an object recognition model to obtain an object category of the object in the image to be identified output by the object recognition model, wherein the object recognition model is a model trained according to any one of claims 1-7.
9. A model training device, characterized in that: The device comprises: The feature acquisition module is used to divide the sample image into regions, extract features from each image block obtained by the division, and obtain image block features; A first prediction module, configured to input the neutral signal vector and the obtained image block features into a preset visual converter model to predict the category of the object in the sample image, and obtain a first object category output by the visual converter model; A second prediction module obtains a category text feature of the first object category, inputs the category text feature and the obtained image block feature into the visual converter model to predict the category of the object, and obtains a second object category output by the visual converter model; a loss determination module, configured to determine a first loss generated by the visual converter model predicting the first object category based on the first object category and the true category label of the object, and to determine a second loss generated by the visual converter model predicting the second object category based on the second object category and the true category label; A weight updating module is used to determine the total loss generated by the visual converter model for category prediction based on the first loss and the second loss, perform reverse gradient propagation based on the total loss, update the weight coefficient of the visual converter model, and obtain an object recognition model.
10. The device according to claim 9, characterized in that The second prediction module is specifically used to input the first category text and the first instruction text of the first object category into a preset large language model to perform a detailed description of the first category text, and obtain the category description text output by the large language model, wherein the first instruction text is: text instructing the large language model to perform a detailed description of the first category text; extracting text features of the category description text as category text features of the first object category; inputting the category text features and the obtained image block features into the visual converter model to predict the category of the object, and obtain the second object category output by the visual converter model.
11. The device according to claim 9, characterized in that The second prediction module is specifically used to extract text features of the first category text of the first object category as category text features of the first object category; input the category text features and the obtained image block features into the visual converter model to predict the category of the object, and obtain a second object category output by the visual converter model.
12. The device according to claim 9, characterized in that The second prediction module comprises: A first feature acquisition submodule, used to obtain initial text features of the first object category; A second feature acquisition submodule is used for converting the initial text feature into a text feature of the second dimension as a category text feature of the first object category if the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature; The second prediction submodule is used to input the category text features and the obtained image block features into the visual converter model to predict the category of the object, so as to obtain a second object category output by the visual converter model.
13. The device according to claim 12, characterized in that The second feature acquisition submodule is specifically used to calculate the product of the initial text feature and the first dimensional conversion matrix to obtain the text feature of the second dimension if the first dimension of the initial text feature is not equal to the second dimension of the obtained image block feature.
14. The device according to any one of claims 9 to 13, characterized in that The neutral signal vector is: a zero vector of the second dimension, and the second dimension is: a feature dimension of the obtained image block feature.
15. The device according to claim 14, characterized in that The device also includes: The neutral signal vector acquisition module is used to calculate the product of the zero vector and the second-dimensional conversion matrix before the first prediction module is triggered, so as to obtain the zero vector of the second dimension as the neutral signal vector.
16. An object recognition device, characterized in that: The device comprises: An image acquisition module, used for acquiring an image to be recognized; An object recognition module is used to input the image to be recognized into an object recognition model to obtain the object category of the object in the image to be recognized output by the object recognition model, wherein the object recognition model is: a model trained according to any one of claims 1-7.
17. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1-7 or 8 when executing a program stored in a memory.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of any one of claims 1-7 or 8 are implemented.