Training Method, Device, Equipment and Storage Medium of Image Recognition Model
By using prototype samples in the training of image recognition model, screening and training neural network models, the problem of introducing a large number of extra parameters during the training process is solved, and the effect of saving storage and computing resources is achieved, reducing training difficulty, and improving model robustness is achieved.
Patent Information
- Application Number
- CN202111583676.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-22
AI Technical Summary
When training open sets to identify models, as the sample categories increase, the model training process will introduce a large number of additional parameters, resulting in an increase in storage space and operation time, as well as an increase in model training difficulty.
By training the image recognition model based on the prototype samples in the training samples, filter out the prototype samples of each image category, and train the neural network model using these prototype samples to avoid introducing additional prototype features and feature distances as training parameters.
This method saves storage space and computing time, reduces the difficulty of training the model, and improves the robustness and open set recognition performance of the image recognition model.
Smart Images

Figure CN114330522B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image recognition, and in particular, to a method, device, equipment and storage medium for training an image recognition model. Background Art
[0002] In the field of image recognition in computer vision, an open-set recognition model can process new sample categories that appear in a complex recognition system, and avoid misjudging unknown category samples as a certain known category. Therefore, the application scenarios are very extensive.
[0003] However, in the related art, when training an open-set recognition model, the prototype features of each category (i.e., the representative feature vectors representing the sample features of each category) are usually introduced, and the distance between the features of each sample and the prototype features is used as an additional parameter for training and learning. This results in a large number of additional parameters being introduced in the model training process as the number of sample categories increases, which not only increases the occupied storage space and computing time, but also increases the training difficulty of the model. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, equipment and storage medium for training an image recognition model, which is used to train an image recognition model based on prototype samples in training samples, so as to save storage space and computing time, and reduce the training difficulty of the model.
[0005] To achieve the above object, the embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, a method for training an image recognition model is provided, including: determining the uncertainty of each image sample in a plurality of image samples based on at least two image pre-recognition models; wherein the uncertainty is used to characterize the uncertainty of at least two image pre-recognition models in recognizing the image category; at least two image pre-recognition models are both used to determine the category of an image to be recognized from the same plurality of image categories; determining the first target image samples with uncertainties less than or equal to a first preset threshold in the plurality of image samples as the prototype samples of the image category to which the first target image samples belong, or determining the second target image samples with uncertainties less than a first preset number of other image samples in the plurality of image samples as the prototype samples of the image category to which the second target image samples belong; training a neural network model based on the prototype samples of each image category in the plurality of image categories to obtain an image recognition model; wherein the image recognition model is used to determine the category of an image to be recognized from the plurality of image categories, or determine that the category of the image to be recognized does not belong to the plurality of image categories.
[0007] In this technical solution, multiple image samples are screened based on the uncertainty of each image sample, and low-quality noise samples among the multiple image samples are filtered out to obtain prototype samples for each image category. Then, a neural network model is trained based on the prototype samples to obtain an image recognition model. By using the prototype samples for model training, it is not necessary to introduce the prototype features of each category, and it is not necessary to use the distance between the features of each sample and the prototype features as an additional parameter for training and learning. Therefore, when the number of sample categories increases, a large number of additional parameters can be avoided, saving storage space and computing time and reducing the training difficulty.
[0008] In a possible implementation, at least two image pre-recognition models include a first image pre-recognition model and a second image pre-recognition model, and the multiple image samples include first image samples. Determining the uncertainty of each image sample among the multiple image samples based on at least two image pre-recognition models includes: extracting a first feature vector of each image sample based on the first image pre-recognition model; wherein, the first feature vector is used for the first image pre-recognition model to determine the category of the image to be recognized from multiple image categories; extracting a second feature vector of each image sample based on the second image pre-recognition model; wherein, the second feature vector is used for the second image pre-recognition model to determine the category of the image to be recognized from multiple image categories; obtaining a first feature distribution vector based on the first feature vector of the first image sample and the first feature vectors of each image sample respectively; wherein, the first feature distribution vector is used to characterize the distance distribution between the first feature vector of the first image sample and the first feature vectors of other image samples; obtaining a second feature distribution vector based on the second feature vector of the first image sample and the second feature vectors of each image sample respectively; wherein, the second feature distribution vector is used to characterize the distance distribution between the second feature vector of the first image sample and the second feature vectors of other image samples; and obtaining the uncertainty of the first image sample based on the first feature distribution vector and the second feature distribution vector.
[0009] This possible implementation provides a specific implementation for determining the uncertainty of each image sample. If a computer device determines the uncertainty according to this method, it will help improve the accuracy of the uncertainty of each image sample, and thus help accurately screen out high-quality prototype samples.
[0010] In a possible implementation, training a neural network model based on the prototype samples of each image category among multiple image categories to obtain an image recognition model includes: determining a target prototype sample set for each image category, the multiple image categories include a first image category, and the target prototype sample set of the first image category includes at least a first prototype sample, and the first prototype sample is the one with the smallest uncertainty among the prototype samples of the first image category; and training a neural network model based on the target prototype sample set of each image category to obtain an image recognition model.
[0011] In this possible implementation, the neural network model is trained based on the prototype samples with the highest sample quality in each image category, reducing the amount of data for model training and the difficulty of model training, and shortening the time for model training.
[0012] In a possible implementation, the method further includes: determining the minimum feature distance between a second prototype sample in the first image category and the set of target prototype samples of the first image category, where the second prototype sample is a prototype sample that does not belong to the set of target prototype samples of the first image category; determining the second prototype sample with the largest minimum feature distance as an element belonging to the set of target prototype samples of the first image category; repeating the above steps until the largest minimum feature distance is less than or equal to a second preset threshold, or the number of elements in the set of target prototype samples of the first image category is equal to a second preset number.
[0013] This possible implementation filters out redundant and duplicate prototype samples among multiple prototype samples while retaining the diversity of prototype samples, improving the effectiveness of model training and the robustness of the image recognition model.
[0014] In a possible implementation, the multiple image categories include a second image category. Training a neural network model based on the prototype samples of each image category in the multiple image categories to obtain an image recognition model includes: training the neural network model based on the prototype samples of each image category and the loss function of the prototype samples of each image category to obtain an image recognition model; where the loss function of the prototype samples of the second image category is determined based on a first feature distance and a second feature distance, the first feature distance is the feature distance between each image sample in the second image category and the prototype sample of the second image category, and the second feature distance is the feature distance between each image sample in the second image category and the prototype samples of other image categories in the multiple image categories.
[0015] In this possible implementation, the loss function of the prototype samples of each image category constrains the training process of the neural network model, which helps the image recognition model better distinguish known category samples and unknown category samples.
[0016] Second aspect: A training device for an image recognition model is provided, including: functional units for performing the functions of any of the methods provided in the first aspect, and the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the training device for the image recognition model may include: an identification unit, a determination unit, and a training unit; the identification unit determines the uncertainty of each image sample in a plurality of image samples based on at least two image pre-recognition models; wherein the uncertainty is used to characterize the uncertainty of at least two image pre-recognition models in recognizing image categories; at least two image pre-recognition models are both used to determine the category of an image to be recognized from the same plurality of image categories; the determination unit is used to determine the first target image samples with uncertainties less than or equal to a first preset threshold in the plurality of image samples as prototype samples of the image category to which the first target image samples belong, or determine the second target image samples with uncertainties less than a first preset number of other image samples in the plurality of image samples as prototype samples of the image category to which the second target image samples belong; the training unit is used to train a neural network model based on the prototype samples of each image category in the plurality of image categories to obtain an image recognition model; wherein the image recognition model is used to determine the category of an image to be recognized from the plurality of image categories, or determine that the category of the image to be recognized does not belong to the plurality of image categories.
[0017] Third aspect: A computer device is provided, including: a processor and a memory. The processor is connected to the memory, and the memory is used to store computer execution instructions. The processor executes the computer execution instructions stored in the memory to implement any of the methods provided in the first aspect.
[0018] Fourth aspect: A chip is provided, including: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; the processor is used to run the code instructions to execute any of the methods provided in the first aspect.
[0019] Fifth aspect: A computer-readable storage medium is provided, including computer execution instructions. When the computer execution instructions run on a computer, the computer is caused to execute any of the methods provided in the first aspect.
[0020] Sixth aspect: A computer program product is provided, including computer execution instructions. When the computer execution instructions run on a computer, the computer is caused to execute any of the methods provided in the first aspect.
[0021] For the technical effects brought by any of the implementation manners in the second aspect to the sixth aspect, reference may be made to the technical effects brought by the corresponding implementation manners in the first aspect, which will not be elaborated here. Description of the Drawings
[0022] Figure 1A schematic diagram of a sample set provided by an embodiment of the present application;
[0023] Figure 2 A schematic diagram of the principle of open-set recognition provided by an embodiment of the present application;
[0024] Figure 3 A schematic diagram of the structure of a computer device provided by an embodiment of the present application;
[0025] Figure 4 A schematic diagram of the process of a method for training an image recognition model provided by an embodiment of the present application;
[0026] Figure 5 A schematic diagram of the process of another method for training an image recognition model provided by an embodiment of the present application;
[0027] Figure 6 A schematic diagram of the composition of a device for training an image recognition model provided by an embodiment of the present application. Detailed implementation manners
[0028] In the description of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more. The terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit to be different.
[0029] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0030] First, for the convenience of understanding the present application, relevant elements involved in the present application are described below.
[0031] Closed-set recognition (CSR): The categories in the training set and the test set are the same. The most common way is to use an open dataset for training, and the categories of all objects in the dataset are known, without objects of unknown categories.
[0032] Open-Set Recognition (OSR): The test set contains classes that are not in the training set. When using the test set for testing, if an image that does not belong to any of the known classes in the training set is input, the output is unknown, or the prediction result is output with a low confidence level.
[0033] Prototype: A representative sample that characterizes each sample class, or a representative feature (vector) that characterizes the features of each sample class. The prototype is equivalent to concepts such as common templates and support points.
[0034] Known-class samples: Samples that belong to the classes included in the training set, and can also be called closed-set samples.
[0035] Unknown-class samples: Samples that do not belong to the classes already included in the training set, and can also be called open-set samples.
[0036] Secondly, a brief introduction to the application scenarios involved in this application is given.
[0037] In the field of image recognition in computer vision, existing general image recognition problems can all be classified as "closed-set recognition" problems, that is, it is assumed that all sample classes in the test set are included in the training set (that is, there are no sample classes in the test set that are not in the training set). However, in some complex recognition applications, sample classes that have not appeared in the training set (i.e., samples of unknown classes / open-set samples) may appear during the testing process. At this time, if general image recognition techniques are still used, samples of unknown classes will inevitably be misjudged as samples of a certain known class, resulting in a decrease in the robustness of the recognition application.
[0038] The open-set recognition technology is proposed to address this problem. The open-set recognition technology can be regarded as a functional extension of general image recognition technology. That is, it not only has the ability to classify known-class samples (i.e., closed-set samples, samples of classes already included in the training set) of general image recognition technology, but also can detect samples of unknown classes, that is, distinguish between known-class samples and unknown-class samples.
[0039] As Figure 1 shown, only three known classes, namely, dog, human, and bird, are included in the training set, while two new classes, namely, car and cat, appear in the test set. The open-set recognition task has two goals: (1) correctly classify samples belonging to known classes ("dog, bird, human"); (2) distinguish whether the sample belongs to a known class or an unknown class ("cat, car"). It should be noted that the open-set recognition task only needs to distinguish between known classes and unknown classes, and does not need to further classify samples of unknown classes (that is, it does not need to classify samples of unknown classes as "cat" or "car").
[0040] AsFigure 2 As shown, the principle of the open-set recognition technology based on prototype features. In the field of open-set recognition technology, the method based on prototype features is a mainstream technology with both practical value and high performance. Specifically, first, the prototype features (representative feature vectors) used to characterize the features of each known-class sample are learned. Then, based on the prototype features, the features of the known-class samples are constrained, making the features of the known-class samples as close as possible to the prototype features of their own classes and as far as possible from the prototype features of other classes. Furthermore, the features of samples of the same class are compactly distributed around the prototype features, while the features of samples of different classes are as far apart as possible. In this way, when an unknown-class sample appears, there is a greater probability of falling into the area between the known-class samples. Then, according to the different feature distribution positions of the known-class samples and the unknown-class samples, the two are distinguished. For example, Figure 2 As shown, it can be seen that the feature distance of the sample points of the known class to the nearest prototype feature is small, while the feature distance of the sample points of the unknown class to all prototype features is large. Based on this distance, the known-class samples and the unknown-class samples can be distinguished.
[0041] The open-set recognition technology has great application potential and value in the field of image recognition because it can handle newly emerging sample classes in complex recognition systems and avoid misclassifying unknown-class samples as a certain known-class sample. Moreover, the application scenarios of the open-set recognition technology are very extensive and can be deployed in a variety of recognition systems, such as tasks like face recognition, pedestrian recognition, vehicle recognition, and so on.
[0042] However, in related technologies, when training an open-set recognition model, the prototype features of each class (i.e., the representative feature vectors representing the features of each class of samples) are usually introduced, and the distance between the features of each sample and the prototype features is used as an additional parameter for training and learning. This leads to a large number of additional parameters being introduced in the training process of the model as the number of sample classes increases, not only increasing the occupied storage space and computing time but also increasing the training difficulty of the model.
[0043] Next, a brief introduction to the implementation environment (implementation architecture) involved in this application is given.
[0044] The embodiments of the present application provide a method for training an image recognition model, which can be applied to a computer device. The embodiments of the present application do not impose any restrictions on the specific form of the computer device. For example, the computer device can specifically be a terminal device or a network device. Among them, the terminal device can be referred to as: terminal, user equipment (UE), terminal device, access terminal, user unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user device, etc. The terminal device can specifically be a mobile phone, an augmented reality (AR) device, a virtual reality (VR) device, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The network device can specifically be a server, etc. Among them, the server can be a physical or logical server, or two or more physical or logical servers sharing different responsibilities and cooperating with each other to implement the various functions of the server.
[0045] In terms of hardware implementation, the above computer device can be implemented by a computer device as shown in Figure 3 the following. As shown in Figure 3 Figure 7 is a schematic diagram of the hardware structure of a computer device 30 provided by the embodiments of the present application. The computer device 30 can be used to implement the functions of the above computer device.
[0046] Figure 3 The computer device 30 shown in Figure 7 may include: a processor 301, a memory 302, a communication interface 303, and a bus 304. The processor 301, the memory 302, and the communication interface 303 can be connected through the bus 304.
[0047] The processor 301 is the control center of the computer device 30, which can be a general-purpose central processing unit (CPU) or other general-purpose processors. Among them, the general-purpose processor can be a microprocessor or any conventional processor.
[0048] As an example, the processor 301 may include one or more CPUs, such as Figure 3 the CPU 0 and CPU 1 shown in Figure 8.
[0049] The memory 302 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0050] In a possible implementation, the memory 302 can exist independently of the processor 301. The memory 302 can be connected to the processor 301 through the bus 304 and is used to store data, instructions, or program code. When the processor 301 calls and executes the instructions or program code stored in the memory 302, the vehicle abnormal behavior discovery method provided by the embodiments of the present application can be implemented.
[0051] In another possible implementation, the memory 302 can also be integrated with the processor 301.
[0052] The communication interface 303 is used for the computer device 20 to connect to other devices through a communication network, and the communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 303 can include a receiving unit for receiving data and a sending unit for sending data.
[0053] The bus 304 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0054] It should be noted that Figure 3 the structure shown in the figure does not constitute a limitation on the computer device 30, except Figure 3In addition to the components shown, the computer device 30 may include more or fewer components than those shown, or combine certain components, or have a different component arrangement.
[0055] For ease of understanding, the following specifically introduces the training method of the image recognition model provided in this application in conjunction with the accompanying drawings.
[0056] As Figure 4 shown, it is a flowchart of a training method of an image recognition model provided in this application. The method includes:
[0057] S401: Determine the uncertainty of each image sample in a plurality of image samples based on at least two image pre-recognition models.
[0058] Among them, the uncertainty is used to characterize the uncertainty of at least two image pre-recognition models in recognizing the image category. At least two image pre-recognition models are both used to determine the category of the image to be recognized from the same plurality of image categories.
[0059] Specifically, the uncertainty of an image sample is used to characterize the uncertainty of at least two image pre-recognition models in recognizing the image category of this image sample. Among them, the image pre-recognition model determines the image category of this image sample based on the features extracted from this image sample. It can be understood that the higher the quality of the image sample with lower uncertainty, that is, the easier it is for the image pre-recognition model to predict the image sample as the category represented by the category identifier of this image sample.
[0060] It should be noted that the plurality of image samples are a training sample set. That is, the plurality of image samples are all samples with known image categories.
[0061] It should be noted that the category of the image to be recognized refers to the category to which the content on the image to be recognized belongs. For example, if the plurality of image categories include cats, dogs, and birds, and the content on the image to be recognized is a cat, then the category of the image to be recognized is "cat".
[0062] Optionally, before S401, the training method of the image pre-recognition model further includes: inputting a plurality of image training samples and the category identifier of each image training sample in the plurality of image training samples into the target neural network model, and training the target neural network model to obtain the image pre-recognition model. Among them, the category of each image training sample predicted by the image pre-recognition model is the same as the category represented by the category identifier of this image training sample.
[0063] In one embodiment, if multiple image training samples include k image classes, the image pre-recognition model is a neural network model that can predict K classes. It should be noted that at least two image pre-recognition models can predict the same multiple image classes. Taking two image pre-recognition models as an example, the first image pre-recognition model can predict three image classes: human, dog, and cat, and the image classes that the second image pre-recognition model can predict are also: human, dog, and cat.
[0064] It should be noted that each of at least two image pre-recognition models can be trained by the above optional method. Among them, the multiple image training samples used to train each image pre-recognition model can be the same or different. In addition, the target neural networks used to train each image pre-recognition model can be the same or different.
[0065] In one embodiment, the target neural network model can be any one of VGG, ResNet, and Transformer. Of course, other neural network models can also be used, and the present application is not limited to a specific recognition model structure.
[0066] S402: Determine the first target image samples with uncertainty less than or equal to the first preset threshold among multiple image samples as the prototype samples of the image classes to which the first target image samples belong, or determine the second target image samples with uncertainty less than the first preset number of other image samples among multiple image samples as the prototype samples of the image classes to which the second target image samples belong.
[0067] Among them, other image samples are other image samples among multiple image samples except the second target image samples.
[0068] In one embodiment, multiple image samples belong to multiple image classes, and each image class includes at least one image sample. Determine the image samples with uncertainty less than or equal to the first preset threshold in each image class as the prototype samples of that image class. For example, 5 image samples belong to 2 image classes. Among them, image class A includes image sample 1 and image sample 2, image class B includes image sample 3, image sample 4, and image sample 5, and the uncertainties of image sample 1, image sample 3, and image sample 5 are all less than the first preset threshold. At this time, determine image sample 1 as the prototype sample of image class A, and image sample 3 and image sample 5 as the prototype samples of image class B.
[0069] In another implementation, multiple image samples are sorted in ascending order of uncertainty to obtain a sequence. Then, the first preset number of image samples in the pre-sequence are all determined as the prototype samples of the image category of each image sample. Of course, the sequence can also sort multiple image samples in descending order of uncertainty, and the last first preset number of image samples in the sequence are all determined as the prototype samples of the image category of each image sample. Through the foregoing method, the first preset number of second target image samples with uncertainty less than other image samples among multiple image samples are determined as the prototype samples of the image category to which the second target image samples belong.
[0070] Optionally, multiple image categories include the first preset number of prototype samples, where the first preset number of prototype samples includes the target preset number of prototype samples for each image category.
[0071] For example, multiple image categories include 10 prototype samples. The multiple image categories include image category 1 and image category 2. Among them, image category 1 includes 4 prototype samples, and image category 2 includes 6 prototype samples. That is, the 10 prototype samples include 4 prototype samples of image category 1 and 6 prototype samples of image category 2.
[0072] In one implementation, multiple image categories include image category D. Multiple image samples belonging to image category D are sorted in ascending order of uncertainty to obtain a sequence. Then, the first target preset number of image samples in the pre-sequence are determined as the prototype samples of image category D. Of course, the sequence can also sort multiple image samples belonging to image category D in descending order of uncertainty, and the last target preset number of image samples in the sequence are all determined as the prototype samples of image category D. Through the foregoing method, the target preset number of image samples with uncertainty less than other image samples in each image category are determined as the prototype samples of that image category. Further, the method for determining the prototype samples of each image category among multiple image categories can refer to the foregoing method, so as to further implement determining the first preset number of second target image samples with uncertainty less than other image samples among multiple image samples as the prototype samples of the image category to which the second target image samples belong; it should be noted that the prototype samples are representative samples in each image category. That is to say, the prototype samples are samples in each image category that "are more likely to determine the image category according to the characteristics of the image samples". For example, if image sample A is a "real photo of a cat" and image sample B is a "simple drawing of a cat", obviously, it is easier to determine the image category according to the characteristics of image sample A. Therefore, image sample A is the prototype sample of the image category "cat".
[0073] By determining the eligible target image samples (the first target image samples or the second target image samples) in each image category as the prototype samples of that image category, low-quality noise samples in each image category can be filtered out, thereby effectively avoiding the interference of low-quality noise samples on the prototype samples, enabling the samples used for training the model to have stronger noise resistance, and further improving the robustness of the finally obtained model.
[0074] Optionally, the multiple image categories include a third image category, which can be any one of the multiple image categories. The third image category can be the same as the first image category or the second image category, or a different image category. In the case where there are no target image samples less than or equal to the first preset threshold in the third image category, the image sample with the smallest uncertainty in the third image category can be determined as the prototype image sample, or the multiple image samples belonging to the third image category are sorted in ascending order of uncertainty to obtain a sequence, and then the first target preset number of image samples in the sequence are all determined as the prototype samples of the third image category. Of course, the sequence can also be sorted in descending order of uncertainty for the multiple image samples belonging to the third image category, and the last target preset number of image samples in the sequence are all determined as the prototype samples of the third image category.
[0075] S403: Based on the prototype samples of each image category in the multiple image categories, train a neural network model to obtain an image recognition model.
[0076] Among them, the image recognition model is used to determine the category of the image to be recognized from the above multiple image categories. Specifically, the image recognition model at least includes a feature extraction layer and a category recognition layer. The feature extraction layer is used to extract the feature vector of the image to be recognized, and the category recognition layer predicts the category of the image to be recognized based on the feature vector extracted by the feature extraction layer.
[0077] It should be noted that the neural network model in S403 is an open-set recognition model. Therefore, the finally obtained image recognition model is also an open-set recognition model.
[0078] Optionally, input the prototype samples of each image category in the multiple image categories into the neural network model to train the neural network model, so that when the image recognition model extracts the feature vectors of each prototype sample, the feature distances of the feature vectors of the prototype samples in the same image category are close, and the feature distances of the feature vectors of the prototype samples in different image categories are pulled apart.
[0079] Optionally, the method further includes: inputting the test image samples into the image recognition model to obtain a prediction result.
[0080] In one embodiment, a confidence threshold of the image recognition model is preset in advance. The prediction result includes a prediction probability, which is used to characterize the probability that the test image sample is a known class sample. Specifically, when the prediction probability is greater than or equal to the confidence threshold, it is determined that the test image sample is a known class sample; when the prediction probability is less than the confidence threshold, it is determined that the test image sample is an unknown class sample.
[0081] In another embodiment, a test sample feature vector of the test image sample and feature vectors of prototype samples of each image class are obtained based on the image recognition model. Based on the feature distance between the test sample feature vector and the feature vectors of the prototype samples of each image class, and based on the relationship between the minimum feature distance and a preset distance threshold, the class of the test image sample is determined. Specifically, when the minimum feature distance is less than the preset distance threshold, it is determined that the test image sample is a known class sample; when the minimum feature distance is greater than the preset distance threshold, it is determined that the test image sample is an unknown class sample.
[0082] In the above embodiments, through the uncertainty of each image sample, multiple image samples are screened, low-quality noise samples among the multiple image samples are filtered out, prototype samples of each image class are obtained, and a neural network model is trained according to the prototype samples to obtain an image recognition model. By using the prototype samples for model training, it is not necessary to introduce the prototype features of each class, and it is not necessary to use the distance between the features of each sample and the prototype features as an additional parameter for training and learning. Furthermore, when the number of sample classes increases, a large number of additional parameters are avoided, storage space and computing time are saved, the training difficulty is reduced, and the image recognition model can achieve a better convergence effect, thereby realizing better open-set recognition performance.
[0083] Optionally, at least two image pre-recognition models include a first image pre-recognition model and a second image pre-recognition model, and the multiple image samples include a first image sample, where the first image sample is any one of the multiple image samples.
[0084] The above S401 may include:
[0085] Step 1: Extract a first feature vector of each image sample based on the first image pre-recognition model. The first feature vector is used for the first image pre-recognition model to determine the class of the image to be recognized from multiple image classes.
[0086] Optionally, the multiple image samples are input into the first image pre-recognition model, and the features of each image sample are extracted to obtain the first feature vector of each image sample.
[0087] For example, the first image pre-recognition model extracts the features of each image sample to obtain the first feature vector of each image sample. For example, for image sample 1 (denoted by x 1 ), image sample 2 (denoted by x 2 ), ……, image sample N (denoted by x N ), feature extraction is performed to obtain the first feature vector (z 1 ) of x 1 , the first feature vector (z 2 ) of x 2 , ……, the first feature vector (z N ) of x N . Exemplarily, the first feature vectors of the multiple image samples can be expressed as Z 1 =(z 1 , z 2 , ……, z N ).
[0088] Step 2: Extract the second feature vector of each image sample based on the second image pre-recognition model; wherein, the second feature vector is used for the second image pre-recognition model to determine the category of the image to be recognized from multiple image categories.
[0089] Optionally, input the multiple image samples into the second image pre-recognition model, extract the features of each image sample to obtain the second feature vector of each image sample.
[0090] For example, the second image pre-recognition model extracts the features of each image sample to obtain the second feature vector of each image sample. For example, for image sample 1 (denoted by x 1 ), image sample 2 (denoted by x 2 ), ……, image sample N (denoted by x N ), feature extraction is performed to obtain the second feature vector (z 1 ’) of x 1 , the second feature vector (z 2 ’) of x 2 , ……, the second feature vector (z N ) of x N . Exemplarily, the first feature vectors of the multiple image samples can be expressed as Z 2 =(z 1 ’, z 2 ’, ……, z N ’).
[0091] It should be noted that the image pre-recognition model at least includes a feature extraction layer and a category recognition layer. The feature extraction layer is used to extract the features of the image sample, and the category recognition layer is used to predict the category of the image sample according to the features extracted by the feature extraction layer.
[0092] Optionally, the feature vector of the image sample is the feature vector extracted by the feature extraction layer of the image pre-recognition model.
[0093] Step 3: Based on the first feature vector of the first image sample and the first feature vectors of each of the multiple image samples, obtain a first feature distribution vector. The first feature distribution vector is used to characterize the distance distribution between the first feature vector of the first image sample and the first feature vectors of other image samples.
[0094] It should be noted that other image samples are the image samples other than the first image sample among the multiple image samples.
[0095] Optionally, respectively determine the first feature distances between the first feature vector of the first image sample and the first feature vectors of each image sample, and obtain the first feature distribution vector of the first image sample based on the determined multiple first feature distances. For example, the multiple image samples include a total of N image samples, that is, image sample 1 (denoted by x 1 ), image sample 2 (denoted by x 2 ), ……, image sample N (denoted by x N ), respectively determine the distances between the first feature vector of image sample 1 and the first feature vectors of image sample 1, ……, the Nth image sample, and a total of N first feature distances are obtained. Based on the N first feature distances, the first feature distribution vector of image sample 1 is obtained.
[0096] In one implementation, the first feature distribution vector is characterized by distri(z 1 ), and distri(z 1 ) satisfies the following formula: distri(z 1 ) = <d(z 1 , z 1 ), d(z 1 , z 2 ), ……, d(z 1 , z N )>. Wherein, d(z 1 , z 2 ) represents the feature distance from the first feature vector (z 1 ) of x 1 to the first feature vector (z 2 ) of x 2 . The calculation method of the feature distance can adopt any calculation method for measuring the distance between vectors, such as Euclidean distance, cosine distance, Mahalanobis distance, KL divergence, etc.
[0097] Step 4: Based on the second feature vectors of the first image sample and the second feature vectors of each of the multiple image samples, obtain a second feature distribution vector. The second feature distribution vector is used to characterize the distance distribution between the second feature vector of the first image sample and the second feature vectors of other image samples.
[0098] It should be noted that the other image samples are the image samples among the multiple image samples except the first image sample.
[0099] Optionally, respectively determine the second feature distances between the second feature vector of the first image sample and the second feature vectors of each image sample, and obtain the second feature distribution vector of the first image sample based on the determined multiple second feature distances. For example, the multiple image samples include N image samples in total, that is, image sample 1 (denoted by x 1 ), image sample 2 (denoted by x 2 ), ……, image sample N (denoted by x N ). Respectively determine the distances between the second feature vector of image sample 1 and the second feature vectors of image sample 1, ……, the Nth image sample, and a total of N second feature distances are obtained. Based on these N second feature distances, the second feature distribution vector of image sample 1 is obtained.
[0100] In one implementation, the second feature distribution vector is characterized by distri(z 1 ’), and distri(z 1 ’) satisfies the following formula: distri(z 1 ’) = <d(z 1 ’, z 1 ’), d(z 1 ’, z 2 ’), …, d(z 1 ’, z N ’)>. Wherein, d(z 1 ’, z 2 ’) represents the feature distance from the second feature vector (z 1 ’) of x 1 to the second feature vector (z 2 ’) of x 2 .
[0101] Step 5: Based on the first feature distribution vector and the second feature distribution vector, obtain the uncertainty of the first image sample.
[0102] Optionally, determine the feature distance between the first feature distribution vector and the second feature distribution vector as the uncertainty of the first image sample.
[0103] In one implementation, the uncertainty of the first image sample is represented by Uncertainty(x1 ) Characterization. Uncertainty(x 1 ) satisfies the following formula: Uncertainty(x 1 ) = d(distri(z 1 ), distri(z 1 ')). Wherein, d(distri(z 1 ), distri(z 1 ')) represents the feature distance between the first feature distribution vector and the second feature distribution vector.
[0104] It should be noted that since the first image sample is any one of the multiple image samples, the uncertainty of each image sample in the multiple image samples can be determined by the methods of the above steps 1 to 5.
[0105] In this application, if there are three or more image pre-recognition models, for example, there are three image pre-recognition models, that is, at least two image pre-recognition models further include a third image pre-recognition model. Then S401 further includes: extracting the third feature vector of each image sample based on the third image pre-recognition model (the method can refer to step 1 or step 2), and obtaining the third feature distribution vector based on the third feature vector of the first image sample and the third feature vector of each image sample of the multiple image samples respectively (the method can refer to step 3 or step 4), and the third feature distribution vector is used to represent the distance distribution between the third feature vector of the first image sample and the third feature vectors of other image samples. Further, based on the first feature distribution vector, the second feature distribution vector, and the third feature distribution vector, the uncertainty of the first image sample is obtained (the method can refer to step 5).
[0106] When there are more than three image pre-recognition models, the execution processes of the fourth image pre-recognition model to the Nth image pre-recognition model can refer to the third image pre-recognition model, which will not be elaborated here. In the above embodiments, the feature vectors of each image sample extracted by at least two image pre-recognition models are used to determine multiple feature distribution vectors of each image sample, and the uncertainty of each image sample is determined based on the multiple feature distribution vectors of each image sample. Since the feature vectors of each image sample are parameters for determining the category to which the image sample belongs, therefore, determining the feature distribution vector of the image sample based on the feature vector of the image sample, and then determining the uncertainty of the image category recognized by the image pre-recognition model for the image sample based on the feature distribution vector can improve the accuracy of the uncertainty of each image sample, so as to accurately screen out high-quality prototype samples.
[0107] Optionally, in combination with Figure 4 , such as Figure 5As shown, S403 includes S403a - S403b.
[0108] S403a: Determine the set of target prototype samples for each image category.
[0109] Among them, the multiple image categories include the first image category. The set of target prototype samples for the first image category includes at least the first prototype sample, and the first prototype sample is the one with the smallest uncertainty among the prototype samples of the first image category.
[0110] For example, the prototype samples of image category A include image sample 1 and the second image sample 2. The uncertainty of image sample 1 is less than that of the second image sample. The prototype samples of image category B include image sample 3 and image sample 4. The uncertainty of image sample 3 is less than that of image sample 4. The prototype samples of image category C include image sample 5, image sample 6, and image sample 7. The uncertainty of image sample 5 is less than that of image sample 6 and image sample 7. Based on this, image sample 1 is determined as an element of the set of target prototype samples belonging to image category A, image sample 3 is determined as an element of the set of target prototype samples belonging to image category B, and image sample 5 is determined as an element of the set of target prototype samples belonging to image category C.
[0111] It should be noted that the first image category is any one of the multiple image categories. Therefore, for the method of determining the set of target prototype samples for each image category among the multiple image categories, the method of determining the set of target prototype samples for the first image category can be referred to.
[0112] S403b: Based on the set of target prototype samples for each image category, train a neural network model to obtain an image recognition model.
[0113] In one implementation, input the set of target prototype samples for each image category into the neural network model and train the neural network model to obtain an image recognition model.
[0114] In the above - mentioned embodiment, by screening the prototype samples for each image category and using the prototype sample with the smallest uncertainty in each image category, that is, the prototype sample with the highest sample quality in each image category, to train the neural network model to obtain an image recognition model, thereby reducing the amount of data in the model training process, reducing the time for model training, and reducing the difficulty of model training.
[0115] Optionally, the training method of the image recognition model further includes:
[0116] Step 1: Determine the minimum feature distance between the second prototype sample in the first image category and the set of target prototype samples of the first image category, where the second prototype sample is a prototype sample that does not belong to the set of target prototype samples of the first image category.
[0117] Among them, the minimum feature distance is used to represent the minimum feature distance among the feature distances between the second prototype sample and each prototype sample in the set of target prototype samples.
[0118] For example, the set of target prototype samples of the first image category includes element 1, element 2, and element 3. The feature distance between the second prototype sample 1 of the first image category and element 1 is X, the feature distance between the second prototype sample 1 and element 2 is Y, and the feature distance between the second prototype sample 1 and element 3 is Z. Among them, X is less than Y and X is less than Z. Therefore, X is the minimum feature distance of the second prototype sample 1.
[0119] It should be noted that in the case where the set of target prototype samples of the first image category only includes the first prototype sample, the feature distance between the second prototype sample 1 and the first prototype sample is the minimum feature distance of the second prototype sample 1.
[0120] Furthermore, referring to the method for determining the minimum feature distance of the second prototype sample 1, determine the minimum feature distance between each second prototype sample in the first image category and the set of target prototype samples of the first image category.
[0121] Step 2: Determine the second prototype sample with the largest minimum feature distance as an element belonging to the set of target prototype samples of the first image category.
[0122] For example, the first image category includes 3 second prototype samples: second prototype sample 1, second prototype sample 2, and second prototype sample 3. The minimum feature distance of the second prototype sample 1 is Z1, the minimum feature distance of the second prototype sample 2 is Z2, and the minimum feature distance of the second prototype sample 3 is Z3. Among them, Z1 is less than Z2 and Z1 is less than Z3. Therefore, the second prototype sample 1 is determined as an element belonging to the set of target prototype samples of the first image category.
[0123] Step 3: Repeat the above Step 1 and Step 2 until the largest minimum feature distance is less than or equal to the second preset threshold, or the number of elements in the set of target prototype samples of the first image category is equal to the second preset number.
[0124] In one implementation, steps one and two are repeatedly executed until the maximum minimum feature distance is less than or equal to a second preset threshold. For example, when step one is repeatedly executed for the Nth time, the minimum feature distance of the second prototype sample N among the multiple second prototype samples of the first image category is the largest. At this time, if the minimum feature distance of the second prototype sample N is less than the second preset threshold, step two is stopped from continuing to be executed, and the task of repeatedly executing steps one and two ends.
[0125] It should be noted that the first image category is any one of the multiple image categories. Therefore, for the determination method of the target prototype sample set of each image category in the multiple image categories, reference can be made to the determination method of the target prototype sample set of the first image category.
[0126] In the above embodiment, by determining the second prototype samples whose minimum feature distance from the target prototype sample set is greater than the second preset threshold as the elements of the target prototype sample set belonging to the first image category, the prototype samples with a relatively small minimum feature distance from the target prototype sample set in the prototype samples are eliminated, that is, the samples with a relatively high similarity to the target prototype sample set are eliminated, so as to filter out the redundant and repeated prototype samples among the multiple prototype samples, retain the diversity of the target prototype sample set, and then improve the effectiveness of model training, and further improve the robustness of the image recognition model.
[0127] In another implementation, steps one and two are repeatedly executed until the number of the target prototype sample set of the first image category is equal to a second preset number. For example, after steps one and two are repeatedly executed M times, the number of elements in the target prototype sample set of the first image category is equal to the second preset number. At this time, the task of repeatedly executing steps one and two ends.
[0128] In the above embodiment, by determining the second prototype samples with the largest minimum feature distance among the second preset number as the elements of the target prototype sample set belonging to the first image category, the prototype samples with a relatively small minimum feature distance from the target prototype sample set in the prototype samples are eliminated, that is, the samples with a relatively high similarity to the target prototype sample set are eliminated, so as to filter out the redundant and repeated prototype samples among the multiple prototype samples, retain the diversity of the target prototype sample set, and then improve the effectiveness of model training, and further improve the robustness of the image recognition model.
[0129] Optionally, the multiple image categories include a second image category, and the second image category can be any one of the multiple image categories. The second image category may be the same as or different from the first image category described above. Based on the prototype samples of each image category in the multiple image categories, a neural network model is trained to obtain an image recognition model, including: training the neural network model based on the prototype samples of each image category and the loss function of the prototype samples of each image category to obtain an image recognition model. Among them, the loss function of the prototype samples of the second image category is determined based on a first feature distance and a second feature distance. The first feature distance is the feature distance between each image sample in the second image category and the prototype sample of the second image category, and the second feature distance is the feature distance between each image sample in the second image category and the prototype samples of other image categories in the multiple image categories.
[0130] Optionally, the prototype samples of each image category are input into the neural network model, the neural network model is trained, and the training process of the neural network model is constrained by the loss function of the prototype samples of each image category to obtain an image recognition model.
[0131] In one implementation, the loss function of the prototype samples of each image category is represented by L p and L p satisfies the following formula:
[0132]
[0133] where N represents the number of multiple image samples. The feature vector of the i-th image sample x i is z i , the i-th image sample x i belongs to the image category m, P m is the prototype sample set of the m-th image category, and z(P m ) represents the feature vectors of all prototype samples in the prototype sample set P m . Among them, z(P m ) can be obtained by inputting the prototype samples in P m into the recognition model for extraction. d(z i , z(P m )) represents the feature distance from the image sample x i to the prototype sample set P m of the same category (i.e., the image category m). For example, the minimum feature distance from the image sample x i to all prototype samples in P m can be used as d(z i , z(P m )) Of course, other methods for measuring the feature distance from a single image sample to multiple prototype sample sets can also be used. Pu The set of prototype samples with the smallest feature distance characterizing the distance image sample x i is a set of prototype samples of other image classes outside the image class m. δ represents the fourth preset threshold.
[0134] Optionally, the training process of the model can also be constrained by the classification loss function L cls . For example, L cls can adopt the SoftMaX loss function or the cross-entropy loss function. By using the classification loss function L cls and the loss functions of the prototype samples of each image class to jointly constrain the model training process, the model training process can be optimized, thereby improving the feature extraction ability of the image recognition model and the accuracy of identifying image classes.
[0135] By constraining the training process of the neural network model through the loss functions of the prototype samples of each image class, the feature distance from the image sample x i to the set of prototype samples P m of the image class m is less than the third preset threshold. For example, the third preset threshold is greater than or equal to 0 and less than δ, and the feature distance from the image sample x i to the set of prototype samples of other image classes (i.e., image classes other than the image class m among multiple image classes) is greater than the fourth preset threshold. For example, the fourth preset threshold is δ. It should be noted that the second image class is any one of the multiple image classes. Therefore, the method for determining the loss functions of the prototype samples of each image class among multiple image classes can refer to the method for determining the loss functions of the prototype samples of the second image class.
[0136] In the above embodiments, by constraining the training process of the neural network model through the loss functions of the prototype samples of each image class, the feature distances of the image samples of known different image classes are maximally separated, so that the image samples of unknown classes have a greater probability of falling in the region between different known image classes, which helps to better distinguish known class samples and unknown class samples.
[0137] The above mainly introduces the solutions of the embodiments of the present application from the perspective of methods. It can be understood that in order for a computer device to implement the above functions, it includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0138] The embodiments of the present application can divide the functional units of the computer device according to the above method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0139] Exemplarily, Figure 6 FIG. shows a possible structural schematic diagram of a training device for an image recognition model (denoted as the training device 60 for the image recognition model) involved in the above embodiments. The training device 60 for the image recognition model includes an identification unit 601, a determination unit 602, and a training unit 603. The identification unit 601 is configured to determine the uncertainty of each image sample in a plurality of image samples based on at least two image pre-recognition models; wherein, the uncertainty is used to characterize the uncertainty of at least two image recognition models in recognizing image categories; at least two image pre-recognition models are both used to determine the category of the image to be recognized from the same plurality of image categories. For example, Figure 4 as shown in S401. The determination unit 602 is configured to determine the first target image sample with an uncertainty less than or equal to a first preset threshold in the plurality of image samples as the prototype sample of the image category to which the first target image sample belongs, or determine the second target image sample with an uncertainty less than a first preset number of other image samples in the plurality of image samples as the prototype sample of the image category to which the second target image sample belongs. For example, Figure 4 as shown in S402. The training unit 603 is configured to train a neural network model based on the prototype sample of each image category in the plurality of image categories to obtain an image recognition model; wherein, the image recognition model is used to determine the category of the image to be recognized from the plurality of image categories, or determine that the category of the image to be recognized does not belong to the plurality of image categories. For example, Figure 4 as shown in S403, andFigure 5 as shown in S403a - S403b.
[0140] Optionally, at least two image pre - recognition models include a first image pre - recognition model and a second image pre - recognition model, and multiple image samples include a first image sample; the recognition unit 601 is specifically configured to: extract a first feature vector of each image sample based on the first image pre - recognition model; wherein, the first feature vector is used for the first image pre - recognition model to determine the category of the image to be recognized from multiple image categories; extract a second feature vector of each image sample based on the second image pre - recognition model; wherein, the second feature vector is used for the second image pre - recognition model to determine the category of the image to be recognized from multiple image categories; obtain a first feature distribution vector based on the first feature vector of the first image sample and the first feature vectors of each image sample respectively; wherein, the first feature distribution vector is used to characterize the distance distribution between the first feature vector of the first image sample and the first feature vectors of other image samples; obtain a second feature distribution vector based on the second feature vector of the first image sample and the second feature vectors of each image sample respectively; wherein, the second feature distribution vector is used to characterize the distance distribution between the second feature vector of the first image sample and the second feature vectors of other image samples; obtain the uncertainty of the first image sample based on the first feature distribution vector and the second feature distribution vector.
[0141] Optionally, the training unit 603 is specifically configured to determine a target prototype sample set for each image category, multiple image categories include a first image category, and the target prototype sample set of the first image category includes at least a first prototype sample, and the first prototype sample is the one with the smallest uncertainty among the prototype samples of the first image category; train a neural network model based on the target prototype sample sets of each image category to obtain an image recognition model.
[0142] Optionally, the training unit 603 is further configured to: determine the minimum feature distance between a second prototype sample in the first image category and the target prototype sample set of the first image category, where the second prototype sample is a prototype sample that does not belong to the target prototype sample set of the first image category; determine the second prototype sample with the largest minimum feature distance as an element belonging to the target prototype sample set of the first image category; repeat the above steps until the largest minimum feature distance is less than or equal to a second preset threshold, or the number of elements in the target prototype sample set of the first image category is equal to a second preset number.
[0143] Optionally, the multiple image categories include a second image category. The training unit 603 is specifically configured to train a neural network model based on the prototype samples of each image category and the loss function of the prototype samples of each image category, so as to obtain an image recognition model. Among them, the loss function of the prototype samples of the second image category is determined based on a first feature distance and a second feature distance. The first feature distance is the feature distance between each image sample in the second image category and the prototype sample of the second image category, and the second feature distance is the feature distance between each image sample in the second image category and the prototype samples of other image categories in the multiple image categories.
[0144] For the specific description of the above optional manner, reference may be made to the foregoing method embodiments, which will not be elaborated here. In addition, the explanations of any of the above-provided training devices 60 for image recognition models and the descriptions of the beneficial effects can refer to the corresponding method embodiments above and will not be elaborated.
[0145] As an example, in combination with Figure 3 , the functions implemented by some or all of the recognition unit 601, the determination unit 602, and the training unit 603 in the training device 60 for the image recognition model can be executed by Figure 3 the processor 301 in Figure 3 and implemented by the program code in the memory 302 in
[0146] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, the computer is enabled to execute the method performed by any of the above-provided computer devices.
[0147] For the explanations and descriptions of the beneficial effects of the relevant content in any of the above-provided computer-readable storage media, reference may be made to the corresponding embodiments above, which will not be elaborated here.
[0148] An embodiment of the present application further provides a chip. The chip integrates a control circuit for implementing the functions of the above computer device and one or more ports. Optionally, the functions supported by the chip can be referred to the above, which will not be elaborated here. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a random access memory, etc. The above processing unit or processor can be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0149] An embodiment of the present application further provides a computer program product containing instructions. When the instructions run on a computer, the computer is made to execute any one of the methods in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, they generate all or part of the processes or functions according to the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.
[0150] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of the present application, such as but not limited to, the above-mentioned memory, computer-readable storage medium, and communication chip, etc., are all non-transitory.
[0151] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0152] Although the present application has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0153] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present application defined by the appended claims, and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
Claims
1. A training method for an image recognition model, characterized in that, it includes: Based on at least two image pre-recognition models, determining the uncertainty of each image sample in a plurality of image samples; wherein, the uncertainty is used to characterize the uncertainty of the at least two image pre-recognition models in recognizing image categories; the at least two image pre-recognition models are both used to determine the category of the image to be recognized from the same plurality of image categories; the plurality of image categories include a second image category; Determining the first target image samples with uncertainty less than or equal to a first preset threshold in the plurality of image samples as the prototype samples of the image category to which the first target image samples belong, or determining the second target image samples with uncertainty less than a first preset number of other image samples in the plurality of image samples as the prototype samples of the image category to which the second target image samples belong; Based on the prototype samples of each image category and the loss function of the prototype samples of each image category, training a neural network model to obtain an image recognition model; wherein, the loss function of the prototype samples of the second image category is determined based on a first feature distance and a second feature distance, the first feature distance is the feature distance between each image sample in the second image category and the prototype sample of the second image category, and the second feature distance is the feature distance between each image sample in the second image category and the prototype samples of other image categories in the plurality of image categories; wherein, the image recognition model is used to determine the category of the image to be recognized from the plurality of image categories, or determine that the category of the image to be recognized does not belong to the plurality of image categories.
2. The method according to claim 1, characterized in that, the at least two image pre-recognition models include a first image pre-recognition model and a second image pre-recognition model, and the plurality of image samples include a first image sample; the determining the uncertainty of each image sample in the plurality of image samples based on at least two image pre-recognition models includes: Extracting a first feature vector of each image sample based on the first image pre-recognition model; wherein, the first feature vector is used for the first image pre-recognition model to determine the category of the image to be recognized from the plurality of image categories; Extracting a second feature vector of each image sample based on the second image pre-recognition model; wherein, the second feature vector is used for the second image pre-recognition model to determine the category of the image to be recognized from the plurality of image categories; Based on the first feature vector of the first image sample and the first feature vectors of each image sample respectively, obtaining a first feature distribution vector; wherein, the first feature distribution vector is used to characterize the distance distribution between the first feature vector of the first image sample and the first feature vectors of other image samples; Based on the second feature vector of the first image sample and the second feature vectors of each image sample respectively, obtaining a second feature distribution vector; wherein, the second feature distribution vector is used to characterize the distance distribution between the second feature vector of the first image sample and the second feature vectors of other image samples; Based on the first feature distribution vector and the second feature distribution vector, the uncertainty of the first image sample is obtained.
3. The method according to claim 1 or 2, wherein, training a neural network model based on prototype samples of each image category among multiple image categories to obtain an image recognition model, including: determining a target prototype sample set of each image category, the multiple image categories including a first image category, the target prototype sample set of the first image category at least including a first prototype sample, and the first prototype sample being the one with the smallest uncertainty among the prototype samples of the first image category; training a neural network model based on the target prototype sample set of each image category to obtain an image recognition model.
4. The method according to claim 3, wherein, the method further includes: determining the minimum feature distance between a second prototype sample in the first image category and the target prototype sample set of the first image category, where the second prototype sample is a prototype sample that does not belong to the target prototype sample set of the first image category; determining the second prototype sample with the largest minimum feature distance as an element belonging to the target prototype sample set of the first image category; repeatedly executing the above steps until the largest minimum feature distance is less than or equal to a second preset threshold, or the number of elements in the target prototype sample set of the first image category is equal to a second preset number.
5. An apparatus for training an image recognition model, wherein, comprising: an identification unit, configured to determine the uncertainty of each image sample among multiple image samples based on at least two image pre-recognition models; wherein, the uncertainty is used to characterize the uncertainty of the at least two image pre-recognition models in recognizing the image category; the at least two image pre-recognition models are both used to determine the category of an image to be recognized from the same multiple image categories; the multiple image categories include a second image category; a determination unit, configured to determine the first target image sample with uncertainty less than or equal to a first preset threshold among the multiple image samples as a prototype sample of the image category to which the first target image sample belongs, or determine the second target image sample with uncertainty less than a first preset number of other image samples among the multiple image samples as a prototype sample of the image category to which the second target image sample belongs; A training unit, configured to train a neural network model based on the prototype samples of each image category and the loss function of the prototype samples of each image category, to obtain an image recognition model; wherein, the loss function of the prototype samples of the second image category is determined based on a first feature distance and a second feature distance, the first feature distance is the feature distance between each image sample in the second image category and the prototype sample of the second image category, and the second feature distance is the feature distance between each image sample in the second image category and the prototype samples of other image categories among the multiple image categories; wherein, the image recognition model is used to determine the category of the image to be recognized from the multiple image categories, or determine that the category of the image to be recognized does not belong to the multiple image categories.
6. The apparatus according to claim 5, wherein, the at least two image pre-recognition models include a first image pre-recognition model and a second image pre-recognition model, and the multiple image samples include first image samples; the recognition unit is specifically configured to: extract a first feature vector of each image sample based on the first image pre-recognition model; wherein, the first feature vector is used for the first image pre-recognition model to determine the category of the image to be recognized from the multiple image categories; extract a second feature vector of each image sample based on the second image pre-recognition model; wherein, the second feature vector is used for the second image pre-recognition model to determine the category of the image to be recognized from the multiple image categories; obtain a first feature distribution vector based on the first feature vector of the first image sample and the first feature vectors of each image sample respectively; wherein, the first feature distribution vector is used to characterize the distance distribution between the first feature vector of the first image sample and the first feature vectors of other image samples; obtain a second feature distribution vector based on the second feature vector of the first image sample and the second feature vectors of each image sample respectively; wherein, the second feature distribution vector is used to characterize the distance distribution between the second feature vector of the first image sample and the second feature vectors of other image samples; obtain the uncertainty of the first image sample based on the first feature distribution vector and the second feature distribution vector.
7. The apparatus according to claim 6, wherein, the training unit is specifically configured to determine a target prototype sample set for each image category, the multiple image categories include a first image category, and the target prototype sample set of the first image category includes at least a first prototype sample, and the first prototype sample is the one with the smallest uncertainty among the prototype samples of the first image category; based on the target prototype sample set of each image category, train a neural network model to obtain an image recognition model; The training unit is further configured to: determine the minimum feature distance between a second prototype sample in the first image category and the set of target prototype samples of the first image category, where the second prototype sample is a prototype sample that does not belong to the set of target prototype samples of the first image category; determine the second prototype sample with the maximum minimum feature distance as an element belonging to the set of target prototype samples of the first image category; repeat the above steps until the maximum minimum feature distance is less than or equal to a second preset threshold, or the number of elements in the set of target prototype samples of the first image category is equal to a second preset number.
8. A computer device, characterized in that it includes: a processor; The processor is connected to a memory, and the memory is used to store computer execution instructions. The processor executes the computer execution instructions stored in the memory so that the computer device implements the method according to any one of claims 1-4.
9. A computer-readable storage medium, characterized in that it is used to store computer instructions, and when the computer instructions are run on a computer, the computer is caused to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Intelligent legal scene classification system and method
CN112765315A
Hyperspectral image semi-supervised classification method based on small sample learning
CN113408605A