Model Training and Scene Recognition Methods, Devices, Equipment, and Media
By training the scene recognition model, combining the similarity between image features and class-center features, the problem of existing systems not being able to identify scene categories that do not belong to the closed image set is solved, and accurate recognition of closed and non-closed image sets is achieved, and recognition accuracy and performance are improved.
Patent Information
- Application Number
- CN202111159087.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-09-30
AI Technical Summary
The existing scene recognition system cannot accurately process scene-categorized images that do not belong to the closed image set, resulting in errors in recognition results and affecting downstream algorithm processing.
By obtaining the scene probability vector and sample features of the sample image, combining scene labels and class-center features, the original scene recognition model is trained to obtain the trained scene recognition model, and the similarity between image features and class-center features is determined, and the scene category to which the image belongs is determined.
It realizes accurate recognition of scene categories in the closed image set, and can process scene category images that do not belong to the closed image set, improving the accuracy and performance of the scene recognition model.
Smart Images

Figure CN113902944B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular, to a method, apparatus, device and medium for training a model and scene recognition. Background Art
[0002] With the development of multimedia technology, the types of video images that people watch daily are increasing, and the products involved in video content are also becoming more and more abundant. Automatically identifying and classifying the scene information of images helps the machine better understand the images and helps the downstream algorithm development to develop functions for different scenes.
[0003] With the development of neural networks in the field of vision, their performance in image classification tasks has also surpassed most traditional algorithms. However, most neural network-based scene recognition systems are trained and tested on a closed image set, that is, the scene recognition system can only recognize the scene categories included in the closed image set. However, in actual applications, since the scene categories to which all images may belong are not enumerable, the scene category to which the current image to be scene-recognized actually belongs may not be the scene category included in the closed image set. However, if the scene recognition system is used to recognize the scene category to which the image belongs, an incorrect result will be obtained, which will affect the processing of the downstream algorithm.
[0004] Therefore, there is an urgent need for a scene recognition system that can not only accurately recognize images belonging to the scene categories included in the closed image set, but also process images that do not belong to the scene categories included in the closed image set. Summary of the Invention
[0005] The present application provides a method, apparatus, device and medium for training a model and scene recognition, which is used to solve the problem that the existing scene recognition system cannot accurately process images that do not belong to the scene categories included in the closed image set.
[0006] The present application provides a method for training a scene recognition model, and the method includes:
[0007] Obtain any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs;
[0008] Determine the scene probability vector corresponding to the sample image and the sample feature of the sample image through the original scene recognition model; wherein, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively;
[0009] Train the original scene recognition model based on the scene probability vector, the scene label, the sample feature, the class center feature corresponding to the first scene category, and the class center feature corresponding to the second scene category to obtain a trained scene recognition model; wherein the second scene category is the scene category other than the first scene category among each scene category.
[0010] This application provides a scene recognition method, and the method includes:
[0011] Determine the image feature of the image to be recognized through a pre-trained scene recognition model;
[0012] Determine the similarities between the image feature and the target class center features of each scene category respectively;
[0013] Determine whether each scene category contains the scene category to which the image to be recognized belongs according to each similarity and the similarity threshold;
[0014] If it is determined that each scene category contains the scene category to which the image to be recognized belongs, then determine the scene category to which the image to be recognized belongs through the scene recognition model;
[0015] If it is determined that each scene category does not contain the scene category to which the image to be recognized belongs, then do not continue to recognize the scene category to which the image to be recognized belongs.
[0016] This application provides a scene recognition model training device, and the device includes:
[0017] An acquisition unit, configured to acquire any sample image in a sample set; wherein the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs;
[0018] A processing unit, configured to determine the scene probability vector corresponding to the sample image and the sample feature of the sample image through an original scene recognition model; wherein the scene probability vector includes the probability values of the sample image belonging to each scene category respectively;
[0019] A training unit, configured to train the original scene recognition model based on the scene probability vector, the scene label, the sample feature, the class center feature corresponding to the first scene category, and the class center feature corresponding to the second scene category to obtain a trained scene recognition model; wherein the second scene category is the scene category other than the first scene category among each scene category.
[0020] The present application provides a scene recognition device, the device comprising:
[0021] A first processing module, configured to determine image features of an image to be recognized through a pre-trained scene recognition model;
[0022] A second processing module, configured to determine similarities between the image features and target class center features of each scene class;
[0023] A third processing module, configured to determine whether each scene class includes the scene class to which the image to be recognized belongs according to each similarity and a similarity threshold; if it is determined that each scene class includes the scene class to which the image to be recognized belongs, determine the scene class to which the image to be recognized belongs through the scene recognition model; if it is determined that each scene class does not include the scene class to which the image to be recognized belongs, do not continue to recognize the scene class to which the image to be recognized belongs.
[0024] The present application provides an electronic device, the electronic device comprising a processor, and the processor is configured to implement the steps of the above-mentioned scene recognition model training method or the steps of the above-mentioned scene recognition method when executing a computer program stored in a memory.
[0025] The present application provides a computer-readable storage medium, which stores a computer program, and the computer program is configured to implement the steps of the above-mentioned scene recognition model training method or the steps of the above-mentioned scene recognition method when executed by a processor.
[0026] During the process of training an original scene recognition model based on sample images in a sample set, through the original scene recognition model, a scene probability vector corresponding to an input sample image and sample features of the sample image can be obtained, so that subsequently, based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene class, and the sample features and the class center features corresponding to the second scene class, the original scene recognition model can be trained to obtain a trained scene recognition model. The trained scene recognition model can make the image features within the same scene class converge towards the class center features of that scene class and move away from the class center features of other scene classes. Further, in combination with the feature level of the image, it can be determined whether the scene class of the image can be recognized and, in the case where the scene class of the image can be recognized, the scene class to which the image belongs. This not only realizes the accurate recognition of scene class images included in a closed image set but also can process scene class images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0028] Figure 1 Schematic diagram of the training process of the scene recognition model provided by some embodiments of the present application;
[0029] Figure 2 Schematic diagram of the specific training process of the scene recognition model provided by some embodiments of the present application;
[0030] Figure 3 Schematic diagram of the structure of an original scene recognition model provided by some embodiments of the present application;
[0031] Figure 4 Schematic diagram of the scene recognition process provided by some embodiments of the present application;
[0032] Figure 5 Schematic diagram of the specific scene recognition process provided by some embodiments of the present application;
[0033] Figure 6 Schematic diagram of the structure of a scene recognition model training device provided by some embodiments of the present application;
[0034] Figure 7 Schematic diagram of the structure of a scene recognition device provided by some embodiments of the present application;
[0035] Figure 8 Schematic diagram of the structure of an electronic device provided by some embodiments of the present application;
[0036] Figure 9 Schematic diagram of the structure of an electronic device provided by some embodiments of the present application. Detailed implementation manners
[0037] In order to make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the present application in detail in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0038] How to enable a scene recognition system to accurately process images of scene categories that do not belong to a closed image set is essentially an open-set recognition problem. The scene recognition system needs to be able to discover and learn the scene categories to which the images of unknown scene categories belong. In summary, the open-set recognition problem is an important and challenging problem in the pattern recognition and multimedia communities.
[0039] Therefore, in order to enable the scene recognition system to accurately process images of scene categories that do not belong to the closed image set, the present application provides a method, device, equipment, and medium for training a model and scene recognition.
[0040] Embodiment 1:
[0041] Figure 1 It is a schematic diagram of the training process of the scene recognition model provided in some embodiments of the present application. The process includes:
[0042] S101: Obtain any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs.
[0043] The method for training the scene recognition model provided by the present application is applied to an electronic device, which can be an intelligent device such as a mobile terminal, or a server such as a home brain. Of course, the electronic device can also be a display device such as a television.
[0044] In order to obtain an accurate scene recognition model, it is necessary to train the original scene recognition model according to each sample image in the pre-obtained sample set. Among them, any sample image in the sample set is obtained in the following ways: determining the collected original image as the sample image; and / or, after adjusting the pixel values of the pixel points in the collected original image, determining the adjusted image as the sample image.
[0045] It should be noted that, for the convenience of training the scene recognition model, any sample image in the sample set corresponds to a scene label, and any scene label is used to identify the scene category to which the sample image belongs (for the convenience of description, it is denoted as the first scene category). For example, the scene category is a live broadcast scene, a game scene, a food live broadcast scene, etc.
[0046] As a possible implementation manner, if the sample set contains a sufficient number of sample images, that is, it contains a large number of original images collected in different environments, the original scene recognition model can be trained according to the sample images in the sample set.
[0047] As another possible implementation, if it is necessary to ensure the diversity of sample images to improve the accuracy of the scene recognition model, the pixel values of the pixel points in the original image can be adjusted. For example, the original image can be blurred, sharpened, contrast-adjusted, etc., to obtain a large number of adjusted images, and the adjusted images are determined as sample images to train the original scene recognition model.
[0048] According to statistics, taking an electronic device as a display device, such as a television, in the working scene of the display device, the relatively common image quality problems in the obtained images include: blurring, overexposure, underexposure, too low contrast, noise in the picture, etc. For example, in a live broadcast scene, the obtained image may have problems such as overexposure. To ensure the diversity of sample images and improve the accuracy of the scene recognition model, the image quality of the collected original images can be adjusted in advance for the possible image quality problems in the images obtained in the working scene of the display device. The pixel values of the pixel points in the collected original image can be adjusted by at least one of the following methods:
[0049] Method 1: Adjust the pixel values of the pixel points in the original image through a preset convolution kernel;
[0050] Method 2: Adjust the contrast of the pixel values of the pixel points in the original image;
[0051] Method 3: Adjust the brightness of the pixel values of the pixel points in the original image;
[0052] Method 4: Add noise to the pixel values of the pixel points in the original image.
[0053] For example, if it is desired to add noise to the original image to obtain adjusted images with different noises, the pixel values of the pixel points in the original image can be added with noise, that is, randomly add noise to the original image. Among them, in the process of adding noise to the original image, the types of noise used should also be as many as possible, such as white noise, salt-and-pepper noise, Gaussian noise, etc., so that the sample images in the sample set are more diverse, thereby improving the accuracy and robustness of the scene recognition model.
[0054] It should be noted that the process of processing the pixel values of the pixel points in the original image belongs to the prior art and will not be elaborated here specifically.
[0055] By the above method, the sample images can be obtained, doubling the number of sample images in the sample set, enabling a large number of sample images to be quickly obtained, and reducing the difficulty, cost, and resources consumed in obtaining the sample images. Subsequently, based on more sample images, the original scene recognition model can be trained, improving the accuracy and robustness of the scene recognition model.
[0056] As another possible implementation, the collected original images and the adjusted images obtained by adjusting the pixel values of the pixel points in the collected original images can both be determined as sample images. Based on the original images and the adjusted images in the sample set, the original scene recognition model is trained together.
[0057] S102: Through the original scene recognition model, determine the scene probability vector corresponding to the sample image and the sample feature of the sample image; wherein, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively.
[0058] After obtaining the sample set for training the original scene recognition model based on the above embodiments, the original scene recognition image can be trained based on each sample image in the sample set.
[0059] In the specific implementation process, any sample image is input into the original scene recognition model. Through the original scene recognition model, the scene probability vector corresponding to the above sample image and the image feature of the sample image (for convenience of description, denoted as sample feature) can be obtained. Among them, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively, and each of these scene categories is determined by the scene categories to which the sample images in the sample set belong. Any sample feature represents a higher-dimensional and more abstract image feature extracted from the sample image.
[0060] Among them, the original scene recognition model can be a decision tree, logistic regression (LR), naive bayes (NB) classification algorithm, random forest (RF) algorithm, support vector machines (SVM) classification algorithm, histogram of oriented gradients (HOG), deep learning algorithm, etc. Among them, the deep learning algorithm can include neural network, deep neural network, Convolutional Neuron Network (CNN), etc.
[0061] In a possible implementation manner, in order to perform scene recognition through a scene recognition model, the original scene recognition model includes a feature extraction layer, a feature output layer, and a classification output layer. The feature extraction layer outputs to the feature output layer, and the feature output layer is connected to the classification output layer. When a sample image is input into the original scene model, the sample features of the input sample image can be obtained through the feature extraction layer in the original scene recognition model. Then, through the feature output layer in the original scene recognition model, the sample features can be output. Through the classification output layer in the original scene recognition model, based on the sample features, the scene probability vector corresponding to the sample image can be obtained and output.
[0062] S103: Based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, train the original scene recognition model to obtain a trained scene recognition model; wherein, the second scene category is the scene category other than the first scene category in each scene category.
[0063] Since any sample image in the sample set corresponds to a scene label, that is, it identifies the scene category to which the sample image actually belongs. Therefore, in this application, after determining the scene probability vector corresponding to the sample image and the sample features of the sample image, based on the scene probability vector, the corresponding scene label, and the sample features, the original scene recognition model can be trained by using the scene recognition model training method provided in this application.
[0064] If the scene classification to which an image belongs is a scene classification that can be recognized by a pre-trained scene recognition model, generally, the measurement distance between the image features of the image and the image features of the sample images belonging to this scene type in the sample set will be larger, and the measurement distance from the image features of the sample images not belonging to this scene type in the sample set will be smaller. Based on this, when recognizing the scene classification of an image, the measurement distance between the image features of the image and the image features of each sample image in the sample set can be determined, and based on the obtained measurement distance, it can be determined whether the scene classification to which the image belongs is a scene classification that can be recognized by a pre-trained scene recognition model.
[0065] Among them, the measurement distance can be obtained through methods such as Euclidean distance, cosine similarity, and KL divergence function.
[0066] Further, since the sample set will contain a large number of sample images, if the metric distances between the image features of a certain image and the image features of each sample image in the sample set are determined, a large amount of computing resources will be consumed, reducing the efficiency of the scene recognition system in determining the scene category to which the image belongs. Based on this, the class center features of the scene categories to which each sample image in the sample set belongs can be obtained, so that the general features of the images of the scene category can be represented by the class center features. Subsequently, when recognizing the scene classification of an image, the metric distances between the image features of the image and the class center features of the scene categories to which each sample image in the sample set belongs can be determined. According to the obtained metric distances, it can be determined whether the scene classification to which the image belongs is a scene classification that the pre-trained scene recognition model can recognize.
[0067] It should be noted that the dimension of the sample features is the same as the dimension of the class center features.
[0068] In a possible implementation manner, in order to accurately obtain the class center features of each scene category, the class center features of each scene category can be obtained in the following manner:
[0069] Method 1: In order to incorporate the process of obtaining the class center features of each scene category into the model training process, for each iterative training of the original scene recognition model, through the scene recognition model of the current iteration, the sample features of one or more sample images of each scene category in the sample set can be obtained. Then, according to each sample feature, the candidate class center features corresponding to each scene category are determined. Based on each candidate class center feature, the class center features of each scene category in the next iterative training are determined.
[0070] In a possible implementation manner, for each scene category, the sample images correctly recognized by the scene recognition model of the current iteration (for the convenience of description, denoted as target sample images) can be determined among the various sample images of the scene category. The being correctly recognized by the scene recognition model of the current iteration can be understood as that the scene category of the sample image determined by the scene recognition model of the current iteration is the same as the first scene classification of the sample image. Then, according to the sample features of the target sample images and the weight values of the target sample images, a weighted average vector is determined, and based on the weighted average vector, the candidate class center features corresponding to the scene category are determined.
[0071] Among them, the weight values of the target sample images can be pre-configured in a manual configuration manner. For example, the weight value of each target sample image is set to 1, or the probability value that the target sample image belongs to the scene category can be obtained through the scene recognition model of the current iteration, and this probability value is determined as the weight value of the target sample image.
[0072] For example, if the weight value of the target sample image is pre-configured manually, then according to the sample features of the target sample image and the weight value of the target sample image, the weighted average vector can be expressed by the following formula:
[0073]
[0074] where C i is the class center feature of scene category i, is the sample feature of the j-th target sample image correctly recognized as scene classification i, is the number of target sample images correctly recognized as scene classification i, and the weight value of the target sample image is 1.
[0075] Again, for example, if the probability value that the target sample image belongs to the scene category is obtained through the scene recognition model of the current iteration, and this probability value is determined as the weight value of the target sample image, then according to the sample features of the target sample image and the weight value of the target sample image, the weighted average vector can be expressed by the following formula:
[0076]
[0077] where C i is the class center feature of scene category i, is the sample feature of the j-th target sample image correctly recognized as scene classification i, is the number of target sample images correctly recognized as scene classification i, is the probability value that the j-th target sample image belongs to the scene category i obtained through the scene recognition model of the current iteration. The higher this weight value, the more accurate the recognition result of the scene recognition model of the current iteration, and the greater the contribution of the sample features of the more accurately recognized sample images to the class center.
[0078] In another possible implementation, for each scene category, it is also possible to determine the target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of this scene category; based on a preset target algorithm, obtain the target features in the sample features of the target sample image, and based on this target feature, determine the candidate class center features corresponding to the scene categories respectively. Among them, the target feature is the principal component feature, or the normalized feature.
[0079] Among them, the process of specifically obtaining the principal component feature, or the normalized feature, in the pattern features of the image through the preset target algorithm belongs to the prior art and will not be elaborated here.
[0080] After obtaining the candidate class center features for each scene category based on the above embodiments, the class center features for each scene category can be determined according to the candidate class center features for each scene category. Specifically, the process of determining the class center features for each scene category according to the candidate class center features for each scene category mainly includes the following two cases:
[0081] Case 1: Since the class center features for each scene category are randomly initialized before training the original scene recognition model, the class center features for each scene category in the current iteration are inaccurate during the first iteration training of the original scene model. Therefore, in this application, if it is determined that the current iteration is the first iteration, each candidate class center feature can be directly determined as the class center feature for each scene category in the next iteration training, that is, the class center features for each scene category in the current iteration are updated according to each candidate class center feature, thereby improving the accuracy of the obtained class center features for each scene category.
[0082] Among them, the dimension of the randomly initialized class center feature is the same as the dimension of the sample feature.
[0083] Case 2: In order to make the class center features for each scene category determined in each iteration more accurate and the changes more stable, in this application, a weight vector is pre-configured, and this weight vector is used to adjust the amplitude of updating the class center feature each time. After determining the candidate class center features for each scene category based on the above embodiments, if it is determined that the current iteration is not the first iteration, for each scene category, the difference vector between the candidate class center feature corresponding to this scene category and the class center feature corresponding to this scene category determined in the current iteration is determined. Then, according to this difference vector and the pre-configured weight vector, this difference vector is adjusted. According to the adjusted difference vector and the class center feature corresponding to this scene category determined in the current iteration, the class center feature corresponding to this scene category in the next iteration training is determined.
[0084] In a possible implementation manner, the product vector of this difference vector and the pre-configured weight vector can be obtained, and this product vector is determined as the adjusted difference vector.
[0085] In a possible implementation manner, the sum vector can be determined according to the adjusted difference vector and the class center feature corresponding to this scene category determined currently, and this sum vector is determined as the class center feature corresponding to this scene category in the next iteration training.
[0086] For example, based on the above Case 1 and Case 2, the process of determining the class center features for each scene category according to the candidate class center features for each scene category can be represented by the following formula:
[0087]
[0088] Among them, is the class center feature corresponding to the scene category i in the next iterative training, is the candidate class center vector corresponding to the scene category i, is the class center feature corresponding to the scene category i determined in the current iteration, and W is a pre-configured weight vector.
[0089] Method 2: It is also possible to obtain the sample features of each sample image in the sample set through a pre-trained feature extraction model. It can be understood that this feature extraction model is also a feature extraction algorithm. Then, a clustering algorithm, such as a fuzzy clustering algorithm, K-means clustering, maximum-minimum distance clustering algorithm, etc., is used to cluster each sample feature, so as to obtain the clusters corresponding to each scene category. Among them, the clusters corresponding to any scene category include the sample features of this scene category. Then, according to the sample features included in the clusters corresponding to each scene category respectively, the class center features in the clusters are determined, that is, the class center features of each scene category are determined.
[0090] Among them, any sample feature included in the cluster can be determined as the class center feature, or the average vector of the various sample features included in the cluster can be determined as the class center feature. In the specific implementation process, it can be flexibly used according to actual needs, and no specific description is given here.
[0091] It should be noted that the process of training the feature extraction model and the process of clustering the sample features according to the clustering algorithm belong to the prior art, and no specific description is given here.
[0092] In order to facilitate the trained scene recognition model to determine the scene to which the image belongs from the feature level of the image, in this application, during the process of training the original scene recognition model, the metric distance between the sample features of the sample image and the class center features of the scene category to which each sample image belongs can be considered. Then, based on this metric distance, the scene probability vector, and the scene label, the original scene recognition model is trained. It can be understood that the original scene recognition model is trained based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category. Among them, the second scene category is the scene category other than the first scene category among the scene categories to which each sample image in the sample set belongs.
[0093] In a possible implementation manner, when determining the metric distance between the sample features and the class center features of the scene category to which each sample image belongs, it can be determined through the following Euclidean distance formula:
[0094]
[0095] where d(x, y i ) represents the metric distance between the sample feature x and the class center feature of the i-th scene category.
[0096] In another possible implementation, since the Euclidean distance represents the degree of closeness between two vectors in terms of absolute distance, and the cosine similarity represents the degree of closeness between two vectors in terms of direction. Therefore, when determining the metric distance between the sample feature and the class center feature of the scene category to which each sample image belongs, it can be determined by the following formula:
[0097]
[0098] where d(x, y i ) represents the metric distance between the sample feature x and the class center feature of the i-th scene category, cos_sim(x, y i ) represents the cosine similarity between the sample feature x and the class center feature of the i-th scene category, α1 represents the weight value corresponding to the Euclidean distance, and α2 represents the weight value corresponding to the cosine similarity.
[0099] In one possible implementation, the loss value (for convenience of description, denoted as the first loss value) can be determined based on the scene probability vector and the scene label; the loss value (for convenience of description, denoted as the second loss value) can be determined based on the sample feature and the class center feature corresponding to the first scene category; the loss value (for convenience of description, denoted as the third loss value) can be determined based on the sample feature and the class center feature corresponding to the second scene category. Then, according to the first loss value and its corresponding first weight value, the second loss value and its corresponding second weight value, and the third loss value and its corresponding third weight value, the comprehensive loss value is determined. Based on the comprehensive loss value, the original scene recognition model is trained to update the parameter values of the parameters in the original scene recognition model, so as to obtain the trained scene recognition model.
[0100] In specific implementation, when training the original scene recognition model according to the comprehensive loss value, the gradient descent algorithm can be used to perform backpropagation on the gradients of the parameters in the original scene recognition model, so as to realize the training of the original scene recognition model.
[0101] It can be understood that the second loss value can be determined by the metric distance between the sample feature and the class center feature corresponding to the first scene category, and the third loss value can also be determined by the metric distance between the sample feature and the class center feature corresponding to the second scene category.
[0102] For example, according to the first loss value and its corresponding first weight value, the second loss value and its corresponding second weight value, and the third loss value and its corresponding third weight value, a comprehensive loss value is determined, which can be determined by the following formula:
[0103]
[0104] Wherein, is the first loss value determined according to the scenario probability vector y and the scenario label , d(x i , C i ) is the metric distance between the sample feature x i and the class center feature C i corresponding to the first scenario category, d(x i , C cls!=i ) is the metric distance between the sample feature x i and the class center feature C cls!=i corresponding to the second scenario category, ω1 is the first weight value, ω2 is the second weight value, and ω3 is the third weight value.
[0105] In the actual application process, the smaller the metric distance between the image features of the same scenario category image, and the larger the metric distance between the image features of different scenario category images. Therefore, when setting the first weight value, the second weight value, and the third weight value, the second loss value can be negative, and the first loss value and the third loss value can be positive, so that when optimizing the comprehensive loss value, the optimization direction of the scenario recognition model is towards minimizing the first loss value, minimizing the second loss value, and maximizing the third loss value, in order to increase the metric distance between different scenario categories and reduce the metric distance between the sample features of the same scenario category. From the perspective of the feature space, it makes the distribution of the sample features of different scenario categories relatively dispersed, but the distribution of the sample features of the same scenario category relatively concentrated.
[0106] By training the scenario recognition model with the above comprehensive loss value, it can not only make the similarity between the sample features of different scenario categories smaller in the feature space, but also make the similarity between the sample features of the same scenario category larger in the feature space, which is beneficial to improving the accuracy of determining whether the scenario category of a certain image is recognizable by the scenario recognition model.
[0107] Since the sample set contains a large number of sample images, the above operations are performed on each sample image. When the preset convergence condition is met, the training of the scenario recognition model is completed.
[0108] Among them, meeting the preset convergence condition can be that the sum of each comprehensive loss value obtained according to the current iterative training is less than the set loss value threshold, the number of iterations for training the model reaches the set maximum number of iterations, etc. In specific implementation, it can be flexibly set and will not be specifically limited here.
[0109] In a possible implementation manner, in order to determine the accuracy of the trained scene recognition model, before the scene recognition model is launched and released, the scene recognition model can be tested to determine whether the scene recognition model can accurately process images that do not belong to the scene category images included in the sample set, and the recognition accuracy of the scene recognition model for the images that can be recognized.
[0110] In the specific implementation process, a test set for testing the trained scene recognition model is obtained. The test set contains test sample images to verify the reliability of the above-mentioned trained scene recognition model based on the test sample images. When obtaining the test sample images included in the test set, it can be obtained by re-collecting the test sample images included in the test set, and / or, it can also be obtained by dividing the sample images in the sample set into training sample images and test sample images. It should be noted that the specific process of collecting the test sample images included in the test set is similar to the process of collecting the sample images included in the sample set above, and the repeated parts will not be elaborated.
[0111] In order to ensure that the ability of the scene recognition model to accurately process images that do not belong to the scene category images included in the sample set can be tested, in each obtained test sample image, there should be at least one scene category to which the test sample image belongs and is different from all the scene categories included in the sample set.
[0112] Among them, each test sample image corresponds to a scene label and a processing label. The scene label is used to identify the scene category to which the test sample image belongs (for the convenience of description, denoted as the third scene category), and the processing label is used to identify whether the scene categories included in the sample set contain the third scene category.
[0113] For each test sample image in the test set, the test sample image is input into the scene recognition model. Through the scene recognition model, the image feature of the test sample image is obtained (for the convenience of description, denoted as the test sample feature). Then, the similarity between the test sample feature and the target class center feature of each scene category is determined. Among them, the target class center feature can be the class center feature of each scene category included in the sample set during the last iterative training of the original scene recognition model. Then, according to the similarity threshold and the obtained similarity, it is determined whether each scene category contains the scene category to which the test sample image belongs.
[0114] In a possible implementation, the similarity threshold can be configured manually, or for each target class center feature, a reference similarity between the target class center feature and other target class center features can be determined. Then, based on the reference similarity corresponding to each target class center feature, the similarity threshold is determined.
[0115] In a possible implementation, if the reference similarity is determined based on a metric distance such as the Euclidean distance, the similarity threshold can be determined according to the minimum value among the respective reference similarities.
[0116] In a possible implementation, if the reference similarity is determined based on a metric distance such as the cosine similarity, the similarity threshold can be determined according to the maximum value among the respective reference similarities.
[0117] In a possible implementation, if the similarity is determined based on a metric distance such as the Euclidean distance, the smaller the similarity between two image features, the higher the similarity between the two image features, and the more likely the two image features belong to the same scene category; the greater the similarity between two image features, the lower the similarity between the two image features, and the less likely the two image features belong to the same scene category. Therefore, when determining whether each scene category contains the scene category to which the test sample image belongs based on each similarity and the similarity threshold, if there is any similarity less than the similarity threshold, it indicates that the image features of the test sample image and the target class center feature corresponding to this similarity are very likely to belong to the same scene category, and it is determined that each scene category included in the sample set contains the scene category to which the test sample image belongs; if each similarity is not less than the similarity threshold, it indicates that the image features of the test sample image and each target class center feature come from different scene categories, and it is determined that each scene category included in the sample set does not contain the scene category to which the test sample image belongs.
[0118] In a possible implementation, if the similarity is determined according to a distance metric such as cosine similarity, the smaller the similarity between two image features, the lower the similarity between the two image features, and the less likely the two image features belong to the same scene category; the smaller the similarity between two image features, the higher the similarity between the two image features, and the more likely the two image features belong to the same scene category. Therefore, when determining whether each scene category contains the scene category to which the test sample image belongs according to each similarity and the similarity threshold, if any similarity is greater than the similarity threshold, it indicates that the image features of the test sample image and the target class center features corresponding to the similarity are very likely to belong to the same scene category, and it is determined that each scene category included in the sample set contains the scene category to which the test sample image belongs; if each similarity is not greater than the similarity threshold, it indicates that the image features of the test sample image and each target class center feature come from different scene categories, and it is determined that each scene category included in the sample set does not contain the scene category to which the test sample image belongs.
[0119] Specifically, if it is determined that each scene category contains the scene category to which the test sample image belongs, it means that the scene category to which the test sample image belongs can be accurately determined by the scene recognition model, that is, the scene category to which the test sample image belongs is known, and then the scene category to which the test sample image belongs is determined through the scene recognition model; if it is determined that each scene category does not contain the scene category to which the test sample image belongs, it means that the scene category to which the test sample image belongs cannot be accurately determined by the scene recognition model, that is, the scene category to which the test sample image belongs is unknown, and then the scene category to which the image belongs is not recognized further.
[0120] Since the test set contains a large number of test sample images, the above operations are performed on each test sample image. Based on the processing results of each test sample image (including the result of whether the scene recognition model recognizes the scene category to which the test sample image belongs, and the scene probability vector of the test sample image obtained when the scene recognition model recognizes the scene category to which the test sample image belongs), the processing label of each test sample image, and the scene label of each test sample image, corresponding calculations are performed to determine various evaluation indicators of the scene recognition model, such as accuracy, error rate, precision, etc. If it is determined that the various evaluation indicators of the scene recognition model meet the preset release requirements, the scene recognition model can be released online. If it is determined that the various evaluation indicators of the scene recognition model do not meet the preset release requirements, the scene recognition model can be further trained based on the sample images in the sample set.
[0121] Since, when training the scene recognition model in this application, the sample features of the sample images of each scene category in the sample set are learned simultaneously, so as to obtain the class center features of each scene category, when using or testing the trained scene recognition model subsequently, the image features of the input image can be obtained through the scene recognition model, and then the metric distance between the image features and the class center features of each scene category can be determined. If it is determined that the metric distance between the image features and the class center features of any scene category is not close, it means that it is determined that none of the known scene categories contains the scene category to which the image belongs, that is, the scene category to which the image belongs is also unknown. Then, subsequent recognition of the scene type of the image is not performed, thereby realizing that the image features extracted by the scene recognition model are discriminative, so as to help the scene recognition model determine whether the scene category to which the image belongs can be recognized by the scene recognition image, and avoid misidentifying the scene category to which the image belongs and affecting the accuracy of the following algorithms.
[0122] Since, in the process of training the original scene recognition model based on the sample images in the sample set, through the original scene recognition model, the scene probability vector corresponding to the input sample image and the sample features of the sample image can be obtained, so that subsequently, based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, the original scene recognition model can be trained to obtain a trained scene recognition model, so that the trained scene recognition model can make the image features within the same scene category approach the class center features of that scene category and at the same time be far from the class center features of other scene categories. Further combining the feature level of the image, it is determined whether the scene category of the image can be recognized and, in the case where the scene category of the image can be recognized, the scene category to which the image belongs, which not only realizes accurately recognizing the scene category images included in the closed image set, but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0123] Embodiment 2:
[0124] Taking the execution entity as a display device as an example, the scene recognition model training method provided in this application will be described in detail through specific embodiments below. Figure 2 The following is a schematic diagram of the specific scene recognition model training process provided in some embodiments of this application. The process includes:
[0125] S201: Construct an original scene recognition model.
[0126] S202: Randomly construct the class center features of each scene category.
[0127] S203: Obtain any sample image in the sample set.
[0128] Among them, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs.
[0129] S204: Through the original scene recognition model, determine the scene probability vector corresponding to the sample image and the sample features of the sample image.
[0130] Among them, the scene probability vector includes the probability values that the sample image belongs to each scene category respectively.
[0131] The following combines Figure 3 to introduce in detail the process of determining the scene probability vector corresponding to the sample image and the sample features of the sample image through the original scene recognition model. Figure 3 FIG. 3 is a schematic structural diagram of an original scene recognition model provided by some embodiments of the present application. After any sample image is input into the original scene recognition model, through the feature extraction layer in the original scene recognition model, the sample features of the input sample image can be obtained. Then, through the feature output layer in the original scene recognition model, the sample features can be output. Through the classification output layer in the original scene recognition model, based on the sample features, the scene probability vector corresponding to the sample image can be obtained and output.
[0132] Since the sample set contains a large number of sample images, the above operations S203-S204 are performed on each sample image.
[0133] S205: Update the class center features of each scene category in the current iteration.
[0134] Among them, if the current iteration is the first iteration, according to each sample feature obtained in the current iteration, determine the candidate class center features corresponding to each scene category; determine each candidate class center feature as the class center feature of each scene category in the next iteration training.
[0135] If the current iteration is the first iteration, according to each sample feature obtained in the current iteration, determine the candidate class center features corresponding to each scene category; for each scene category, determine the difference vector between the candidate class center feature corresponding to the scene category and the class center feature corresponding to the scene category determined in the current iteration; according to the difference vector, the pre-configured weight vector, and the class center feature corresponding to the scene category determined in the current iteration, determine the class center feature corresponding to the scene category in the next iteration training.
[0136] S206: For each sample image, determine the comprehensive loss value according to the scene probability vector and scene label of the sample image, the sample features of the sample image, the class center features corresponding to the first scene category to which the sample image belongs, and the class center features corresponding to the second scene category of the sample image.
[0137] S207: Determine whether the sum of each comprehensive loss value is less than a preset loss value threshold. If it is less, execute S208; otherwise, execute S209.
[0138] S208: Obtain the trained scene recognition model and save it.
[0139] S209: Adjust the parameter values of the parameters of the original scene recognition model, and execute S203.
[0140] Embodiment 3:
[0141] The present application also provides a scene recognition method. Figure 4 As shown in the schematic diagram of the scene recognition process provided by some embodiments of the present application, the process includes:
[0142] S401: Determine the image features of the image to be recognized through a pre-trained scene recognition model.
[0143] S402: Determine the similarity between the image features and the target class center features of each scene category.
[0144] S403: According to each similarity and the similarity threshold, determine whether each scene category contains the scene category to which the image to be recognized belongs.
[0145] S404: If it is determined that each scene category contains the scene category to which the image to be recognized belongs, then determine the scene category to which the image to be recognized belongs through the scene recognition model.
[0146] S405: If it is determined that each scene category does not contain the scene category to which the image to be recognized belongs, then do not continue to recognize the scene category to which the image to be recognized belongs.
[0147] The scene recognition method provided by the present application is applied to an electronic device, which can be a smart device such as a mobile terminal or a server. Of course, the electronic device can also be a display device such as a television.
[0148] In a possible application scenario, taking an electronic device as a TV set as an example, for the scenario of real-time scene classification of the video picture played on the TV, in order to better analyze the video picture, the TV program can first perform scene recognition on the images included in the video, and then, according to the scene category to which the video belongs, combine with downstream algorithms to process the video picture. For example, optimize the picture quality of the video picture, etc.
[0149] In a possible implementation manner, after determining that the electronic device receives a processing request for scene recognition of an image in a certain video, the image is determined as an image to be recognized, and based on the image to be recognized, the scene recognition method provided in this application is used for corresponding processing.
[0150] Among them, the electronic device for performing scene recognition receives a processing request for scene recognition of an image in a certain video, which mainly includes at least one of the following situations:
[0151] Situation 1: When scene recognition is required, the user can input a service processing request for scene recognition to the intelligent device. After the intelligent device receives the service processing request, it can send a processing request for scene recognition of the image in the video to the electronic device for performing scene recognition.
[0152] Situation 2: When the intelligent device determines that a video is recorded, it generates a processing request for scene recognition of the image in the recorded video and sends it to the electronic device for performing scene recognition.
[0153] Situation 3: When the user needs to perform scene recognition on a certain specific video, the user can input a service processing request for scene recognition of the video to the intelligent device. After the intelligent device receives the service processing request, it can send a processing request for scene recognition of the image in the video to the electronic device for performing scene recognition.
[0154] It should be noted that the electronic device for performing scene recognition can be the same as or different from the intelligent device.
[0155] As a possible implementation manner, scene recognition conditions can also be preset. For example, when receiving a video sent by a display device, perform scene recognition on the images in the video; when receiving a preset number of frames of images in a certain video sent by a display device, perform scene recognition on the preset number of frames of images; perform scene recognition on the images in the currently acquired video at a preset period, etc. When the electronic device determines that the current time meets the preset scene recognition conditions, perform scene recognition on the images in a certain video.
[0156] In the present application, when obtaining images from a video, some video frames can be extracted from the video according to a preset frame extraction strategy, and the extracted video frames can be converted into corresponding images. Alternatively, all video frames in the video can be converted into corresponding images in a full-frame extraction manner.
[0157] To accurately determine the scene to which an image belongs, a scene recognition model is pre-trained. When an electronic device for scene recognition needs to perform scene recognition on a to-be-recognized image, the to-be-recognized image can be input into the pre-trained scene recognition model, so as to determine the scene category to which the input to-be-recognized image belongs through the pre-trained scene recognition model.
[0158] Among them, the process of training the scene recognition model has been described in the above embodiments, and repeated parts will not be elaborated. For the above training method of the scene recognition model, during the process of training the original scene recognition model based on the sample images in the sample set, through the original scene recognition model, a scene probability vector corresponding to the input sample image and the sample features of the sample image can be obtained, so that subsequently, based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, the original scene recognition model can be trained to obtain a trained scene recognition model. The trained scene recognition model can make the image features of images within the same scene category approach the class center features of that scene category and move away from the class center features of other scene categories. Further combining the feature level of the image, it can be determined whether the scene category of the image can be recognized, and in the case where the scene category of the image can be recognized, the scene category to which the image belongs. This not only realizes the accurate recognition of scene category images included in a closed image set but also can process scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0159] It should be noted that the electronic device for training the scene recognition model and the electronic device for scene recognition can be the same or different.
[0160] Since the scene category to which the image to be recognized belongs is unpredictable and has a certain diversity, and the actual scene category to which the image to be recognized belongs may not be the scene category included in the sample set used to train the scene recognition model. However, if the scene recognition model is directly used to recognize the scene category to which the image belongs, incorrect results will be obtained, which will in turn affect the processing of downstream algorithms. Moreover, if the scene classification to which an image belongs is a scene classification that can be recognized by a pre-trained scene recognition model, then the image features of this image generally have a greater metric distance from the image features of the sample images belonging to this scene type in the sample set and a smaller metric distance from the image features of the sample images not belonging to this scene type in the sample set. Therefore, in order to ensure that the scene recognition model can accurately recognize the scene category images included in the closed image set, in this application, the image features of the image to be recognized can be obtained through the pre-trained scene recognition model, and for the scene categories included in the sample set used to train the scene recognition model, the target class center features of this scene category are pre-obtained. After the image to be recognized is input into the pre-trained scene recognition model based on the above embodiment, the image features of the image to be recognized can be obtained through the pre-trained scene recognition model. Then, the similarities between the determined image features and the target class center features of each scene category are determined. According to each similarity, it is determined whether the scene category to which the image to be recognized belongs is any scene category included in the sample set.
[0161] Among them, the target class center feature can be the class center feature of each scene category included in the sample set during the last iteration training of the original scene recognition model.
[0162] In the specific implementation process, the image features of the input image to be recognized can be obtained through the feature extraction layer in the scene recognition model. Then, through the feature output layer in the scene recognition model, the image features can be output. Then, the similarities between the determined image features and the target class center features of each scene category are determined. According to each similarity, it is determined whether the scene category to which the image to be recognized belongs is any scene category included in the sample set.
[0163] It should be noted that during the last iteration training of the original scene recognition model, the method for obtaining the class center feature of each scene category included in the sample set can refer to the obtaining methods in Case 1 and Case 2, and the repeated parts will not be elaborated.
[0164] In a possible implementation manner, the similarities between the image features and the target class center features of each scene category can be determined according to the metric distances between the image features and the target class center features of each scene category. Among them, the metric distance can be obtained through methods such as Euclidean distance, cosine similarity, and KL divergence function.
[0165] In a possible implementation, when determining the metric distance between the image feature and the target class center feature of the scene category to which each sample image belongs, it can be determined by the following Euclidean distance formula:
[0166]
[0167] where d(x, y i ) represents the metric distance between the image feature x and the target class center feature of the i-th scene category.
[0168] In another possible implementation, since the Euclidean distance represents the degree of closeness between two vectors in terms of absolute distance, and the cosine similarity represents the degree of closeness between two vectors in terms of direction. Therefore, when determining the metric distance between the image feature and the target class center feature of the scene category to which each sample image belongs, it can be determined by the following formula:
[0169]
[0170] where d(x, y i ) represents the metric distance between the image feature x and the target class center feature of the i-th scene category, cos_sim(x, y i ) represents the cosine similarity between the image feature x and the target class center feature of the i-th scene category, α1 represents the weight value corresponding to the Euclidean distance, and α2 represents the weight value corresponding to the cosine similarity.
[0171] In a possible implementation, the similarity threshold can be configured manually, or for each target class center feature, the reference similarity between the target class center feature and other target class center features can be determined. Then, according to the reference similarity corresponding to each target class center feature, the similarity threshold can be determined.
[0172] In a possible implementation, if the reference similarity is determined based on metric distances such as the Euclidean distance, the similarity threshold can be determined according to the minimum value among the various reference similarities.
[0173] In a possible implementation, if the reference similarity is determined based on metric distances such as the cosine similarity, the similarity threshold can be determined according to the maximum value among the various reference similarities.
[0174] In a possible implementation, if the similarity is determined based on a metric distance such as the Euclidean distance, the smaller the similarity between two image features, the higher the similarity between the two image features, and the more likely the two image features belong to the same scene category; the greater the similarity between two image features, the lower the similarity between the two image features, and the less likely the two image features belong to the same scene category. Therefore, when determining whether each scene category contains the scene category to which the image to be recognized belongs based on each of the similarities and a similarity threshold, if any similarity is less than the similarity threshold, it indicates that the image features of the image to be recognized and the target class center features corresponding to the similarity are very likely to belong to the same scene category, and it is determined that each scene category included in the sample set contains the scene category to which the image to be recognized belongs; if each similarity is not less than the similarity threshold, it indicates that the image features of the image to be recognized and each target class center feature come from different scene categories, and it is determined that each scene category included in the sample set does not contain the scene category to which the image to be recognized belongs.
[0175] In a possible implementation, if the similarity is determined based on a metric distance such as the cosine similarity, the smaller the similarity between two image features, the lower the similarity between the two image features, and the less likely the two image features belong to the same scene category; the smaller the similarity between two image features, the higher the similarity between the two image features, and the more likely the two image features belong to the same scene category. Therefore, when determining whether each scene category contains the scene category to which the image to be recognized belongs based on each of the similarities and a similarity threshold, if any similarity is greater than the similarity threshold, it indicates that the image of the image features to be recognized and the target class center features corresponding to the similarity are very likely to belong to the same scene category, and it is determined that each scene category included in the sample set contains the scene category to which the image to be recognized belongs; if each similarity is not greater than the similarity threshold, it indicates that the image features of the image to be recognized and each target class center feature come from different scene categories, and it is determined that each scene category included in the sample set does not contain the scene category to which the image to be recognized belongs.
[0176] In the specific implementation process, if it is determined that each scene category contains the scene category to which the image to be recognized belongs, it means that the scene category to which the image to be recognized belongs can be accurately determined by the scene recognition model, that is, the scene category to which the image to be recognized belongs is known, and the scene category to which the image to be recognized belongs is determined through the pre-trained scene recognition model; if it is determined that each scene category does not contain the scene category to which the image to be recognized belongs, it means that the scene category to which the image to be recognized belongs cannot be accurately determined by the scene recognition model, that is, the scene category to which the image to be recognized belongs is unknown, and the scene category to which the image to be recognized belongs is not recognized further.
[0177] Further, if it is determined that each scene category includes the scene category to which the image to be recognized belongs, then through the classification output layer in the scene recognition model, based on the image features of the image to be recognized, the scene category to which the image to be recognized belongs can be obtained and output.
[0178] Since a scene recognition model is pre-trained, and the scene recognition model is obtained by training the original scene recognition model based on the scene probability vector of the sample image, the scene label of the sample image, the sample features of the sample image, the class center features corresponding to the first scene category of the sample image, and the sample features of the sample image and the class center features corresponding to the second scene category of the sample image, during the process of recognizing the scene category to which the image to be recognized belongs based on this scene recognition model, it is possible to make the image features within the same scene category approach the class center features of this scene category, while moving away from the class center features of other scene categories. Further combining with the feature level of the image, it is determined whether the scene category of this image can be recognized, and in the case where the scene category of this image can be recognized, the scene category to which this image belongs. This not only realizes the accurate recognition of the scene category images included in the closed image set, but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0179] Embodiment 4:
[0180] Taking the electronic device for scene recognition as a television as an example, the scene recognition method provided by the present application will be described in detail through specific embodiments. Figure 5 The following is a schematic diagram of the specific scene recognition process provided by some embodiments of the present application. The process includes:
[0181] S501: Obtain a pre-trained scene recognition model.
[0182] S502: Determine the image features of the image to be recognized through the pre-trained scene recognition model.
[0183] S503: Determine the similarity between the image features and the target class center features of each scene category respectively.
[0184] S504: If the similarity is the Euclidean distance, determine whether there is any similarity less than the similarity threshold. If so, execute S505; otherwise, execute S506.
[0185] S505: Determine the scene category to which the image to be recognized belongs through the scene recognition model.
[0186] S506: Do not continue to recognize the scene category to which the image to be recognized belongs.
[0187] Embodiment 5:
[0188] The present application provides a scene recognition model training device. Figure 6 As a schematic structural diagram of a scene recognition model training device provided in some embodiments of the present application, the device includes:
[0189] An acquisition unit 61, configured to acquire any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify a first scene category to which the sample image belongs;
[0190] A processing unit 62, configured to determine a scene probability vector corresponding to the sample image and a sample feature of the sample image through an original scene recognition model; wherein, the scene probability vector includes probability values of the sample image belonging to each scene category respectively;
[0191] A training unit 63, configured to train the original scene recognition model based on the scene probability vector, the scene label, the sample feature, a class center feature corresponding to the first scene category, the sample feature, and a class center feature corresponding to a second scene category, so as to obtain a trained scene recognition model; wherein, the second scene category is a scene category other than the first scene category among each scene category.
[0192] In some possible implementation manners, the training unit 63 is further configured to, for each iterative training of the original scene recognition model, obtain sample features of sample images of each scene category in the sample set through the scene recognition model of the current iteration; determine candidate class center features corresponding to each scene category according to each sample feature; determine class center features of each scene category in the next iterative training based on each candidate class center feature; or, obtain sample features of each sample image in the sample set through a pre-trained feature extraction model; perform clustering on the sample features of each sample image to determine class center features of each scene category.
[0193] In some possible implementation manners, the training unit 63 is specifically configured to, for each scene category, determine the target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of the scene category; determine a weighted average vector according to the sample features of the target sample images and the weight values of the target sample images, and determine a candidate class center feature corresponding to the scene category based on the weighted average vector; wherein the weight value of the target sample image is preset or determined according to the probability value that the target sample image belongs to the scene category obtained by the scene recognition model of the current iteration; or, for each scene category, determine the target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of the scene category; obtain target features in the sample features of the target sample images based on a preset target algorithm, and determine a candidate class center feature corresponding to the scene category based on the target features; wherein the target features are principal component features or normalized features.
[0194] In some possible implementation manners, the training unit 63 is specifically configured to, if the current iteration is the first iteration, determine each candidate class center feature as the class center feature of each scene category in the next iteration training; if the current iteration is not the first iteration, for each scene category, determine a difference vector between the candidate class center feature corresponding to the scene category and the class center feature corresponding to the scene category determined in the current iteration; and determine the class center feature corresponding to the scene category in the next iteration training according to the difference vector, a pre-configured weight vector, and the class center feature corresponding to the scene category determined in the current iteration.
[0195] In some possible implementation manners, the training unit 63 is specifically configured to determine a first loss value, a second loss value, and a third loss value; wherein the first loss value is determined based on the scene probability vector and the scene label; the second loss value is determined based on the sample features and the class center feature corresponding to the first scene category; the third loss value is determined based on the sample features and the class center feature corresponding to the second scene category; determine a comprehensive loss value according to the first loss value and its corresponding first weight value, the second loss value and its corresponding second weight value, and the third loss value and its corresponding third weight value; and train the original scene recognition model based on the comprehensive loss value.
[0196] During the process of training the original scene recognition model based on the sample images in the sample set, through the original scene recognition model, the scene probability vector corresponding to the input sample image and the sample features of the sample image can be obtained, so that subsequently, based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, the original scene recognition model can be trained to obtain a trained scene recognition model. The trained scene recognition model can make the image features of the images within the same scene category approach the class center features of that scene category and at the same time move away from the class center features of other scene categories. Further combining the feature level of the image, it can be determined whether the scene category of the image can be recognized and, in the case where the scene category of the image can be recognized, the scene category to which the image belongs. This not only realizes the accurate recognition of the scene category images included in the closed image set but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0197] Embodiment 6:
[0198] Figure 7 The following is a schematic structural diagram of a scene recognition device provided by some embodiments of the present application. The present application provides a scene recognition device, including:
[0199] A first processing module 71, configured to determine the image features of the image to be recognized through a pre-trained scene recognition model;
[0200] A second processing module 72, configured to determine the similarity between the image features and the target class center features of each scene category;
[0201] A third processing module 73, configured to determine whether each scene category includes the scene category to which the image to be recognized belongs according to each similarity and a similarity threshold; if it is determined that each scene category includes the scene category to which the image to be recognized belongs, then determine the scene category to which the image to be recognized belongs through the scene recognition model; if it is determined that each scene category does not include the scene category to which the image to be recognized belongs, then do not continue to recognize the scene category to which the image to be recognized belongs.
[0202] Since the scene recognition model is pre-trained, and the scene recognition model is obtained by training the original scene recognition model based on the scene probability vector of the sample image, the scene label of the sample image, the sample feature of the sample image, the class center feature corresponding to the first scene category of the sample image, the sample feature of the sample image, and the class center feature corresponding to the second scene category of the sample image, during the process of recognizing the scene category to which the image to be recognized belongs based on the scene recognition model, it is possible to make the image features within the same scene category approach the class center feature of that scene category and at the same time move away from the class center features of other scene categories. Further combining with the feature level of the image, it is determined whether the scene category of the image can be recognized, and in the case where the scene category of the image can be recognized, the scene category to which the image belongs, which not only realizes the accurate recognition of the scene category images included in the closed image set, but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0203] Embodiment 7:
[0204] As Figure 8 is a schematic structural diagram of an electronic device provided in some embodiments of the present application. On the basis of the above embodiments, the present application also provides an electronic device, as Figure 8 shown, including: a processor 81, a communication interface 82, a memory 83, and a communication bus 84, where the processor 81, the communication interface 82, and the memory 83 complete mutual communication through the communication bus 84;
[0205] A computer program is stored in the memory 83, and when the program is executed by the processor 81, the processor 81 is caused to execute the following steps:
[0206] Obtain any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs;
[0207] Through the original scene recognition model, determine the scene probability vector corresponding to the sample image and the sample feature of the sample image; wherein, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively;
[0208] Based on the scene probability vector, the scene label, the sample feature, the class center feature corresponding to the first scene category, the sample feature, and the class center feature corresponding to the second scene category, train the original scene recognition model to obtain a trained scene recognition model; wherein, the second scene category is the scene category other than the first scene category among each scene category.
[0209] Since the principle of the above electronic device for solving problems is similar to the method for training a scene recognition model, the implementation of the above electronic device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0210] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0211] The communication interface 82 is used for communication between the above electronic device and other devices.
[0212] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0213] The above processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0214] During the process of training the original scene recognition model based on the sample images in the sample set, through the original scene recognition model, the scene probability vector corresponding to the input sample image and the sample features of the sample image can be obtained, enabling subsequent training of the original scene recognition model based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the class center features corresponding to the second scene category to obtain a trained scene recognition model. The trained scene recognition model can make the image features within the same scene category approach the class center features of that scene category and move away from the class center features of other scene categories. Further combining with the feature level of the image, it can determine whether the scene category of the image can be recognized and, in the case where the scene category of the image can be recognized, the scene category to which the image belongs. This not only realizes the accurate recognition of the scene category images included in the closed image set but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0215] Embodiment 8:
[0216] As Figure 9 is a schematic structural diagram of an electronic device provided by some embodiments of the present application. On the basis of the above embodiments, the present application further provides an electronic device, as Figure 9 shown, including: a processor 91, a communication interface 92, a memory 93, and a communication bus 94. Among them, the processor 91, the communication interface 92, and the memory 93 complete communication with each other through the communication bus 94;
[0217] A computer program is stored in the memory 93. When the program is executed by the processor 91, the processor 91 is caused to execute the following steps:
[0218] Determine the image features of the image to be recognized through a pre-trained scene recognition model;
[0219] Determine the similarities between the image features and the target class center features of each scene category;
[0220] According to each of the similarities and a similarity threshold, determine whether each scene category includes the scene category to which the image to be recognized belongs;
[0221] If it is determined that each scene category includes the scene category to which the image to be recognized belongs, then determine the scene category to which the image to be recognized belongs through the scene recognition model;
[0222] If it is determined that each of the scene categories does not include the scene category to which the image to be recognized belongs, then the recognition of the scene category to which the image to be recognized belongs is not continued.
[0223] Since the principle of the above electronic device for solving problems is similar to that of the scene recognition method, the implementation of the above electronic device can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0224] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0225] The communication interface 92 is used for communication between the above electronic device and other devices.
[0226] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0227] The above processor may be a general-purpose processor, including a central processing unit, a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0228] Since the scene recognition model is pre-trained and the scene recognition model is obtained by training the original scene recognition model based on the scene probability vector of the sample image, the scene label of the sample image, the sample feature of the sample image, the class center feature corresponding to the first scene category of the sample image, and the sample feature of the sample image and the class center feature corresponding to the second scene category of the sample image, during the process of recognizing the scene category to which the image to be recognized belongs based on the scene recognition model, it is possible to make the image features within the same scene category approach the class center feature of the scene category and move away from the class center features of other scene categories. Further, by combining the feature level of the image, it is determined whether the scene category of the image can be recognized, and when the scene category of the image can be recognized, the scene category to which the image belongs. This not only realizes the accurate recognition of the scene category images included in the closed image set but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0229] Embodiment 9:
[0230] Based on the above embodiments, the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor is caused to perform the following steps when executing:
[0231] Obtain any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs;
[0232] Determine the scene probability vector corresponding to the sample image and the sample feature of the sample image through the original scene recognition model; wherein, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively;
[0233] Based on the scene probability vector, the scene label, the sample feature, the class center feature corresponding to the first scene category, the sample feature, and the class center feature corresponding to the second scene category, train the original scene recognition model to obtain a trained scene recognition model; wherein, the second scene category is the scene category other than the first scene category among each scene category.
[0234] Since the principle of solving problems by the above-provided computer-readable medium is similar to that of the scene recognition model training method, after the processor executes the computer program in the above computer-readable medium, the steps implemented can refer to the implementation of the method, and the repeated parts will not be described again.
[0235] During the process of training the original scene recognition model based on the sample images in the sample set, through the original scene recognition model, the scene probability vector corresponding to the input sample image and the sample features of the sample image can be obtained, so that subsequently, based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, the original scene recognition model can be trained to obtain a trained scene recognition model. The trained scene recognition model can make the image features of images within the same scene category converge towards the class center features of that scene category and at the same time move away from the class center features of other scene categories. Further, by combining the feature level of the image, it can be determined whether the scene category of the image can be recognized and, in the case where the scene category of the image can be recognized, the scene category to which the image belongs. This not only realizes the accurate recognition of the scene category images included in the closed image set but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0236] Embodiment 10:
[0237] Based on the above embodiments, the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the following steps are implemented when the processor executes:
[0238] Determine the image features of the image to be recognized through a pre-trained scene recognition model;
[0239] Determine the similarity between the image features and the target class center features of each scene category;
[0240] According to each similarity and the similarity threshold, determine whether each scene category contains the scene category to which the image to be recognized belongs;
[0241] If it is determined that each scene category contains the scene category to which the image to be recognized belongs, then determine the scene category to which the image to be recognized belongs through the scene recognition model;
[0242] If it is determined that each scene category does not contain the scene category to which the image to be recognized belongs, then do not continue to recognize the scene category to which the image to be recognized belongs.
[0243] Since the principle of solving problems by the above-provided computer-readable medium is similar to that of the scene recognition method, the steps implemented after the processor executes the computer program in the above computer-readable medium can refer to the implementation of the method, and the repeated parts will not be elaborated.
[0244] Since the scene recognition model is pre-trained, and the scene recognition model is obtained by training the original scene recognition model based on the scene probability vector of the sample image, the scene label of the sample image, the sample feature of the sample image, the class center feature corresponding to the first scene category of the sample image, and the class center feature corresponding to the second scene category of the sample image. During the process of recognizing the scene category to which the image to be recognized belongs based on this scene recognition model, it is possible to make the image features within the same scene category approach the class center feature of this scene category and at the same time move away from the class center features of other scene categories. Further combining with the feature level of the image, it is determined whether the scene category of this image can be recognized, and when the scene category of this image can be recognized, the scene category to which this image belongs. It not only realizes accurately recognizing the scene category images included in the closed image set, but also can process the scene category images that do not belong to the closed image set, improving the accuracy, performance, and naturalness of the scene recognition model.
[0245] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0246] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0247] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.
[0248] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one box or more boxes.
[0249] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for training a scene recognition model, characterized in that The method includes: Obtaining any sample image in the sample set; wherein, the sample image corresponds to a scene label, and the scene label is used to identify the first scene category to which the sample image belongs; Determining, by the original scene recognition model, the scene probability vector corresponding to the sample image and the sample features of the sample image; wherein, the scene probability vector includes the probability values of the sample image belonging to each scene category respectively; Training the original scene recognition model based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, the sample features, and the class center features corresponding to the second scene category, so as to obtain a trained scene recognition model; wherein, the second scene category is the scene category other than the first scene category among each scene category; Wherein, obtaining the class center features of each scene category includes: For each iterative training of the original scene recognition model, obtaining the sample features of the sample images of each scene category in the sample set through the scene recognition model of the current iteration; determining the candidate class center features corresponding to each scene category according to each of the sample features; and determining the class center features of each scene category in the next iterative training based on each of the candidate class center features; Wherein, determining the candidate class center features corresponding to each scene category according to each of the sample features includes: For each scene category, determining the target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of the scene category; determining a weighted average vector according to the sample features of the target sample images and the weight values of the target sample images, and determining the candidate class center features corresponding to the scene category based on the weighted average vector; wherein, the weight value of the target sample image is preset, or is determined according to the probability value that the target sample image belongs to the scene category obtained through the scene recognition model of the current iteration.
2. The method according to claim 1, wherein Determining the class center features of each scene category in the next iterative training based on each of the candidate class center features includes: If the current iteration is the first iteration, determining each of the candidate class center features as the class center features of each scene category in the next iterative training; If the current iteration is not the first iteration, then for each scene category, determining the difference vector between the candidate class center features corresponding to the scene category and the class center features corresponding to the scene category determined in the current iteration; and determining the class center features corresponding to the scene category in the next iterative training according to the difference vector, the pre-configured weight vector, and the class center features corresponding to the scene category determined in the current iteration.
3. The method according to claim 1, characterized in that, Training the original scene recognition model based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, the sample features, and the class center features corresponding to the second scene category includes: Determine a first loss value, a second loss value, and a third loss value; wherein, the first loss value is determined based on the scenario probability vector and the scenario label; the second loss value is determined based on the sample feature and the class center feature corresponding to the first scenario category; the third loss value is determined based on the sample feature and the class center feature corresponding to the second scenario category; Determine a comprehensive loss value according to the first loss value and its corresponding first weight value, the second loss value and its corresponding second weight value, and the third loss value and its corresponding third weight value; Train the original scenario recognition model based on the comprehensive loss value.
4. A scene recognition method, characterized in that, The method includes: Determine the image feature of the image to be recognized through a pre-trained scenario recognition model; Determine the similarity between the image feature and the target class center feature of each scenario category; According to each similarity and a similarity threshold, determine whether each scenario category includes the scenario category to which the image to be recognized belongs; If it is determined that each scenario category includes the scenario category to which the image to be recognized belongs, determine the scenario category to which the image to be recognized belongs through the scenario recognition model; If it is determined that each scenario category does not include the scenario category to which the image to be recognized belongs, do not continue to recognize the scenario category to which the image to be recognized belongs; Wherein, the target class center feature of each scenario category is the class center feature of each scenario category included in the sample set during the last iterative training of the original scenario recognition model; Wherein, obtaining the class center feature of each scenario category includes: For each iterative training of the original scenario recognition model, obtain the sample features of the sample images of each scenario category in the sample set through the scenario recognition model of the current iteration; according to each sample feature, determine the candidate class center feature corresponding to each scenario category; based on each candidate class center feature, determine the class center feature of each scenario category in the next iterative training; Wherein, determining the candidate class center feature corresponding to each scenario category according to each sample feature includes: For each scenario category, determine the target sample images correctly recognized by the scenario recognition model of the current iteration among the respective sample images of this scenario category; determine a weighted average vector according to the sample feature of the target sample image and the weight value of the target sample image, and determine the candidate class center feature corresponding to this scenario category based on the weighted average vector; wherein, the weight value of the target sample image is pre-set or determined according to the probability value that the target sample image belongs to this scenario category obtained through the scenario recognition model of the current iteration.
5. A scene recognition model training device, characterized in that, The device includes: An acquisition unit, configured to acquire any sample image in the sample set; wherein, the sample image corresponds to a scenario label, and the scenario label is used to identify the first scenario category to which the sample image belongs; A processing unit, configured to determine a scene probability vector corresponding to the sample image and sample features of the sample image through an original scene recognition model; wherein, the scene probability vector includes probability values of the sample image belonging to each scene category respectively; A training unit, configured to train the original scene recognition model based on the scene probability vector, the scene label, the sample features, the class center features corresponding to the first scene category, and the sample features and the class center features corresponding to the second scene category, so as to obtain a trained scene recognition model; wherein, the second scene category is a scene category other than the first scene category in each of the scene categories; Wherein, the training unit is specifically configured to, for each iterative training of the original scene recognition model, obtain the sample features of the sample images of each scene category in the sample set through the scene recognition model of the current iteration; determine candidate class center features corresponding to each of the scene categories according to each of the sample features; and determine the class center features of each of the scene categories in the next iterative training based on each of the candidate class center features; The training unit is specifically configured to, for each scene category, determine target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of the scene category; determine a weighted average vector according to the sample features of the target sample images and the weight values of the target sample images, and determine candidate class center features corresponding to the scene category based on the weighted average vector; wherein, the weight values of the target sample images are preset, or are determined according to the probability values of the target sample images belonging to the scene category obtained through the scene recognition model of the current iteration.
6. A scene recognition device, characterized in that, The apparatus includes: A first processing module, configured to determine image features of an image to be recognized through a pre-trained scene recognition model; A second processing module, configured to determine similarities between the image features and target class center features of each scene category; A third processing module, configured to determine whether each of the scene categories includes the scene category to which the image to be recognized belongs according to each of the similarities and a similarity threshold; if it is determined that each of the scene categories includes the scene category to which the image to be recognized belongs, determine the scene category to which the image to be recognized belongs through the scene recognition model; if it is determined that each of the scene categories does not include the scene category to which the image to be recognized belongs, do not continue to recognize the scene category to which the image to be recognized belongs; Wherein, the target class center features of each scene category are the class center features of each scene category included in the sample set during the last iterative training of the original scene recognition model; Wherein, obtaining the class center features of each scene category includes: For each iterative training of the original scene recognition model, sample features of the sample images of each scene category in the sample set are obtained through the scene recognition model of the current iteration; according to each of the sample features, candidate class center features corresponding to each of the scene categories are determined; based on each of the candidate class center features, class center features of each of the scene categories in the next iterative training are determined. Among them, the determining of the candidate class center features corresponding to each of the scene categories according to each of the sample features includes: For each of the scene categories, target sample images correctly recognized by the scene recognition model of the current iteration among the respective sample images of the scene category are determined; according to the sample features of the target sample images and the weight values of the target sample images, a weighted average vector is determined, and based on the weighted average vector, the candidate class center feature corresponding to the scene category is determined; wherein, the weight value of the target sample image is preset or determined according to the probability value that the target sample image belongs to the scene category obtained through the scene recognition model of the current iteration.
7. An electronic device, characterized in that, The electronic device includes a processor, and the processor is configured to implement the steps of the scene recognition model training method as described in any one of claims 1-3, or implement the steps of the scene recognition method as described in claim 4 when executing a computer program stored in a memory.
8. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements the steps of the scene recognition model training method as described in any one of claims 1-3, or implements the steps of the scene recognition method as described in claim 4.
Citation Information
Patent Citations
Image classification and neural network training method and device, equipment and storage medium
CN111259967A