An image recognition method, device, computer device and storage medium
The proposed image recognition method improves training efficiency and accuracy by clustering images, filtering strong and weak samples, and refining labels, addressing the inefficiencies and inaccuracies of manual labeling and single-method training.
Patent Information
- Application Number
- CN202011447629.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-12-09
AI Technical Summary
The prior art requires a lot of manpower to manually mark categories in image classification, resulting in large human errors, low model recognition accuracy, and a single training method, which reduces the accuracy of image recognition.
By clustering multiple sample images, obtaining image sets, and using the image recognition model to predict categories, cleaning and correcting sample images in the image set, setting category labels, training the image recognition model, and obtaining the trained image recognition model.
It improves the accuracy and reliability of image recognition model training, and enhances the accuracy of image recognition by the model after training.
Smart Images

Figure CN113392867B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to an image recognition method, apparatus, computer device, and storage medium. Background Art
[0002] Currently, when classifying images, images can be classified manually or by a model. For example, in the process of classifying images by a model, generally the model needs to be trained first so that the trained model can classify images. Specifically, multiple images are collected, and the categories of the images are defined manually, and then the model is trained based on the images with defined categories, and the trained model is used to classify the images.
[0003] Since it is necessary to manually define the categories of images, it requires a large amount of human cost and is affected by human subjective factors, so it is relatively easy to miss or make mistakes. For example, when manually defining the categories of images of multiple shots in a video, the multiple images of the same shot are very similar and belong to the same category. If each image is labeled one by one, the human input is very large and there is duplicate labeling, resulting in repeated human input. Moreover, the training method of the model is relatively single, resulting in a relatively low recognition accuracy of the trained model, reducing the accuracy of the trained model in classifying images. Summary of the Invention
[0004] Embodiments of this application provide an image recognition method, apparatus, computer device, and storage medium, which can improve the accuracy and reliability of training an image recognition model, so as to improve the accuracy of the trained image recognition model in image recognition.
[0005] To solve the above technical problems, the embodiments of this application provide the following technical solutions:
[0006] Embodiments of this application provide an image recognition method, including:
[0007] Obtain multiple sample images, perform clustering on the multiple sample images to obtain at least one set of image sets;
[0008] Use an image recognition model to perform category prediction on the sample images in the image set to obtain the category prediction probability corresponding to each sample image in the image set;
[0009] Clean the sample images in the image set whose category prediction probability is greater than a first threshold to obtain strong sample images;
[0010] Correct the sample images in the image set whose category prediction probability is less than a second threshold to obtain weak sample images;
[0011] Set class labels for the strong sample images and the weak sample images respectively. Train the image recognition model based on the strong sample images, the weak sample images and the class labels to obtain a trained image recognition model, so as to perform class recognition on images through the trained image recognition model.
[0012] According to one aspect of the present application, there is also provided an image recognition device, including:
[0013] A clustering unit, configured to obtain multiple sample images, perform clustering on the multiple sample images, and obtain at least one set of image sets;
[0014] A prediction unit, configured to perform class prediction on the sample images in the image set through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image set;
[0015] A cleaning unit, configured to clean the sample images in the image set whose class prediction probability is greater than a first threshold to obtain strong sample images;
[0016] A correction unit, configured to correct the sample images in the image set whose class prediction probability is less than a second threshold to obtain weak sample images;
[0017] A training unit, configured to set class labels for the strong sample images and the weak sample images respectively. Train the image recognition model based on the strong sample images, the weak sample images and the class labels to obtain a trained image recognition model, so as to perform class recognition on images through the trained image recognition model.
[0018] According to one aspect of the present application, there is also provided a computer device, including a processor and a memory. A computer program is stored in the memory. When the processor calls the computer program in the memory, it executes any one of the image recognition methods provided in the embodiments of the present application.
[0019] According to one aspect of the present application, there is also provided a storage medium, which is used to store a computer program. The computer program is loaded by a processor to execute any one of the image recognition methods provided in the embodiments of the present application.
[0020] Embodiments of the present application can obtain multiple sample images, cluster the multiple sample images to obtain at least one set of image sets, and then can predict the categories of the sample images in the image sets through an image recognition model to obtain the category prediction probabilities corresponding to each sample image in the image sets; clean the sample images in the image sets with category prediction probabilities greater than a first threshold to obtain strong sample images, and correct the sample images in the image sets with category prediction probabilities less than a second threshold to obtain weak sample images. At this time, category labels can be set for the strong sample images and the weak sample images respectively, and the image recognition model can be trained according to the strong sample images, the weak sample images and the category labels to obtain a trained image recognition model, so as to perform category recognition on images through the trained image recognition model. This solution clusters multiple sample images and sets category labels, and cleans and corrects the sample images based on the category prediction probabilities, etc., so as to train the image recognition model based on the strong sample images obtained by cleaning, the weak sample images obtained by correction and the category labels, and obtain a trained image recognition model with high recognition accuracy, improving the accuracy and reliability of training the image recognition model, so as to improve the accuracy of the trained image recognition model in image recognition. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 is a schematic diagram of the scenario to which the image recognition method provided by the embodiments of the present application is applied;
[0023] Figure 2 is a schematic flowchart of the image recognition method provided by the embodiments of the present application;
[0024] Figure 3 is a schematic diagram of obtaining multiple sample images by splitting scenes provided by the embodiments of the present application;
[0025] Figure 4 is another schematic diagram of obtaining multiple sample images by splitting scenes provided by the embodiments of the present application;
[0026] Figure 5 is a schematic diagram of setting category labels for sample images in an image set provided by the embodiments of the present application;
[0027] Figure 6 is another schematic diagram of setting category labels for sample images in an image set provided by the embodiments of the present application;
[0028] Figure 7It is another schematic diagram for setting category labels for sample images in the image set provided by the embodiments of the present application;
[0029] Figure 8 It is a schematic diagram of the image recognition device provided by the embodiments of the present application;
[0030] Figure 9 It is a schematic structural diagram of a computer device provided by the embodiments of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0032] The embodiments of the present application provide an image recognition method, device, computer device, and storage medium.
[0033] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of the scenario of the image recognition system provided by the embodiments of the present application. The image recognition system may include an image recognition device, which may be specifically integrated in a computer device. The computer device may be a terminal or a server, etc. Among them, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto. The terminal may be a mobile phone, a tablet computer, a notebook computer, a desktop computer, or a wearable device, etc.
[0034] Among them, the computer device can be used to obtain multiple sample images, cluster the multiple sample images to obtain at least one set of image sets, and then can predict the categories of the sample images in the image sets through an image recognition model to obtain the category prediction probabilities corresponding to each sample image in the image sets; clean the sample images in the image sets with category prediction probabilities greater than the first threshold to obtain strong sample images, and correct the sample images in the image sets with category prediction probabilities less than the second threshold to obtain weak sample images. At this time, category labels can be set for the strong sample images and the weak sample images respectively, and the image recognition model can be trained according to the strong sample images, the weak sample images and the category labels to obtain a trained image recognition model. At this time, an image to be recognized can be obtained, and the category probability of the image to be recognized can be calculated through the trained image recognition model to obtain a target category probability, and the category to which the image to be recognized belongs can be determined according to the target category probability. This solution clusters multiple sample images and sets category labels, and cleans and corrects sample images based on category prediction probabilities, etc., so as to train the image recognition model based on the strong sample images obtained by cleaning, the weak sample images obtained by correction and the category labels, and obtain a trained image recognition model with high recognition accuracy, improving the accuracy and reliability of training the image recognition model, and improving the accuracy of the trained image recognition model for image recognition.
[0035] It should be noted that Figure 1 The scene schematic diagram of the image recognition system shown is only an example. The image recognition system and the scene described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the image recognition system and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0036] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0037] The image recognition method provided by the embodiments of the present application may involve technologies such as machine learning technology in artificial intelligence. First, the artificial intelligence technology and the machine learning technology will be described below.
[0038] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formal teaching learning.
[0041] In this embodiment, a description will be made from the perspective of an image recognition device, which can be specifically integrated in computer devices such as servers or terminals.
[0042] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image recognition method provided by an embodiment of this application. The image recognition method may include:
[0043] S101. Obtain multiple sample images, perform clustering on the multiple sample images, and obtain at least one set of image sets.
[0044] Among them, the acquisition method, type, and content included in the sample images can be flexibly set according to actual needs. For example, the sample images can include a target object, which can be a person, an item, an animal, a plant, etc. Multiple sample images can be obtained from a local database, or multiple sample images can be collected through a preset camera or camera, or multiple sample images sent by a server or a terminal can be received, and so on.
[0045] In one embodiment, obtaining multiple sample images may include: obtaining a sample video containing multiple shots; performing shot segmentation on the sample video to obtain multiple sample images corresponding to each shot respectively.
[0046] To improve the convenience and flexibility of obtaining sample images and facilitate subsequent training of the image recognition model, a sample video containing multiple shots can be obtained. The sample video can include one or more. For example, a sample video containing multiple shots can be obtained from a local database, or a sample video containing multiple shots can be collected through a preset camera or camera, etc. Then, the sample video can be subjected to shot segmentation to obtain multiple sample images corresponding to each shot respectively. For example, the sample video can be segmented through SceneDetect v5.0 in the open-source video segmentation library python to obtain multiple shots contained in the sample video. Each shot can correspond to one or more sample images, and the target objects contained in the sample images corresponding to the same shot can be the same. For example, as Figure 3 shown, for shot A, Figure 3 (a), Figure 3 (b), and Figure 3 (c) and other multiple sample images can be obtained. For example, as Figure 4 shown, for shot B, Figure 4 (d), Figure 4 (e), and Figure 4 (f) and other multiple sample images can be obtained.
[0047] After obtaining multiple sample images, the multiple sample images can be clustered to obtain at least one set of image sets. For example, the multiple sample images can be clustered through a k-means clustering algorithm or a similarity model to obtain at least one set of image sets. By clustering and grouping the multiple sample images, the speed of subsequent annotation (such as setting labels) of the sample images can be improved, saving time costs.
[0048] In one embodiment, clustering multiple sample images to obtain at least one set of image sets may include: calculating the similarity values between every two of the multiple sample images through a trained similarity model to obtain the similarity values between every two sample images; clustering the multiple sample images according to the similarity values to obtain at least one set of image sets.
[0049] To improve the accuracy and efficiency of clustering multiple sample images, the multiple sample images can be clustered through a trained similarity model. The specific type of the similarity model can be flexibly set according to actual needs. For example, the similarity model can be a residual network resnet50, an image classification network googlenet, or a support vector machine (SVM), etc. Specifically, the similarity values between every two of the multiple sample images can be calculated through the trained similarity model to obtain the similarity values between every two sample images, and then the multiple sample images can be clustered according to the similarity values to obtain at least one set of image sets. For example, the sample images with similarity values greater than a preset similarity threshold can be clustered into the same set of image sets, and the sample images in the same set of image sets can be multiple sample images corresponding to the same shot.
[0050] In one embodiment, before calculating the similarity values between every two of the multiple sample images through a trained similarity model to obtain the similarity values between every two sample images, the image recognition method may further include: obtaining an initial image, performing enhancement processing on the initial image to obtain multiple training sample images; predicting the similarity values between every two of the multiple training sample images through an initial similarity model to obtain the similarity prediction values between every two training sample images; training the initial similarity model based on the similarity prediction values and the pre-annotated similarity values through a preset cross-entropy loss function to adjust the parameters of the initial similarity model and obtain a trained similarity model.
[0051] To improve the accuracy of the similarity model in clustering sample images, the similarity model can be pre-trained. Specifically, an initial image can be obtained. The initial image can include one or more. For example, the initial image can be obtained from a local database, or can be collected through a preset camera or camera, or the imagenet open-source data or openimage open-source data, etc. can be used as the initial image.
[0052] Then, in order to enrich the training sample images and make the similarity model more effective in subsequent clustering, the initial images can be enhanced to obtain multiple training sample images. For example, the initial images can be enhanced by adding noise, rotating, adding borders, and cropping to obtain multiple training sample images. Secondly, the initial similarity model can be used to predict the similarity between every two of the multiple training sample images, and the similarity prediction values between every two training sample images can be obtained. At this time, a cross-entropy loss function can be constructed. Through this cross-entropy loss function, based on the similarity prediction values and the pre-annotated similarity values, the initial similarity model can be trained to adjust the parameters of the initial similarity model. After multiple iterative trainings, the trained similarity model can be obtained.
[0053] Among them, the cross-entropy loss function H(p,q) can be shown as follows:
[0054] H(p,q) = -∑ p (x)logq(x)
[0055] p can represent the pre-annotated similarity value (i.e., the correct answer), and q can represent the predicted similarity prediction value (i.e., the prediction value). The smaller the cross-entropy value calculated by the cross-entropy loss function, the closer p and q are. On the contrary, the larger the cross-entropy value, the greater the gap between p and q.
[0056] It should be noted that the open-source imagenet pre-trained weights can be used as the initial parameters of the similarity model, and multiple training sample images can be divided into multiple batches (i.e., multiple batches) and the standard stochastic gradient descent (SGD) optimization method can be used to update the parameters of the similarity model.
[0057] After obtaining the trained similarity model, the multiple sample images can be clustered by the trained similarity model. For example, the multiple sample images corresponding to each shot in the sample video can be clustered: one sample image is selected from each shot (i.e., each sub-shot) to represent this sub-shot, and all the multiple sample images corresponding to each shot are input into the trained similarity model for feature extraction to obtain the depth features corresponding to the sample images (which can be 1*2048-dimensional vectors). Kmeans clustering is performed based on the depth features. For example, clustering can be performed based on the similarity values between every two sample images to obtain an image set. The clustering categories of the image set can be determined by the quantity of the sample video. For N sample videos, the clustering categories can be between N / 2 and N / 5. After clustering, 1 / 3 of the sample images closest to the class center in each cluster (i.e., the sample images with larger similarities) can be taken as the initial samples of the cluster (i.e., the obtained image set).
[0058] It should be noted that the class centers can also be manually selected and merged, etc. to obtain an image set. For example, for 100 videos, with 15 sub-shots per video and 10 sample images corresponding to each sub-shot, N / 5 clustering centers can be selected, and the first 1 / 3 of the sample images before clustering can be selected as the initial clustering samples. At this time, only 100*15 / 5*1 / 3 = 100 sample images need to be screened manually, instead of the original 100*15*10 = 15,000 sample images. The initial annotation quantity (i.e., the quantity of class label settings) is compressed to a maximum of 1 / 30 of the original, improving the efficiency of subsequent class label settings.
[0059] After obtaining the image set, class labels can be set for each sample image in the image set based on the clustering results. The class labels of the sample images in the same image set can be the same. For example, the feature information of the target object included in the sample images in the image set can be extracted, and class labels can be set for each sample image in the image set based on the feature information of the target object. Another example is that since the sample images in the same group of image sets can be multiple sample images corresponding to the same shot (i.e., the same sub-shot), class labels can be set by video sub-shot and using the sample images corresponding to the same sub-shot as a labeling unit, reducing the quantity of class label settings and improving the efficiency of class label settings.
[0060] For example, as Figure 5 shown, the class labels of sample images such as sample image (g), sample image (h), and sample image (i) in image set A can be set to "human upper body". Another example is that as Figure 6 shown, the class labels of sample images such as sample image (j), sample image (k), and sample image (l) in image set B can be set to "full body". Another example is that as Figure 7 shown, the class labels of sample images such as sample image (m), sample image (n), and sample image (o) in image set C can be set to "crowd".
[0061] S102. Perform class prediction on the sample images in the image set through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image set.
[0062] Among them, the image recognition model can be flexibly set according to actual needs. For example, the image recognition model can be a deep convolutional neural network or a residual network, etc. For example, a series of operations such as convolution, residual connection, and pooling can be performed on the sample images in the image set through the image recognition model to perform class prediction and obtain the class prediction probability corresponding to each sample image in the image set.
[0063] S103. Clean the sample images in the image set with class prediction probabilities greater than the first threshold to obtain strong sample images.
[0064] Among them, the strong sample images can be a set of sample images obtained by deleting the sample images in the image set whose predicted class of the sample image does not match the class label of the image set.
[0065] In one embodiment, cleaning the sample images in the image set whose class prediction probability is greater than the first threshold to obtain the strong sample images may include: determining the first target class to which the sample images in the image set whose class prediction probability is greater than the first threshold belong; deleting the sample images whose first target class to which the sample images belong does not match the class label of the image set where the sample images are located, so as to obtain the strong sample images.
[0066] To improve the accuracy of training the image recognition model, the sample images in the image set can be further screened based on the class prediction results, so as to continue training the image recognition model based on the screened sample images. For example, the sample images in the image set whose class prediction probability is greater than the first threshold can be screened out. The first threshold can be flexibly set according to actual needs. For example, the first threshold can be set to 0.9. If the class prediction probability of the sample image is greater than the first threshold, it means that the class prediction probability of the sample image is relatively high and it is a sample with strong confidence. Then, the first target class to which the sample images in the image set whose class prediction probability is greater than the first threshold belong can be determined. At this time, the first target class to which the sample images belong can be compared with the class label of the image set where the sample images are located to determine whether the first target class to which the sample images belong matches the class label of the image set where the sample images are located (for example, whether they are both in the "crowd" class). If the first target class to which the sample images belong does not match the class label of the image set where the sample images are located, the unmatched sample images are deleted, and only the sample images corresponding to the match between the first target class to which the sample images belong and the class label of the image set where the sample images are located are retained in the image set, so that the strong sample images can be obtained.
[0067] For example, for sample images a, b, c, d, e, f, etc. in image set A with the class label "crowd", when the class prediction probability corresponding to sample image a is 0.9, the class prediction probability corresponding to sample image b is 0.99, the class prediction probability corresponding to sample image c is 0.97, the class prediction probability corresponding to sample image d is 0.6, the class prediction probability corresponding to sample image e is 0.98, and the class prediction probability corresponding to sample image f is 0.96, sample images with a class prediction probability greater than 0.9 can be screened out from image set A, obtaining sample images a, b, c, e, f, etc. It is determined that the class to which sample image a belongs is "crowd", the class to which sample image b belongs is "crowd", the class to which sample image c belongs is "crowd", the class to which sample image e belongs is "", and the class to which sample image f belongs is "half body of a person". Since the class to which sample image f belongs does not match the class label of image set A, sample image f can be deleted at this time, and the obtained strong sample images include sample images a, b, c, and e, etc.
[0068] S104. Modify the sample images in the image set whose class prediction probability is less than the second threshold to obtain weak sample images.
[0069] Among them, the weak sample image can be a sample image obtained by modifying the class of the sample image whose predicted class in the image set does not match the class label of the image set.
[0070] In one embodiment, modifying the sample images in the image set whose class prediction probability is less than the second threshold to obtain weak sample images may include: determining the second target class to which the sample images in the image set with a class prediction probability less than the second threshold belong; when the second target class to which the sample image belongs does not match the class label of the image set where the sample image is located, modifying the second target class to which the sample image belongs to the class label of the image set where the sample image is located to obtain weak sample images.
[0071] To improve the reliability of training an image recognition model, sample images in an image set can be corrected based on the class prediction results, so as to continue training the image recognition model based on the corrected sample images. Specifically, sample images in the image set with a class prediction probability less than a second threshold can be screened out. The second threshold can be flexibly set according to actual needs. For example, the second threshold can be set to 0.9 or 0.8, etc. If the class prediction probability of a sample image is less than the second threshold, it means that the class prediction probability of this sample image is low, and it is a sample with low confidence. Then, the second target class to which the sample images in the image set with a class prediction probability less than the second threshold belong can be determined. At this time, the second target class to which the sample image belongs can be compared with the class label of the image set where the sample image is located to determine whether the second target class to which the sample image belongs matches the class label of the image set where the sample image is located. If the second target class to which the sample image belongs does not match the class label of the image set where the sample image is located, the second target class to which the sample image belongs is corrected to the class label of the image set where the sample image is located, so that the second target class to which the sample image belongs is the correct class, and thus weak sample images can be obtained. Since the sample images in the same set of image sets can be multiple sample images corresponding to the same shot (i.e., the same sub-shot), the classes of the obtained weak sample images can be the correct classes corresponding to the same sub-shot. Thus, through error correction and leakage detection, etc., the annotation quality of the classes of the sample images can be improved. The performance of the subsequent image recognition model can be improved by optimizing the screening strategies for strong sample images and weak sample images, as well as optimizing the sub-shot effect, and by setting appropriate first and second thresholds to expand the sample images, etc., so that there is a way to optimize the image recognition model.
[0072] For example, for the sample images g, h, i, j, k, etc. in the image set B with the class label "crowd", when the class prediction probability corresponding to the sample image g is 0.9, the class prediction probability corresponding to the sample image h is 0.8, the class prediction probability corresponding to the sample image i is 0.7, the class prediction probability corresponding to the sample image j is 0.6, and the class prediction probability corresponding to the sample image k is 0.5, the sample images with a class prediction probability less than 0.9 can be screened out from the image set B, obtaining the sample images h, i, j, k, etc. It is determined that the class to which the sample image h belongs is "crowd", the class to which the sample image i belongs is "half body of a person", the class to which the sample image j belongs is "whole body of a person", and the class to which the sample image k belongs is "half body of a person". Since the classes to which the sample images i, j, and k belong do not match the class label of the image set B, at this time, the class labels of the sample images i, j, and k can be set to "crowd", and the obtained weak sample images including the sample images h, i, j, k, etc. all have the class label of "crowd".
[0073] S105. Set class labels for the strong sample images and the weak sample images respectively, and train an image recognition model according to the strong sample images, the weak sample images, and the class labels to obtain a trained image recognition model, so as to perform class recognition on images through the trained image recognition model.
[0074] In an embodiment, setting class labels for the strong sample images and the weak sample images respectively, and training an image recognition model according to the strong sample images, the weak sample images, and the class labels to obtain a trained image recognition model may include: performing class prediction on the strong sample images and the weak sample images respectively through the image recognition model to obtain a first class prediction probability corresponding to each strong sample image and a second class prediction probability corresponding to each weak sample image; determining a first class corresponding to the strong sample image according to the first class prediction probability, and determining a second class corresponding to the weak sample image according to the second class prediction probability; respectively converging the first class and the second class with the class label to adjust the parameters of the image recognition model to obtain a trained image recognition model.
[0075] After obtaining the strong sample images and weak sample images, the class labels of each strong sample image and weak sample image can be reset. For example, the feature information of the target object contained in the strong sample image can be extracted, and the class label of the strong sample image can be set based on the feature information of the target object. Also, the feature information of the target object contained in the weak sample image can be extracted, and the class label of the weak sample image can be set based on the feature information of the target object. Another example is that the class label of the image set where the strong sample image is located can be used to set the class label of the strong sample image, and the class label of the weak sample image can be obtained based on the above correction method. Of course, the class labels of each strong sample image and weak sample image can also be set manually, etc. Then, the image recognition model can be used to perform class prediction on the strong sample images to obtain the first class prediction probability corresponding to each strong sample image, and perform class prediction on the weak sample images to obtain the second class prediction probability corresponding to each weak sample image. The first class corresponding to the strong sample image can be determined according to the first class prediction probability. For example, since the first class prediction probability corresponding to the strong sample image can include multiple values (such as the prediction probability of class A is 0.9, the prediction probability of class B is 0.6, etc.), the class with the highest probability value among the multiple first class prediction probabilities can be used as the first class corresponding to the strong sample image (such as class A). And the second class corresponding to the weak sample image can be determined according to the second class prediction probability. For example, the class with the highest probability value among the multiple second class prediction probabilities can be used as the second class corresponding to the weak sample image. At this time, the first class corresponding to the strong sample image can be converged with the class label corresponding to the strong sample image, and the second class corresponding to the weak sample image can be converged with the class label corresponding to the weak sample image, and iterative training can be continuously performed to adjust the parameters of the image recognition model by backpropagation until the test accuracy of the image recognition model on the target data reaches a predetermined value, or the specified number of iterations is completed, etc., so as to obtain the trained image recognition model.
[0076] In one embodiment, after training the image recognition model according to the strong sample images and weak sample images to obtain the trained image recognition model, the image recognition method may further include: obtaining the image to be recognized; calculating the class probability of the image to be recognized through the trained image recognition model to obtain the target class probability; and determining the class to which the image to be recognized belongs according to the target class probability.
[0077] After obtaining the trained image recognition model, the trained image recognition model can be used to recognize images. For example, an image to be recognized can be obtained. For example, the image to be recognized can be obtained from a local database, or the image to be recognized can be captured through a preset camera or camera, etc., or the image to be recognized sent by a server or a terminal, etc. can be received, and so on. Then, the trained image recognition model can be used to calculate the class probability of the image to be recognized to obtain the target class probability. For example, it can be obtained that the probability that the image belongs to class A is 0.98, and the probability that the image belongs to class B is 0.1, etc. At this time, the class to which the image to be recognized belongs can be determined according to the target class probability. For example, since the probability that the image belongs to class A is the highest, it can be determined that the class to which the image to be recognized belongs is class A.
[0078] In the embodiments of the present application, multiple sample images can be obtained, the multiple sample images can be clustered to obtain at least one set of image sets, and then the sample images in the image sets can be predicted for their classes through the image recognition model to obtain the class prediction probability corresponding to each sample image in the image sets; the sample images in the image sets with a class prediction probability greater than a first threshold are cleaned to obtain strong sample images, and the sample images in the image sets with a class prediction probability less than a second threshold are corrected to obtain weak sample images. At this time, class labels can be set for the strong sample images and the weak sample images respectively, and the image recognition model can be trained according to the strong sample images, the weak sample images and the class labels to obtain the trained image recognition model. So that subsequently, an image to be recognized can be obtained, the trained image recognition model can be used to calculate the class probability of the image to be recognized to obtain the target class probability, and the class to which the image to be recognized belongs can be determined according to the target class probability. This solution clusters multiple sample images and sets class labels, and cleans and corrects the sample images based on the class prediction probability, etc., so as to train the image recognition model based on the strong sample images obtained by cleaning, the weak sample images obtained by correction and the class labels, to obtain a trained image recognition model with higher recognition accuracy, improving the accuracy and reliability of training the image recognition model, and improving the accuracy of the trained image recognition model in image recognition.
[0079] To facilitate better implementation of the image recognition method provided by the embodiments of the present application, the embodiments of the present application also provide a device based on the above image recognition method. The meanings of the nouns are the same as those in the above image recognition method, and the specific implementation details can refer to the description in the method embodiments.
[0080] Please refer to Figure 8 , Figure 8The figure is a schematic structural diagram of the image recognition device provided by the embodiment of the present application. The image recognition device may include a clustering unit 301, a prediction unit 302, a cleaning unit 303, a correction unit 304, a training unit 305, etc.
[0081] Among them, the clustering unit 301 is configured to obtain multiple sample images, perform clustering on the multiple sample images, and obtain at least one set of image sets.
[0082] The prediction unit 302 is configured to perform class prediction on the sample images in the image set through an image recognition model, and obtain the class prediction probability corresponding to each sample image in the image set.
[0083] The cleaning unit 303 is configured to clean the sample images in the image set whose class prediction probability is greater than a first threshold, and obtain strong sample images.
[0084] The correction unit 304 is configured to correct the sample images in the image set whose class prediction probability is less than a second threshold, and obtain weak sample images.
[0085] The training unit 305 is configured to set class labels for the strong sample images and the weak sample images respectively, and train the image recognition model according to the strong sample images, the weak sample images and the class labels, so as to obtain a trained image recognition model for classifying images through the trained image recognition model.
[0086] In an embodiment, the training unit 305 may specifically be configured to: perform class prediction on the strong sample images and the weak sample images respectively through the image recognition model, obtain the first class prediction probability corresponding to each strong sample image, and the second class prediction probability corresponding to each weak sample image; determine the first class corresponding to the strong sample image according to the first class prediction probability, and determine the second class corresponding to the weak sample image according to the second class prediction probability; respectively converge the first class and the second class with the class label to adjust the parameters of the image recognition model, and obtain a trained image recognition model.
[0087] In an embodiment, the clustering unit 301 may specifically be configured to: obtain a sample video including multiple shots; perform shot segmentation on the sample video to obtain multiple sample images corresponding to each shot respectively; perform clustering on the multiple sample images to obtain at least one set of image sets.
[0088] In an embodiment, the clustering unit 301 may specifically be configured to: calculate the similarity between every two sample images in the multiple sample images through a trained similarity model, and obtain the similarity value between every two sample images; perform clustering on the multiple sample images according to the similarity value to obtain at least one set of image sets.
[0089] In one embodiment, the image recognition device may further include:
[0090] A processing unit, configured to obtain an initial image, perform enhancement processing on the initial image to obtain multiple training sample images;
[0091] A similarity prediction unit, configured to perform similarity prediction between every two of the multiple training sample images through an initial similarity model to obtain similarity prediction values between every two of the multiple training sample images;
[0092] An adjustment unit, configured to train the initial similarity model based on the similarity prediction values and pre-annotated similarity values through a preset cross-entropy loss function to adjust the parameters of the initial similarity model and obtain a trained similarity model.
[0093] In one embodiment, the cleaning unit 303 may specifically be used to: determine a first target category to which a sample image in the image set with a category prediction probability greater than a first threshold belongs; delete sample images whose first target category to which they belong does not match the category label of the image set where the sample images are located to obtain strong sample images.
[0094] In one embodiment, the correction unit 304 may specifically be used to: determine a second target category to which a sample image in the image set with a category prediction probability less than a second threshold belongs; when the second target category to which the sample image belongs does not match the category label of the image set where the sample image is located, correct the second target category to which the sample image belongs to the category label of the image set where the sample image is located to obtain weak sample images.
[0095] In one embodiment, the image recognition device may further include:
[0096] A calculation unit, configured to obtain an image to be recognized, calculate a target category probability for the image to be recognized through a trained image recognition model;
[0097] An identification unit, configured to determine the category to which the image to be recognized belongs according to the target category probability.
[0098] In an embodiment of the present application, the clustering unit 301 may obtain multiple sample images, perform clustering on the multiple sample images to obtain at least one set of image sets, and then the prediction unit 302 may perform class prediction on the sample images in the image set through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image set; the cleaning unit 303 may clean the sample images in the image set with a class prediction probability greater than a first threshold to obtain strong sample images, and the correction unit 304 may correct the sample images in the image set with a class prediction probability less than a second threshold to obtain weak sample images. At this time, the training unit 305 may set class labels for the strong sample images and the weak sample images respectively, and train the image recognition model according to the strong sample images, the weak sample images, and the class labels to obtain a trained image recognition model, so as to perform class recognition on images through the trained image recognition model. This solution performs clustering and class label setting on multiple sample images, and performs cleaning and correction on sample images based on class prediction probabilities, etc., so as to train the image recognition model based on the strong sample images obtained by cleaning, the weak sample images obtained by correction, and the class labels, and obtain a trained image recognition model with high recognition accuracy, improving the accuracy and reliability of training the image recognition model, so as to improve the accuracy of the trained image recognition model in image recognition.
[0099] An embodiment of the present application also provides a computer device, which may be a server or a terminal, etc. For example, Figure 9 as shown, it shows a schematic structural diagram of the computer device involved in the embodiment of the present application. Specifically:
[0100] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art can understand that Figure 9 the computer device structure shown in does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0101] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by calling the data stored in the memory 402, it performs various functions of the computer device and processes data, thereby conducting an overall detection of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0102] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the computer device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0103] The computer device also includes a power supply 403 that powers each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0104] The computer device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0105] Although not shown, the computer device may also include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to achieve various functions as follows:
[0106] Obtain multiple sample images, perform clustering on the multiple sample images to obtain at least one set of image sets; perform class prediction on the sample images in the image sets through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image sets; clean the sample images in the image sets with class prediction probabilities greater than a first threshold to obtain strong sample images; correct the sample images in the image sets with class prediction probabilities less than a second threshold to obtain weak sample images; set class labels for the strong sample images and the weak sample images respectively, and train the image recognition model according to the strong sample images, the weak sample images and the class labels to obtain a trained image recognition model, so as to perform class recognition on images through the trained image recognition model.
[0107] In one embodiment, when setting class labels for the strong sample images and the weak sample images respectively, and training the image recognition model according to the strong sample images, the weak sample images and the class labels to obtain a trained image recognition model, the processor 401 can be used to execute: perform class prediction on the strong sample images and the weak sample images respectively through the image recognition model to obtain the first class prediction probability corresponding to each strong sample image and the second class prediction probability corresponding to each weak sample image; determine the first class corresponding to the strong sample images according to the first class prediction probability and determine the second class corresponding to the weak sample images according to the second class prediction probability; respectively converge the first class and the second class with the class labels to adjust the parameters of the image recognition model to obtain a trained image recognition model.
[0108] In one embodiment, when obtaining multiple sample images, the processor 401 can be used to execute: obtain a sample video containing multiple shots; perform shot segmentation on the sample video to obtain multiple sample images corresponding to each shot respectively.
[0109] In one embodiment, when performing clustering on multiple sample images to obtain at least one set of image sets, the processor 401 can be used to execute: calculate the similarity value between every two sample images in the multiple sample images through a trained similarity model to obtain the similarity value between every two sample images; perform clustering on the multiple sample images according to the similarity value to obtain at least one set of image sets.
[0110] In one embodiment, before calculating the similarity value between every two sample images among multiple sample images through the trained similarity model, the processor 401 can be used to perform the following operations: obtain an initial image, perform enhancement processing on the initial image to obtain multiple training sample images; perform similarity prediction between every two training sample images among the multiple training sample images through the initial similarity model to obtain the similarity prediction value between every two training sample images; train the initial similarity model based on the similarity prediction value and the pre-annotated similarity value through a preset cross-entropy loss function to adjust the parameters of the initial similarity model and obtain the trained similarity model.
[0111] In one embodiment, when cleaning the sample images in the image set whose class prediction probability is greater than the first threshold to obtain strong sample images, the processor 401 can be used to perform the following operations: determine the first target class to which the sample images in the image set whose class prediction probability is greater than the first threshold belong; delete the sample images whose first target class does not match the class label of the image set where the sample images are located to obtain strong sample images.
[0112] In one embodiment, when correcting the sample images in the image set whose class prediction probability is less than the second threshold to obtain weak sample images, the processor 401 can be used to perform the following operations: determine the second target class to which the sample images in the image set whose class prediction probability is less than the second threshold belong; when the second target class to which the sample images belong does not match the class label of the image set where the sample images are located, correct the second target class to which the sample images belong to the class label of the image set where the sample images are located to obtain weak sample images.
[0113] In one embodiment, after training the image recognition model according to the strong sample images and weak sample images to obtain the trained image recognition model, the processor 401 can be used to perform the following operations: obtain the image to be recognized; calculate the class probability of the image to be recognized through the trained image recognition model to obtain the target class probability; determine the class to which the image to be recognized belongs according to the target class probability.
[0114] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not described in detail in a certain embodiment, reference can be made to the detailed description of the image recognition method above, which will not be elaborated here.
[0115] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations in the foregoing embodiments.
[0116] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the foregoing embodiments can be completed by computer instructions, or by controlling related hardware through computer instructions. The computer instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present application provides a storage medium in which a computer program is stored. The computer includes computer instructions, and the computer program can be loaded by a processor to execute any one of the image recognition methods provided by the embodiments of the present application.
[0117] For the specific implementation of each of the above operations, reference may be made to the foregoing embodiments and will not be elaborated herein.
[0118] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0119] Since the instructions stored in the storage medium can execute the steps in any one of the image recognition methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the image recognition methods provided by the embodiments of the present application can be realized. For details, refer to the foregoing embodiments and will not be elaborated herein.
[0120] The foregoing has introduced in detail an image recognition method, apparatus, computer device, and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image recognition method, characterized in that, Including: Obtain multiple sample images, perform clustering on the multiple sample images to obtain at least one set of image sets; Perform class prediction on the sample images in the image set through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image set; Clean the sample images in the image set with class prediction probabilities greater than a first threshold to obtain strong sample images; Correct the sample images in the image set with class prediction probabilities less than a second threshold to obtain weak sample images; Set class labels for the strong sample images and the weak sample images respectively, and train the image recognition model according to the strong sample images, the weak sample images and the class labels to obtain a trained image recognition model for classifying images through the trained image recognition model, including: Perform class prediction on the strong sample images and the weak sample images respectively through the image recognition model to obtain the first class prediction probability corresponding to each strong sample image and the second class prediction probability corresponding to each weak sample image; Determine the first class corresponding to the strong sample image according to the first class prediction probability and determine the second class corresponding to the weak sample image according to the second class prediction probability; Converge the first class and the second class with the class label respectively to adjust the parameters of the image recognition model to obtain a trained image recognition model.
2. The image recognition method according to claim 1, wherein The obtaining of the multiple sample images includes: Obtain a sample video including multiple shots; Perform shot segmentation on the sample video to obtain multiple sample images corresponding to each shot respectively.
3. The image recognition method according to claim 1, characterized in that The performing clustering on the multiple sample images to obtain at least one set of image sets includes: Calculate the similarity value between every two sample images in the multiple sample images through a trained similarity model to obtain the similarity value between every two sample images; Perform clustering on the multiple sample images according to the similarity value to obtain at least one set of image sets.
4. The image recognition method according to claim 3, wherein Before calculating the similarity value between every two sample images in the multiple sample images through a trained similarity model to obtain the similarity value between every two sample images, the image recognition method further includes: Obtain an initial image, perform enhancement processing on the initial image to obtain multiple training sample images; Perform similarity prediction between every two training sample images in the multiple training sample images through an initial similarity model to obtain the similarity prediction value between every two training sample images; Train the initial similarity model through a preset cross-entropy loss function based on the similarity prediction value and the pre-annotated similarity value to adjust the parameters of the initial similarity model to obtain a trained similarity model.
5. The image recognition method according to claim 1, wherein The cleaning of the sample images in the image set with class prediction probabilities greater than a first threshold to obtain strong sample images includes: Determine the first target class to which the sample images in the image set with class prediction probabilities greater than the first threshold belong; Delete the sample images whose first target class to which they belong does not match the class label of the image set where the sample images are located to obtain strong sample images.
6. The image recognition method according to claim 1, wherein Revising the sample images in the image set whose class prediction probability is less than the second threshold to obtain weak sample images includes: Determining the second target class to which the sample images in the image set with class prediction probability less than the second threshold belong; When the second target class to which the sample image belongs does not match the class label of the image set where the sample image is located, correcting the second target class to which the sample image belongs to the class label of the image set where the sample image is located to obtain weak sample images.
7. The image recognition method according to any one of claims 1 to 6, characterized in that, After training the image recognition model according to the strong sample images and the weak sample images to obtain the trained image recognition model, the image recognition method further includes: Obtaining an image to be recognized; Calculating the class probability of the image to be recognized through the trained image recognition model to obtain the target class probability; Determining the class to which the image to be recognized belongs according to the target class probability.
8. An image recognition device, characterized in that, Including: A clustering unit for obtaining multiple sample images, clustering the multiple sample images to obtain at least one set of image sets; A prediction unit for performing class prediction on the sample images in the image set through an image recognition model to obtain the class prediction probability corresponding to each sample image in the image set; A cleaning unit for cleaning the sample images in the image set with class prediction probability greater than the first threshold to obtain strong sample images; A revising unit for revising the sample images in the image set with class prediction probability less than the second threshold to obtain weak sample images; A training unit for respectively setting class labels for the strong sample images and the weak sample images, and training the image recognition model according to the strong sample images, the weak sample images and the class labels to obtain the trained image recognition model for performing class recognition on images, including: Performing class prediction on the strong sample images and the weak sample images respectively through the image recognition model to obtain the first class prediction probability corresponding to each strong sample image and the second class prediction probability corresponding to each weak sample image; Determining the first class corresponding to the strong sample image according to the first class prediction probability and determining the second class corresponding to the weak sample image according to the second class prediction probability; Converging the first class and the second class with the class label respectively to adjust the parameters of the image recognition model to obtain the trained image recognition model.
9. A storage medium, characterized in that, The storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the image recognition method according to any one of claims 1 to 6.
10. A computer device, including a processor and a memory, where a computer program is stored in the memory, and when the processor calls the computer program in the memory, it executes the image recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Object recognition model optimization method, device and electronic device
CN109063790A
Model training method and device, computer equipment and storage medium
CN111210024A