Image recognition method, image dataset merging method, and related device
By performing attribute classification and feature merging on the face dataset, the problem of high time complexity in dataset merging during face recognition model training is solved, achieving efficient dataset merging and model training, and improving the accuracy and efficiency of face recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOE TECHNOLOGY GROUP CO LTD
- Filing Date
- 2023-02-28
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies suffer from high time complexity when acquiring large amounts of training data in face recognition model training, especially when there are many image labels, merging datasets is extremely time-consuming.
By classifying the multiple datasets to be merged, dividing them into multiple data subsets based on image attributes, and merging them using the target features of image labels, the time complexity is reduced. Euclidean distance or cosine distance is used to calculate similarity, and image labels that meet the similarity conditions are merged.
It effectively reduces the time complexity of data merging, improves the accuracy and training efficiency of face recognition models, and can quickly merge publicly available face datasets in academia to form a large training set containing more IDs.
Smart Images

Figure CN116342935B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to an image recognition method, an image dataset merging method, and related equipment. Background Technology
[0002] Currently, there are two main methods to improve the performance (e.g., accuracy) of computer vision classification algorithms or models: one is to design a good recognition model structure, and the other is to collect more training data. Regarding the method of collecting more training data, for example in face recognition, it is often necessary to collect datasets containing tens of thousands of face images. Google, for instance, used 200 million face images with 80 million IDs to train its model. However, acquiring such a massive amount of training data often presents time complexity issues. Therefore, reducing time complexity when acquiring such large amounts of data becomes particularly important. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose an image recognition method, an image dataset merging method, and related equipment.
[0004] In view of the above objectives, firstly, this application provides an image recognition method, comprising:
[0005] Acquire the image to be recognized;
[0006] Determine the classification of the image to be identified;
[0007] Determine the target dataset corresponding to the classification of the image to be identified, wherein the target dataset is a subset of the image dataset that corresponds to the classification of the image to be identified;
[0008] Based on the image to be identified and the target dataset, the image recognition result is obtained.
[0009] Secondly, embodiments of this application also provide a method for merging image datasets, including:
[0010] Obtain at least two datasets to be merged; each dataset includes multiple images, and each image includes an image identifier;
[0011] Based on the attributes of the images, each dataset is classified to obtain multiple subsets of data to be merged, each corresponding to a different classification.
[0012] Determine at least two subsets of target data corresponding to the target category among the multiple categories;
[0013] Calculate the target features of at least two images with the same image identifier in each of the target data subsets to obtain the target features corresponding to each image identifier;
[0014] Based on the target features corresponding to each image identifier, the images in the at least two target data subsets are merged to obtain a merged target data subset;
[0015] The image dataset is determined based on the merged target data subset.
[0016] Thirdly, embodiments of this application also provide an image recognition device, including:
[0017] The image acquisition module is used to acquire the image to be recognized;
[0018] An image classification module is used to determine the classification of the image to be identified;
[0019] The target dataset determination module is used to determine the target dataset corresponding to the classification of the image to be identified, wherein the target dataset is a subset of the image dataset that corresponds to the classification of the image to be identified;
[0020] The image recognition module obtains the image recognition result based on the image to be recognized and the target dataset.
[0021] Fourthly, embodiments of this application also provide an image dataset merging apparatus, comprising:
[0022] A dataset acquisition module is used to acquire at least two datasets to be merged; each dataset includes multiple images, and each image includes an image identifier;
[0023] The dataset classification module is used to classify each dataset according to the attributes of the image, and obtain multiple data subsets to be merged corresponding to multiple classifications respectively;
[0024] The target data subset determination module determines at least two target data subsets corresponding to the target category among the multiple categories;
[0025] The target feature calculation module is used to calculate the target features of at least two images with the same image identifier in each of the target data subsets, so as to obtain the target features corresponding to each image identifier;
[0026] The merging module is used to merge the images in the at least two target data subsets according to the target features corresponding to each image identifier, so as to obtain a merged target data subset;
[0027] The image dataset determination module is used to determine multiple merged target data subsets as the image dataset.
[0028] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method of the first aspect or the method of the second aspect as described in any of the preceding claims.
[0029] This application also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of the first aspect or the method of the second aspect as described above.
[0030] As can be seen from the above, the image recognition method, image dataset merging method, and related equipment provided in this application reduce the time complexity of data merging by classifying individual image datasets. Using a larger training set can significantly improve the performance of the trained face recognition algorithm and increase the speed of retrieving the image identifier corresponding to the image to be recognized from the dataset. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 A flowchart illustrating an exemplary image dataset merging method according to an embodiment of this application is shown.
[0033] Figure 2a This is a schematic diagram of a classification type of dataset in an embodiment of this application;
[0034] Figure 2b This is a schematic diagram illustrating another classification type of the dataset in an embodiment of this application;
[0035] Figure 3 This is a schematic diagram illustrating a process for classifying images according to an embodiment of this application;
[0036] Figure 4 This is another schematic diagram illustrating the image classification process in an embodiment of this application;
[0037] Figure 5 This is a schematic diagram illustrating the classification of images corresponding to the same image identifier according to an embodiment of this application;
[0038] Figure 6 This is a schematic diagram illustrating the calculation of the average features of all images corresponding to each image identifier in the data subset, according to an embodiment of this application.
[0039] Figure 7 This is a schematic diagram of the first process for image merging according to an embodiment of this application;
[0040] Figure 8 This is a schematic diagram of image merging in an embodiment of this application;
[0041] Figure 9 A schematic flowchart illustrating the image recognition method provided in this application embodiment;
[0042] Figure 10 This is a schematic diagram of an image dataset merging apparatus according to an embodiment of this application;
[0043] Figure 11 This is a schematic diagram of an image recognition device according to an embodiment of this application.
[0044] Figure 12 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0046] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by those skilled in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.
[0047] Facial recognition is typically quite complex, exceeding that of gesture recognition, thus requiring hundreds of millions of training images. As is well known, the amount of training data directly impacts the performance of a facial recognition model; for example, Google used 200 million faces from 80 million IDs to train its model. Currently, the best way to collect facial training data is through web scraping, but scraped data generally contains significant noise, requiring considerable effort to clean it later—a huge undertaking for those wanting to train models quickly. Another approach is to use publicly available academic facial datasets. By increasing the amount of data (i.e., images) in publicly available facial datasets, more training data can be obtained, which is beneficial for improving the performance of facial recognition models.
[0048] Currently, the simplest algorithm for increasing training data is to iterate through all data in the dataset for each new query data (i.e., training images), calculate the similarity or distance between the query data and all elements, and then accurately return the Top K similar (i.e., the data with the highest similarity). This method is effective when there are relatively few image identifiers (e.g., IDs) in the face dataset. However, when there are many image identifiers (e.g., IDs) in the face dataset, or when there are many datasets to be merged, such as tens of millions of images, the time consumption becomes enormous.
[0049] Based on this, embodiments of this application provide an image recognition method and an image dataset merging method. By classifying the multiple datasets to be merged (e.g., classifying them according to image attributes), each dataset is divided into multiple data subsets of different types. The data subsets of the same type in the multiple datasets are merged according to the target features of the image identifiers. This reduces the time consumed logarithmically and can solve the problem of huge time consumption when there are many datasets to be merged to a certain extent.
[0050] Figure 1 A flowchart illustrating an exemplary image dataset merging method according to an embodiment of this application is shown.
[0051] like Figure 1 As shown in the figure, this application embodiment provides a method for merging image datasets, including:
[0052] In step S100, at least two datasets to be merged are obtained. Each dataset includes multiple images, each image includes an image identifier, and at least two images share the same image identifier. This can be understood as each dataset containing multiple different image identifiers, each image identifier corresponding to multiple images, and each image corresponding to only one unique image identifier. Typically, the image identifier can be in the form of a numerical sequence number (e.g., image ID), which can correspond to a specific person (e.g., Zhang San).
[0053] The images can be from various recognition scenarios, such as face images, gesture images, and flower images. In some embodiments, the images are face images, and the dataset to be merged can be a publicly available and commonly used academic face dataset. Typically, the number of images corresponding to each image identifier can be in the hundreds or thousands. In practical applications, the data format of the dataset can be as shown in Table 1 below. Here, IDs can be understood as image identifiers, and the number of images is the sum of the number of images corresponding to all image identifiers in the dataset.
[0054] Table 1 shows the data format of the datasets to be merged.
[0055] Public dataset name IDs Number of images CASIAWebface 10575 494414 VGGFace2 8631 3141890 MS1M 93431 51790 Glint360 360232 17091657 Million Celebs 719613 22804708 WebFace260M 2059906 42474990
[0056] In step S200, each dataset is classified according to the attributes of the image, resulting in multiple data subsets to be merged corresponding to each classification. This can be understood as follows: after classifying the datasets, each dataset yields multiple data subsets. For multiple datasets, the number of data subsets to be merged corresponding to any one classification is the same as the number of datasets, i.e., at least two.
[0057] Typically, for face datasets, image attributes can include the face's gender and skin color. Gender can include male and female. Skin color can include a first skin tone (e.g., light tone), a second skin tone (e.g., dark tone), a third skin tone (e.g., yellow tone), and a fourth skin tone (e.g., dark tone). These classifications can be based on the face's gender and skin color, for example, including eight categories: male first skin color type, male second skin color type, male third skin color type, male fourth skin color type; female first skin color type, female second skin color type, female third skin color type, and female fourth skin color type. Figure 2a As shown.
[0058] In some embodiments, the classification of multiple datasets can be a tree-like classification, such as... Figure 2b As shown. For face dataset classification models, existing software can be used, such as Quividi's AI application and the Android app AgeBot, to identify gender, male skin color, and female skin color. Existing classification models, such as the Gaussian Mixture Model (GMM), can be used, pre-trained for first, second, third, and fourth skin color recognition, to identify facial skin color, etc.
[0059] In some embodiments, classifying each dataset may include inputting multiple images from the dataset into a pre-trained classification model and outputting the classification of the multiple images. For face datasets, a single classification model may also be used, such as one capable of recognizing multiple attributes, including gender and skin color attributes. In applications, such as... Figure 3 and Figure 4As shown, a pre-trained classification model can be used to identify two attributes: gender and skin color, directly obtaining the performance and skin color classification of the face image. This classification model is a multi-scale, multi-label face attribute model with good accuracy. During training, each training face image simultaneously possesses two attributes: gender and ethnicity. Depthwise separable convolutions can be used in this classification model to achieve fast attribute classification. Thus, at least two datasets to be merged can be classified separately, dividing each dataset into multiple subsets corresponding to the classification categories.
[0060] In some embodiments, classifying the dataset based on the attributes of the images may further include, in response to determining that at least two images with the same image identifier are classified into multiple categories, determining the category with the highest weight as the category of the at least two images. This can be understood as follows: in each dataset to be merged, for at least two images with the same image identifier (i.e., for all images corresponding to the same image identifier), when the classification results obtained by the classification model for the at least two images are not unique and multiple classification results exist, the category of the at least two images can be determined by voting. Specifically, as... Figure 5 As shown, the weight (proportion) of each classification result in multiple classification results can be calculated separately, and the classification with the highest weight is determined as the classification of the at least two images. This improves the accuracy of image classification and avoids the influence of factors such as dim or bright lighting during shooting, or the effects of facial makeup on actual skin tone.
[0061] In step S300, at least two target data subsets corresponding to the target category among the multiple categories are determined. Here, the target category can be understood as any one of the multiple categories, such as male first skin color type, male second skin color type, male third skin color type, or male fourth skin color type, or female first skin color type, female second skin color type, female third skin color type, or female fourth skin color type. The at least two target data subsets corresponding to the target category among the multiple categories can be understood as all data subsets corresponding to one of the multiple categories. For example, at least two data subsets corresponding to female first skin color. Through this step, for each category, at least two target data subsets corresponding to that category can be categorized, that is, categorized by the target category within the category.
[0062] In step S400, the target features of at least two images with the same image identifier in each of the target data subsets are calculated to obtain the target features corresponding to each image identifier. This can be understood as a calculation performed on each image identifier in a specific data subset corresponding to one of the multiple categories determined in the preceding steps. Through this step, for each category, the features of all images corresponding to all image identifiers in at least two target subsets (i.e., all subsets) corresponding to that category can be calculated separately, thus obtaining the target features corresponding to all image identifiers.
[0063] In some embodiments, the target feature can be an average feature. That is, the target feature of at least two images with the same image identifier in each of the target data subsets is the average feature of at least two images with the same image identifier in each of the target data subsets. This can be understood as calculating the average feature of all images corresponding to each image identifier in the data subset to be merged, using this average feature to represent the comprehensive feature of that image identifier. Figure 6 As shown, a typical approach using a face recognition model, such as the InsightFace model, is to extract features (e.g., f1, f2, f3, f4, and f5) from each image in all data subsets (at least two target data subsets) within the target category (i.e., the same category). Then, the average feature (e.g., f) of all images with the same image identifier is calculated. This average feature is then used to represent the comprehensive facial features of this ID. This feature can typically be multi-dimensional and can take the form of a feature vector, etc. Since the number of images with the same image identifier is usually in the hundreds or thousands, using the average feature of images with the same image identifier as the target feature of the image corresponding to that identifier can effectively represent the features of the image of a specific person corresponding to that image identifier.
[0064] In step S500, the images in the at least two target data subsets are merged according to the target features corresponding to each image identifier to obtain a merged target data subset. In this step, for at least two data subsets corresponding to the same category obtained after classifying the at least two datasets to be merged, the target features corresponding to all image identifiers in all data subsets to be merged are merged based on similarity. For each group of two or more image identifiers that meet the similarity criteria, all images corresponding to those two or more image identifiers are merged to form a merged target data subset. In some embodiments, remaining image identifiers that do not meet the similarity criteria can be directly added to the merged target data subset to make the merged target data subset have more image identifiers, thereby improving the accuracy of the obtained model after subsequent training.
[0065] like Figure 7 As shown, in some embodiments, merging images in the at least two target data subsets based on the target features corresponding to each image identifier typically includes:
[0066] S510, for all image identifiers corresponding to the target features in the at least two target data subsets, calculate the similarity pairwise. Typically, Euclidean distance or cosine distance can be used to calculate the similarity.
[0067] S520, merge the images corresponding to at least two image identifiers that meet the similarity condition, such as... Figure 8 As shown in the figure. The similarity condition is satisfied when the similarity is greater than a preset threshold.
[0068] In this way, by merging images based on the average features of all images corresponding to image identifiers, images corresponding to at least two image identifiers that meet the similarity criteria are merged. Compared to traversing every image in the dataset, this reduces the amount of similarity calculation logarithmically, significantly lowering the time complexity of data merging. Furthermore, after merging similar images, multi-dimensional face images can be obtained for the same face, thereby improving the accuracy of recognition through subsequent model training.
[0069] In some embodiments, the method may further include renaming at least two image identifiers in the resulting target data subset to the same image identifier, thereby obtaining a merged image identifier. That is, the image identifier obtained after merging multiple image identifiers can be named after any one of the image identifiers. Alternatively, after obtaining the merged target data subset, all image identifiers in the target data subset may be uniformly renamed (e.g., numbered).
[0070] In some embodiments, the method may further include calculating the target features of all images after image merging to obtain the target features corresponding to the merged image identifier. For example, calculating the average features of all images after image merging to obtain the average features corresponding to the merged image identifier.
[0071] In step S600, the image dataset is determined based on the merged target data subset. Here, the target data subset corresponding to a single category obtained after merging can be used as the image dataset obtained after merging at least two datasets to be merged. Alternatively, all merged target data subsets corresponding to all categories obtained after merging can be used as the image dataset obtained after merging at least two datasets to be merged, resulting in a richer image dataset.
[0072] In some embodiments, the method may further include: training an image model using the image dataset. For example, the Iresnt126 network can be trained. After merging, the number of images corresponding to the same image identifier increases, allowing for more images of the same face from different shooting angles, distances, or lighting conditions, thereby improving the recognition accuracy of the image model obtained after training.
[0073] Application examples
[0074] The image dataset obtained after merging using the image dataset merging method provided in this application embodiment, along with the original data before merging, are used to train the Iresnt126 network to obtain the corresponding face recognition model. Three face test sets for surveillance scenarios are established: RGB, IR (infrared), and a masked dataset. Each dataset contains 3000 pairs of positive samples (two faces of the same person) and 3000 pairs of negative samples (two faces of different people). The construction method can refer to the classic face dataset LFW.
[0075] The accurate results of the obtained face recognition model are shown in Table 2 below.
[0076] Table 2. Accuracy of face recognition models trained on different datasets
[0077]
[0078] As can be seen, the image dataset obtained by the image dataset merging method in this application embodiment reduces the time complexity of the merged dataset logarithmically compared to merging individual images. For example, the time complexity is reduced from O(n) to O(ln(n)). Furthermore, by increasing the number of images corresponding to the same image identifier, and by incorporating factors such as the image shooting angle, shooting distance, lighting type, age of the face corresponding to the image identifier, and shooting scene, the trained face recognition model achieves good accuracy. In addition, it can easily and quickly merge currently publicly available face datasets into a large face training set containing more IDs, solving the problems of difficult data collection and limited training data when training face models. Moreover, using a larger training set can significantly improve the performance of the face recognition algorithm.
[0079] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0080] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0081] Based on the same inventive concept, this application also provides an image recognition method. Figure 9 A flowchart illustrating an exemplary image recognition method according to an embodiment of this application is shown.
[0082] like Figure 9 As shown, the image recognition method provided in this application embodiment may include:
[0083] In step S620, an image to be recognized is acquired. This image can be from various recognition scenarios, such as a face image, a gesture image, and a flower image. In some embodiments, the image is a face image. That is, the image to be recognized can be a face image.
[0084] In step S640, the classification of the image to be identified is determined. Typically, determining the classification of the image to be identified may include: inputting the image to be identified into a pre-trained classification model and outputting the classification of the image to be identified. For the classification model of face images, the specific details can be as described in step S200 of the aforementioned image dataset merging method, and will not be repeated here. For face images, the classification of the image to be identified can be a classification corresponding to the gender and skin color of the face. Gender can include male and female. Skin color can include a first skin color (e.g., light tone), a second skin color (e.g., dark tone), a third skin color (e.g., yellow tone), and a fourth skin color (e.g., dark tone). For example, the classification of the image to be identified can be male first skin color type, male second skin color type, male third skin color type, male fourth skin color type, female first skin color type, female second skin color type, female third skin color type, or female fourth skin color type.
[0085] In step S660, the target dataset corresponding to the classification of the image to be identified is determined. The target dataset is a subset of multiple subsets of the image dataset that corresponds to the classification of the image to be identified. This can be understood as follows: the image dataset has multiple subsets, each corresponding to an image classification, and the target subset is one of these subsets. The specific classification type of the image classification is the same as the specific classification type of the image to be identified. The classification type corresponding to the target dataset (i.e., the target subset) is the same as the classification type of the image to be identified. For face images, the multiple subsets of the image dataset correspond to multiple classification types, which are categorized according to the gender and skin color of the face. The specific content of the multiple classifications can be referred to in the aforementioned image dataset merging method, and will not be repeated here.
[0086] In some embodiments, the image dataset can be obtained using the aforementioned image dataset merging method.
[0087] Typically, the target dataset may include multiple data subsets, each of which corresponds to an image identifier. That is, each data subset contains all the images corresponding to each image identifier. In a face dataset, the image identifier can be in the form of a numerical sequence number (e.g., image ID), which can correspond to a specific person (e.g., Zhang San). Each data subset contains all the face images corresponding to each image identifier.
[0088] In some embodiments, obtaining the image recognition result based on the image to be recognized and the target dataset includes:
[0089] Calculate the similarity between the image to be identified and the multiple data subsets respectively.
[0090] Based on the similarity between the image to be identified and the plurality of data subsets, a target data subset is determined from the plurality of data subsets. Typically, the data subset with the highest similarity can be identified as the target data subset.
[0091] The image recognition result is obtained based on the image identifiers of the target data subset.
[0092] In some embodiments, each subset of data includes multiple images; calculating the similarity between the image to be identified and the multiple subsets of data includes:
[0093] The features of the image to be identified and the average features of multiple images in each subset of data are calculated. The image features can be extracted using a face recognition model, such as the InsightFace model.
[0094] Calculate the similarity between the features of the image to be identified and the average features of each of the data subsets.
[0095] In another embodiment, the average features of each data subset in the image dataset have been pre-calculated and can be directly obtained. The step of calculating the similarity between the image to be identified and the plurality of data subsets includes:
[0096] Calculate the features of the image to be identified;
[0097] Calculate the similarity between the features of the image to be identified and the average features of each of the data subsets.
[0098] In some embodiments, the method may further include sorting the obtained similarities to obtain the highest similarity and determining it as the highest similarity.
[0099] In step S680, an image recognition result is obtained based on the image to be recognized and the target dataset. This can be understood as determining the image identifier corresponding to the target dataset as the image identifier corresponding to the image to be recognized. In face data, this means obtaining the image identifier corresponding to the face image, such as an image ID.
[0100] The image recognition method provided in this application can identify the image identifiers corresponding to the images to be merged in an image dataset. By comparing the image to be recognized with a target subset in the image dataset that corresponds to the category of the image to be recognized, the image recognition result is obtained, which can quickly and accurately obtain the recognition result of the image to be recognized. Compared with directly traversing all images in the image dataset and calculating the similarity between the image to be recognized and all images in the dataset, the image recognition speed can be significantly improved.
[0101] Based on the same inventive concept, corresponding to the image dataset merging method of any of the above embodiments, this application also provides an image dataset merging device.
[0102] refer to Figure 10 The device for merging the image dataset includes:
[0103] The dataset acquisition module 710 is used to acquire at least two datasets to be merged; each dataset includes multiple images, each image includes an image identifier, and the multiple images include at least two images with the same image identifier;
[0104] The dataset classification module 720 is used to classify the dataset according to the attributes of the image, and obtain multiple data subsets to be merged corresponding to multiple classifications respectively;
[0105] The target data subset determination module 730 determines at least two target data subsets corresponding to the target category among the multiple categories;
[0106] The target feature calculation module 740 is used to calculate the target features of at least two images with the same image identifier in each of the target data subsets, so as to obtain the target features corresponding to each image identifier;
[0107] The merging module 750 is used to merge the images in the at least two target data subsets according to the target features corresponding to each image identifier, so as to obtain a merged target data subset;
[0108] Image dataset determination module 760 is used to determine multiple merged target data subsets as the image dataset.
[0109] In some embodiments, merging images in the at least two target data subsets based on the target features corresponding to each image identifier includes:
[0110] For all image identifiers corresponding to the target features in the at least two target data subsets, calculate the similarity between each pair of them;
[0111] Merge images corresponding to at least two image identifiers that meet the similarity criteria.
[0112] In some embodiments, a renaming module is also included, which is used to rename the at least two image identifiers to the same image identifier, so as to obtain a merged image identifier.
[0113] In some embodiments, classifying the dataset based on the attributes of the image includes:
[0114] The dataset contains multiple images, each of which is input into a pre-trained classification model. The model then outputs the classification of the images.
[0115] In some embodiments, classifying the dataset based on the attributes of the image further includes:
[0116] In response to determining that at least two images with the same image identifier are classified into multiple categories, the category with the highest weight is determined as the category of the at least two images.
[0117] In some embodiments, a training module is also included for training an image model using the image dataset.
[0118] In some embodiments, the target feature is an average feature; the similarity satisfies the similarity condition including similarity greater than a preset threshold; the dataset is a face dataset, and the multiple classifications are classified according to the gender and skin color corresponding to the face.
[0119] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0120] The apparatus of the above embodiments is used to implement the merging method of the corresponding image dataset in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0121] Based on the same inventive concept, corresponding to the image recognition methods of any of the above embodiments, this application also provides an image recognition device.
[0122] refer to Figure 11 An image recognition device, comprising:
[0123] Image acquisition module 810 is used to acquire the image to be recognized;
[0124] Image classification module 820 is used to determine the classification of the image to be identified;
[0125] The target dataset determination module 830 is used to determine the target dataset corresponding to the classification of the image to be identified. The target dataset is a target subset corresponding to the classification of the image to be identified from multiple subsets of the image dataset. The target dataset includes multiple image identifiers.
[0126] The image recognition module 840 obtains the image recognition result based on the image to be recognized and the target dataset.
[0127] In some embodiments, the target dataset includes multiple data subsets, each data subset corresponding to an image identifier; obtaining the image recognition result based on the image to be recognized and the target dataset includes:
[0128] Calculate the similarity between the image to be identified and each of the multiple data subsets;
[0129] Based on the similarity between the image to be identified and the plurality of data subsets, the target data subset is determined from the plurality of data subsets;
[0130] The image recognition result is obtained based on the image identifiers of the target data subset.
[0131] In some embodiments, each subset of data includes multiple images; calculating the similarity between the image to be identified and the multiple subsets of data includes:
[0132] Calculate the features of the image to be identified and the average features of multiple images in each subset of data;
[0133] Calculate the similarity between the features of the image to be identified and the average features of each of the data subsets.
[0134] In some embodiments, determining the classification of the image to be identified includes:
[0135] The image to be identified is input into a pre-trained classification model, which outputs the classification of the image to be identified.
[0136] In some embodiments, the image to be identified is a face image, and multiple subsets of the image dataset correspond to multiple categories, which are classified according to the gender and skin color of the face.
[0137] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0138] The apparatus of the above embodiments is used to implement the corresponding image recognition method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0139] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the image dataset merging method or image recognition method described in any of the above embodiments.
[0140] Figure 12 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0141] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0142] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0143] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0144] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0145] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0146] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0147] The electronic devices described in the above embodiments are used to implement the image dataset merging method or image recognition method described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0148] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the image dataset merging method or image recognition method as described in any of the above embodiments.
[0149] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0150] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the image dataset merging method or image recognition method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0151] Based on the same inventive concept, corresponding to the image dataset merging method or image recognition method described in any of the above embodiments, this disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the image dataset merging method or image recognition method. Corresponding to the execution entity for each step in each embodiment of the image dataset merging method or image recognition method, the processor executing the corresponding step may belong to the corresponding execution entity.
[0152] The computer program product of the above embodiments is used to cause the computer and / or the processor to execute the image dataset merging method or image recognition method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0153] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0154] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0155] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0156] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. An image recognition method, characterized in that, include: Acquire the image to be recognized; Determine the classification of the image to be identified; Determine the target dataset corresponding to the classification of the image to be identified, wherein the target dataset is a subset of the image dataset that corresponds to the classification of the image to be identified; Based on the image to be identified and the target dataset, the image recognition result is obtained; The target dataset includes multiple data subsets, each data subset corresponding to an image identifier; obtaining the image recognition result based on the image to be recognized and the target dataset includes: Calculate the similarity between the image to be identified and each of the multiple data subsets; Based on the similarity between the image to be identified and the plurality of data subsets, the target data subset is determined from the plurality of data subsets; The image recognition result is obtained based on the image identifiers of the target data subset.
2. The image recognition method according to claim 1, characterized in that, Each of the data subsets includes multiple images; calculating the similarity between the image to be identified and the multiple data subsets includes: Calculate the features of the image to be identified and the average features of multiple images in each subset of data; Calculate the similarity between the features of the image to be identified and the average features of each of the data subsets.
3. The image recognition method according to claim 1, characterized in that, Determining the classification of the image to be identified includes: The image to be identified is input into a pre-trained classification model, which outputs the classification of the image to be identified.
4. The image recognition method according to claim 1, characterized in that, The image to be identified is a face image, and multiple subsets of the image dataset correspond to multiple categories, which are classified according to the gender and skin color of the face.
5. A method for merging image datasets, characterized in that, include: Obtain at least two datasets to be merged; each dataset includes multiple images, and each image includes an image identifier; Based on the attributes of the images, each dataset is classified to obtain multiple subsets of data to be merged, each corresponding to a different classification. Determine at least two subsets of target data corresponding to the target category among the multiple categories; Calculate the target features of at least two images with the same image identifier in each of the target data subsets to obtain the target features corresponding to each image identifier; Based on the target features corresponding to each image identifier, the images in the at least two target data subsets are merged to obtain a merged target data subset; The image dataset is determined based on the merged target data subset.
6. The method for merging image datasets according to claim 5, characterized in that, Merging images in the at least two target data subsets based on the target features corresponding to each image identifier includes: For all image identifiers corresponding to the target features in the at least two target data subsets, calculate the similarity between each pair of them; Merge images corresponding to at least two image identifiers that meet the similarity criteria.
7. The method for merging image datasets according to claim 6, characterized in that, The method further includes: The at least two image identifiers are renamed to the same image identifier to obtain the merged image identifier.
8. The method for merging image datasets according to claim 5, characterized in that, The step of classifying each dataset according to the attributes of the image includes: Multiple images from each dataset are input into a pre-trained classification model, which outputs the classification of the multiple images.
9. The method for merging image datasets according to claim 8, characterized in that, The step of classifying each dataset according to the attributes of the image further includes: In response to determining that at least two images with the same image identifier are classified into multiple categories, the category with the highest weight is determined as the category of the at least two images.
10. The method for merging image datasets according to claim 6, characterized in that, The target feature is the average feature; The similarity condition includes a similarity greater than a preset threshold; the dataset is a face dataset, and the multiple classifications are based on the gender and skin color corresponding to the face.
11. An image recognition device, characterized in that, include: The image acquisition module is used to acquire the image to be recognized; An image classification module is used to determine the classification of the image to be identified; The target dataset determination module is used to determine the target dataset corresponding to the classification of the image to be identified, wherein the target dataset is a subset of the image dataset that corresponds to the classification of the image to be identified; The image recognition module obtains the image recognition result based on the image to be recognized and the target dataset; The target dataset includes multiple data subsets, each data subset corresponding to an image identifier; obtaining the image recognition result based on the image to be recognized and the target dataset includes: Calculate the similarity between the image to be identified and each of the multiple data subsets; Based on the similarity between the image to be identified and the plurality of data subsets, the target data subset is determined from the plurality of data subsets; The image recognition result is obtained based on the image identifiers of the target data subset.
12. A device for merging image datasets, characterized in that, include: A dataset acquisition module is used to acquire at least two datasets to be merged; each dataset includes multiple images, and each image includes an image identifier; The dataset classification module is used to classify each dataset according to the attributes of the image, and obtain multiple data subsets to be merged corresponding to multiple classifications respectively; The target data subset determination module is used to determine at least two target data subsets corresponding to the target category among the multiple categories; The target feature calculation module is used to calculate the target features of at least two images with the same image identifier in each of the target data subsets, so as to obtain the target features corresponding to each image identifier; The merging module is used to merge the images in the at least two target data subsets according to the target features corresponding to each image identifier, so as to obtain a merged target data subset; The image dataset determination module is used to determine multiple merged target data subsets as the image dataset.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in any one of claims 1 to 4 or the method as claimed in any one of claims 5 to 10.
14. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method of any one of claims 1 to 4 or the method of any one of claims 5 to 10.