Role recognition method and related device
By extracting multiple feature data in the human body detection area in role recognition for clustering processing, the problem of low face recognition accuracy in the prior art is solved, and a higher character recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111653903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The existing character recognition methods mainly rely on face recognition, resulting in low recognition accuracy and it is difficult to distinguish different roles played by the same actor.
By determining the images in the human detection area, extracting facial feature data and clothing and accessories feature data, and using a pre-trained character feature extraction model for clustering to form clustering clusters to identify characters.
It improves the accuracy of character recognition and can better distinguish different roles played by the same actor.
Smart Images

Figure CN114333060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and in particular, to a method for character recognition and related devices. Background Art
[0002] In order to analyze a video, it is usually necessary to analyze the characters in the video, determine the characters corresponding to different images in the video, and then determine the time points when each character appears.
[0003] Currently, character recognition is usually based on face recognition technology. Since the face region contains less feature information and face features can be affected by the makeup and hairstyle of the character and change. The current character recognition method only recognizes the face in the image, and the recognized features can only represent the face information in the image, making it difficult to distinguish different characters played by the same actor, and thus resulting in a low accuracy of character recognition. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method for character recognition and related devices to improve the accuracy of character recognition. The specific technical solutions are as follows:
[0005] In the first aspect of the present invention, first, a method for character recognition is provided, including:
[0006] Determine the human detection region corresponding to each of the multiple images to be recognized;
[0007] Determine a partial image within the human detection region corresponding to the image to be recognized as a human image;
[0008] Input the human image into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data;
[0009] Perform clustering processing on the multiple human images based on the feature data to obtain clustering clusters; each clustering cluster includes at least one of the human images corresponding to a character category.
[0010] Optionally, after performing clustering processing on the multiple human images based on the feature data to obtain clustering clusters, the method further includes:
[0011] Determine the character category corresponding to each human image based on the character category corresponding to the clustering cluster where each human image is located;
[0012] Determine the character category of the image to be recognized based on the character category corresponding to the human image.
[0013] Optionally, determining a partial image within the human detection region corresponding to the image to be recognized as a human image includes:
[0014] Determining a feature detection region corresponding to each of the images to be recognized among the multiple images to be recognized; the feature detection region includes at least one of a face feature detection region, a clothing feature detection region, and an accessory feature detection region; wherein, the image within the face feature detection region is a face feature image, the image within the clothing feature detection region is a clothing feature image, and the image within the accessory feature detection region is an accessory feature image;
[0015] When the overlap degree between the human detection region and the feature detection region is greater than a first preset value, determining a partial image within the human detection region corresponding to the image to be recognized as a human image.
[0016] Optionally, the range of the first preset value is 0.5 to 0.7.
[0017] Optionally, determining a feature detection region corresponding to each of the images to be recognized among the multiple images to be recognized includes:
[0018] Inputting the multiple images to be recognized into a pre-trained target image detection model for image detection to obtain a feature detection region corresponding to each of the images to be recognized;
[0019] wherein, the target image detection model includes at least one of the following: a face image detection model, a clothing image detection model, and an accessory image detection model.
[0020] Optionally, performing clustering processing on the multiple human images based on the feature data to obtain clustering clusters includes:
[0021] Performing M clustering operations on the multiple human images based on the feature data to obtain an Mth target clustering cluster;
[0022] wherein, the (N + 1)th clustering operation includes:
[0023] Calculating the similarity between a first clustering cluster and a second clustering cluster based on first feature data and second feature data, where the first feature data is the feature data corresponding to the first clustering cluster, and the second feature data is the feature data corresponding to the second clustering cluster; the first clustering cluster and the second clustering cluster are any two Nth clustering clusters;
[0024] When the similarity between the first clustering cluster and the second clustering cluster is less than a second preset value, merging the first clustering cluster and the second clustering cluster into one Nth clustering cluster; N and M are both positive integers greater than 1, and N is less than or equal to M.
[0025] Optionally, before inputting the human body image into a pre-trained character feature extraction model for feature extraction to obtain feature data, the method further includes:
[0026] Iteratively training a character feature extraction model to be trained with a plurality of pre-acquired sample images to obtain the trained character feature extraction model;
[0027] Wherein, the Nth iterative training includes:
[0028] Inputting a plurality of pre-acquired sample images into the character feature extraction model to be trained for the Nth character feature extraction to obtain Nth feature data;
[0029] Performing clustering processing on the plurality of sample images based on the Nth feature data to obtain at least two sample clusters;
[0030] Judging whether a loss value meets a loss convergence condition;
[0031] In the case that the loss value does not meet the loss convergence condition, adjusting parameters of the character feature extraction model to be trained based on the loss value; in the case that the loss value meets the loss convergence condition, determining the currently trained character feature extraction model to be trained as the pre-trained character feature extraction model.
[0032] Optionally, the iteratively training the character feature extraction model to be trained with a plurality of pre-acquired sample images includes:
[0033] Obtaining a plurality of sample images;
[0034] Performing preset processing on the plurality of sample images to obtain augmented sample images;
[0035] Iteratively training the character feature extraction model to be trained with the plurality of sample images and the augmented sample images.
[0036] In a second aspect of the implementation of the present invention, there is also provided a character recognition device, including:
[0037] A first determination module, configured to determine a human body detection area corresponding to each of a plurality of images to be recognized;
[0038] A second determination module, configured to determine a partial image within the human body detection area corresponding to the image to be recognized as a human body image;
[0039] A feature extraction module for inputting the human body image into a pre-trained character feature extraction model to extract features and obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data;
[0040] A clustering processing module for performing clustering processing on multiple human body images based on the feature data to obtain clustering clusters; each clustering cluster includes at least one of the human body images corresponding to a character category.
[0041] In a third aspect of the embodiments of the present invention, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0042] The memory is used to store programs;
[0043] The processor is configured to implement the method steps as described in the first aspect when executing the programs stored on the memory.
[0044] In a fourth aspect of the embodiments of the present invention, a readable storage medium is further provided, on which a program is stored, and when the program is executed by a processor, the method as described in the first aspect is implemented.
[0045] In the embodiments of the present invention, a human body detection area corresponding to each of the multiple images to be recognized is determined; a partial image within the human body detection area corresponding to the image to be recognized is determined as a human body image; the human body image is input into a pre-trained character feature extraction model to extract features, and feature data is obtained. The feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data. Since the feature data includes not only face feature data but also clothing feature data and / or accessory feature data, when recognizing a character, recognition can be performed based on multiple aspects of information such as the face, clothing, and / or accessories, improving the accuracy of character recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.
[0047] Figure 1 It is one of the flow diagrams of a character recognition method in an embodiment of the present invention;
[0048] Figure 2 It is another flow diagram of a character recognition method in an embodiment of the present invention;
[0049] Figure 3 It is a schematic flowchart of the Nth iterative training of a role feature extraction model to be trained in an embodiment of the present invention;
[0050] Figure 4 It is a schematic structural diagram of a role recognition device in an embodiment of the present invention;
[0051] Figure 5 It is a schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention.
[0053] As Figure 1 shown, an embodiment of the present invention provides a role recognition method, including:
[0054] Step 101, determining a human detection area corresponding to each of the multiple images to be recognized in the multiple images to be recognized.
[0055] It should be understood that the specific method for determining the human detection area corresponding to each of the multiple images to be recognized is not limited herein. For example, in some embodiments, the position of the human area in each of the images to be recognized is detected to obtain a corresponding human detection frame, and the area within the human detection frame is the human detection area.
[0056] Step 102, determining a partial image within the human detection area corresponding to the image to be recognized as a human image.
[0057] It should be understood that determining a partial image within the human detection area corresponding to the image to be recognized as a human image can be understood as intercepting the image within the human detection area to obtain the human image. In step 101, the human detection area corresponding to each of the images to be recognized is determined. Therefore, for each of the images to be recognized, the image within its corresponding human detection area can be determined as a human image, and multiple human images can be obtained.
[0058] Optionally, in some embodiments, step 102 includes:
[0059] Determining a feature detection area corresponding to each of the multiple images to be recognized; the feature detection area includes at least one of a face feature detection area, a clothing feature detection area, and an accessory feature detection area; wherein, the image within the face feature detection area is a face feature image, the image within the clothing feature detection area is a clothing feature image, and the image within the accessory feature detection area is an accessory feature image;
[0060] When the overlap degree (Intersection over Union, IOU) between the human body detection area and the feature detection area is greater than a first preset value, a partial image within the human body detection area corresponding to the image to be recognized is determined as a human body image.
[0061] It should be understood that, in some embodiments, the feature detection area includes a face feature detection area. In some other embodiments, the feature detection area includes a clothing feature detection area. In some other embodiments, the feature detection area includes an accessory feature detection area. In some other embodiments, the feature detection area includes a face feature detection area and a clothing feature detection area. In some other embodiments, the feature detection area includes a face feature detection area and an accessory feature detection area. In some other embodiments, the feature detection area includes a clothing feature detection area and an accessory feature detection area. In some other embodiments, the feature detection area includes a face feature detection area, a clothing feature detection area, and an accessory feature detection area.
[0062] It should be understood that, in some embodiments, the feature detection area includes two or three of the face feature detection area, the clothing feature detection area, and the accessory feature detection area. In this embodiment, the IOU between the human body detection area and the feature detection area being greater than the first preset value can be understood as the IOU between the human body detection area and each feature detection area being greater than the first preset value. The IOU between the human body detection area and the feature detection area being greater than the first preset value can also be understood as the IOU between the human body detection area and any one of the feature detection areas being greater than the first preset value.
[0063] It should be understood that, in some embodiments, the IOU between the human body detection area and the feature detection area is as follows:
[0064]
[0065] It should be understood that A can be understood as the area of the human body detection area, and B can be understood as the area of the feature detection area; then A∩B can be understood as the intersection area of A and B, and A∪B can be understood as the union area of A and B.
[0066] It should be understood that the value range of the first preset value is not limited here. In actual implementation, the value of the first preset value can be adjusted according to actual needs. Optionally, in some embodiments, the range of the first preset value is 0.5 to 0.7. Further, in some embodiments, the first preset value is 0.6.
[0067] It should be understood that the validity of the human body detection area can be verified by the IOU between the human body detection area and the feature detection area. Hereinafter, an example in which the feature detection area includes a clothing feature detection area will be described. For the to-be-recognized image corresponding to the human body feature detection area and the clothing feature detection area, when the IOU between the human body feature detection area and the clothing feature detection area is less than a first preset value, it can be considered that the image in the human body feature detection area only includes part of the clothing, and further, it can be considered that the image intercepted by the human body feature detection area is not a complete human body image. Therefore, the validity of this human body detection area is poor, that is, the human body image corresponding to this human body detection area is not sufficient to represent the role information.
[0068] It should be understood that when the IOU between the human body detection area and the feature detection area is less than or equal to the first preset value, the human body detection area corresponding to the to-be-recognized image can be re-determined, and the IOU between the newly determined human body detection area and the feature detection area can be calculated. In some other embodiments, when the IOU between the human body detection area and the feature detection area is less than or equal to the first preset value, the to-be-recognized image corresponding to the human body detection area can also be excluded.
[0069] In this embodiment, the feature detection area corresponding to each of the multiple to-be-recognized images is determined, and the validity of the human body detection area is verified by the IOU between the human body detection area and the feature detection area, which improves the reliability of the human body detection area. By excluding the human body images with poor validity, the result of role recognition is made more accurate.
[0070] It should be understood that the specific method for determining the feature detection area corresponding to each of the multiple to-be-recognized images is not limited herein. For example, in some embodiments, the method for obtaining the feature detection area corresponding to each of the multiple to-be-recognized images is the same as the method for determining the human body detection area corresponding to each of the multiple to-be-recognized images.
[0071] Optionally, in some embodiments, the determining the feature detection area corresponding to each of the multiple to-be-recognized images includes:
[0072] Inputting the multiple to-be-recognized images into a pre-trained target image detection model for image detection to obtain the feature detection area corresponding to each to-be-recognized image;
[0073] Wherein, the target image detection model includes at least one of the following: a face image detection model, a clothing image detection model, and an accessory image detection model.
[0074] It should be understood that the image detection model is pre-trained. By inputting multiple images to be recognized into the pre-trained target image detection model for image detection, a feature detection region corresponding to each image to be recognized can be obtained.
[0075] Among them, the target image detection model includes at least one of the following: a face image detection model, a clothing image detection model, and an accessory image detection model. When the target image detection model includes a face image detection model, by inputting the multiple images to be recognized into the pre-trained target image detection model, a face feature detection region corresponding to each image to be recognized can be obtained. When the target image detection model includes a clothing image detection model, by inputting the multiple images to be recognized into the pre-trained target image detection model, a clothing feature detection region corresponding to each image to be recognized can be obtained. When the target image detection model includes an accessory image detection model, by inputting the multiple images to be recognized into the pre-trained target image detection model, an accessory feature detection region corresponding to each image to be recognized can be obtained.
[0076] It should be understood that the specific method for performing image detection on the multiple images to be recognized is not limited herein. In some embodiments, the image detection method may be an object detection method, such as a method based on a deep learning algorithm.
[0077] Step 103: Input the human body image into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data.
[0078] It should be understood that the character feature extraction model is pre-trained, and the character feature extraction model has the ability to recognize and extract features for characterizing a character, where the features for characterizing a character include at least one of the following: face features, styling features, clothing features, accessory features, and background features. Therefore, through the character feature extraction model, the feature data corresponding to the human body image can be extracted. The feature data is used to characterize the human body image. In some embodiments, the feature data includes face feature data and clothing feature data. In other embodiments, the feature data includes face feature data and accessory feature data. In other embodiments, the feature data includes face feature data, clothing feature data, and accessory feature data.
[0079] Optionally, in some embodiments, before step 103, the method further includes:
[0080] Iteratively train the to-be-trained character feature extraction model using a plurality of pre-acquired sample images to obtain the trained character feature extraction model;
[0081] Among them, as Figure 3 shown, the Nth iterative training includes:
[0082] Input a plurality of pre-acquired sample images into the to-be-trained character feature extraction model for the Nth character feature extraction to obtain the Nth feature data;
[0083] Perform clustering processing on the plurality of sample images based on the Nth feature data to obtain at least two sample clusters;
[0084] Determine whether the loss value satisfies the loss convergence condition;
[0085] In the case where the loss value does not satisfy the loss convergence condition, adjust the parameters of the to-be-trained character feature extraction model based on the loss value; in the case where the loss value satisfies the loss convergence condition, determine the currently trained to-be-trained character feature extraction model as the pre-trained character feature extraction model.
[0086] It should be understood that the loss convergence condition is not limited here. For example, in some embodiments, the convergence condition is that the loss value is less than a threshold. In other embodiments, the loss convergence condition is that the number of iterative training reaches a preset value.
[0087] It should be understood that in this embodiment, the plurality of pre-acquired sample images can be images in an existing dataset. The existing dataset used can be any dataset. In this embodiment, the existing dataset used is the pedestrian dataset Market1051. Market1051 is annotated with a total of 1501 personal identity numbers (Identity document, ID) and 32668 training images. Among them, the training set has 751 people and contains 12936 training images; the test set has 750 people and contains 19732 training images.
[0088] It should be understood that in some embodiments, a preset number of training images can be randomly selected from Market1051, perform filter transformations on the above training image data, and the filter transformations include color transformation, adding border effects, and blur transformation, etc., and add the training images obtained by the filter transformation to the plurality of sample images. Through the above settings, the to-be-trained character feature extraction model can still correctly identify the same sample after the above filter transformation.
[0089] In this embodiment, the specific process of training the to-be-trained character feature extraction model using Market1501 is as follows:
[0090] First, use the to-be-trained character feature extraction model to extract features from the above 12,936 images, obtaining feature data corresponding to each image. Among them, the to-be-trained character feature extraction model can be any residual network model capable of performing image feature representation on images, such as a residual network model trained based on ImageNet.
[0091] Then, use an unsupervised clustering method to cluster the feature data, such as a density-based clustering method. Using the clustered person ID as the key, the average value of all features under this clustering center is used as the representative feature value and stored in a dictionary, thereby establishing a person feature library; use a loss function to detect the result output by the to-be-trained character feature extraction model to determine the loss value; if the loss value meets the preset convergence condition, then determine the currently trained to-be-trained character feature extraction model as the character feature extraction model. In specific implementation, the image features extracted on the convolutional layer of the trained character feature extraction model can be used as character clustering features.
[0092] In specific implementation, the character feature extraction model can extract features from images. The image features extracted by the character feature extraction model obtained through the above training process can be used to represent the character information in the images. Use the character feature extraction model to extract features from the human body image to obtain the feature data. Among them, the feature data can be used to represent the human body information in the corresponding human body image. The more similar the feature data of any two human body images are, the more likely the two human body images represent the same character. Based on the feature data, perform clustering processing on the human body images, and the human body images that may represent the same character can be clustered into the same cluster. The character category corresponding to each cluster where the human body image is located can be used to determine the character category corresponding to each to-be-identified image. In specific implementation, one or more videos can be analyzed to obtain the time points when the same character category appears in one or more videos.
[0093] In this embodiment, an existing database is used to iteratively train the to-be-trained character feature extraction model to obtain the character feature extraction model. Through the above settings, during the training process of the to-be-trained character feature extraction model, there is no need to manually label new business data, which improves the convenience of training the to-be-trained character feature extraction model. Optionally, in some embodiments, the iterative training of the to-be-trained character feature extraction model using a plurality of pre-acquired sample images includes:
[0094] Obtain a plurality of sample images;
[0095] Perform preset processing on multiple sample images to obtain augmented sample images;
[0096] Use the multiple sample images and the augmented sample images to iteratively train the role feature extraction model to be trained.
[0097] It should be understood that the specific manner of performing preset processing on the sample images is not limited herein. For example, in some embodiments, performing preset processing on the sample images includes performing color transformation on the sample images. In some other embodiments, performing preset processing on the sample images further includes performing blurring transformation on the sample images. In some other embodiments, performing preset processing on the sample images further includes adding a border effect to the sample images. In some other embodiments, performing preset processing on the sample images further includes performing cropping on the sample images. In some other embodiments, performing preset processing on the sample images further includes performing filter transformation on the sample images. In some other embodiments, performing preset processing on the sample images further includes adding various subtitles to the sample images. It should be understood that after performing preset processing on the sample images to obtain augmented sample images, the role category corresponding to the augmented sample image obtained by performing preset processing on any one of the sample images is the same as that of this sample image.
[0098] It should be understood that the number of the sample images is multiple. In some embodiments, preset processing can be performed on each of the sample images, and for any one of the sample images, multiple preset processes can be performed to obtain multiple corresponding augmented sample images. In some other embodiments, preset processing can be used for a preset number of the sample images, where the preset number can be adjusted according to actual needs. For example, the preset number is one-tenth of the total number of the sample images.
[0099] In the embodiment of the present invention, by performing preset processing on the sample images to obtain the augmented sample images, and using the multiple sample images and the augmented sample images to iteratively train the role feature extraction model to be trained. Through the above method, on the one hand, the number of training samples used for training the role feature extraction model to be trained can be increased, and on the other hand, the role feature extraction model to be trained can still correctly identify the same sample image after the above preset processing. The role feature extraction model trained by the method provided in this embodiment can be used for searching for roles under different filters and editing styles, improving the recognition accuracy of the role feature extraction model, and further enhancing the recognition robustness of the role feature extraction model for roles from different video sources.
[0100] It should be understood that in step 103, multiple said human body images are all input into a pre-trained character feature extraction model for feature extraction. Therefore, the said feature data includes sub-feature data corresponding to each said human body image.
[0101] Step 104, perform clustering processing on multiple said human body images based on the said feature data to obtain clustering clusters; each said clustering cluster includes at least one of said human body images corresponding to a character category.
[0102] It should be understood that the number of said clustering clusters is at least two, and each said clustering cluster includes at least one of said human body images corresponding to a character category.
[0103] It should be understood that the specific manner of performing clustering processing on multiple said human body images based on the said feature data is not limited herein. For example, in some embodiments, performing clustering processing on multiple said human body images based on the said feature data to obtain clustering clusters can be understood as performing clustering processing on multiple said human body images based on the said feature data using a density-based clustering method to obtain clustering clusters.
[0104] In some other embodiments, performing clustering processing on multiple said human body images based on the said feature data to obtain clustering clusters can be understood as performing clustering processing on multiple said human body images based on the said feature data using a grid-based clustering method to obtain clustering clusters.
[0105] In some other embodiments, performing clustering processing on multiple said human body images based on the said feature data to obtain clustering clusters can be understood as performing clustering processing on multiple said human body images based on the said feature data using a bottom-up hierarchical clustering method to obtain clustering clusters.
[0106] Optionally, in some embodiments, step 104 includes:
[0107] Perform M clustering operations on multiple said human body images based on the said feature data to obtain the Mth target clustering cluster;
[0108] Wherein, the (N + 1)th clustering operation includes:
[0109] Calculate the similarity between a first clustering cluster and a second clustering cluster based on first feature data and second feature data, where the first feature data is the feature data corresponding to the first clustering cluster, and the second feature data is the feature data corresponding to the second clustering cluster; the first clustering cluster and the second clustering cluster are any two Nth clustering clusters;
[0110] When the similarity between the first clustering cluster and the second clustering cluster is less than a second preset value, merge the first clustering cluster and the second clustering cluster into an Nth clustering cluster; both N and M are positive integers greater than 1, and N is less than or equal to M.
[0111] It should be understood that when the similarity between the first clustering cluster and the second clustering cluster is greater than or equal to the second preset value, the first clustering cluster and the second clustering cluster are not merged.
[0112] It should be understood that in the case of performing the first clustering operation, the first clustering cluster and the second clustering cluster can be understood as any two human body images among the multiple human body images. In the case of performing the (N + 1)th clustering operation, the first clustering cluster can be understood as the human body images that have not been clustered in the Nth clustering operation, or the Nth clustering cluster obtained in the Nth clustering operation.
[0113] It should be understood that in specific implementation, the clustering operation can be repeatedly executed until all clustering clusters are merged. Among them, in the process of performing the (N + 1)th clustering operation, when the similarity between any two Nth clustering clusters is greater than or equal to the second preset value, it can be considered that all clustering clusters have been merged at this time. It should be understood that the calculation method of the similarity is not limited herein. For example, in some embodiments, the similarity between the first clustering cluster and the second clustering cluster can also be understood as the average distance d(u, v) of the squares of the distances between the elements included in the first clustering cluster and the second clustering cluster pairwise. Among them, in this embodiment, the element can be understood as the human body image; then the d(u, v) satisfies:
[0114]
[0115] Among them, u can be understood as the feature of the first clustering cluster, u[i] can be understood as the feature of the ith human body image in the first clustering cluster, v can be understood as the feature of the second clustering cluster, v[j] can be understood as the feature of the jth human body image in the second clustering cluster, i and j are positive integers, and the maximum value of i is the total number of human body images in the first clustering cluster, and the maximum value of j is the total number of human body images in the second clustering cluster.
[0116] Among them, the feature of the first clustering cluster is not limited herein. For example, in some embodiments, the feature of the first clustering cluster can be understood as the average value of all features under the first clustering cluster. The feature of the second clustering cluster is not limited herein. For example, in some embodiments, the feature of the second clustering cluster can be understood as the average value of all features under the second clustering cluster.
[0117] In this embodiment, based on the feature data, multiple human body images are clustered according to the similarity between the multiple human body images to obtain clustering clusters. Since the similarity between the multiple human body images is used as the basis for clustering when clustering the multiple human body images, highly similar human body images can be merged into one clustering cluster, so that the same character corresponding to multiple images in a continuous time in a video appears in one clustering cluster.
[0118] Optionally, in some embodiments, after step 104, the method further includes:
[0119] Determining the role category corresponding to each human body image based on the role category corresponding to the clustering cluster where each human body image is located;
[0120] Determining the role category of the image to be recognized based on the role category corresponding to the human body image.
[0121] It should be understood that each human body image is located in one clustering cluster. Therefore, according to the role category corresponding to the clustering cluster where the human body image is located, it is the role category corresponding to the human body image.
[0122] It should be understood that the human body image is a partial image of the corresponding image to be recognized. Therefore, each human body image corresponds to one image to be recognized. Determining the role category of the image to be recognized based on the role category corresponding to the human body image can be understood as determining the role category corresponding to the human body image as the role category of its corresponding image to be recognized.
[0123] In the embodiments of the present invention, the role category corresponding to each image to be recognized can be determined by the role category corresponding to the clustering cluster where each human body image is located. In specific implementation, one or more videos can be analyzed. Multiple images to be recognized corresponding to one or more videos are input into the role feature extraction model. Through the role recognition method provided by the embodiments of the present invention, the time points when the same role category appears in one or more videos can be obtained.
[0124] In an embodiment of the present invention, a human detection region corresponding to each of a plurality of images to be recognized is determined; a partial image within the human detection region corresponding to the image to be recognized is determined as a human image; the human image is input into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data. Since the feature data includes not only face feature data but also clothing feature data and / or accessory feature data. Therefore, in the process of clustering a plurality of the human images based on the feature data to obtain clustering clusters, more of the feature data can be utilized, so the similarity between the human images included in the same clustering cluster is higher and the clustering effect is better.
[0125] Therefore, when recognizing a character, recognition can be performed based on information from multiple aspects such as the face, clothing, and / or accessories, improving the accuracy of character recognition.
[0126] Next, a specific embodiment will be used as an example to illustrate the character recognition method. As Figure 2 shown, Figure 2 is a schematic flowchart of the character recognition method provided in this embodiment. It should be noted that in this embodiment, the feature detection region is a clothing detection region.
[0127] First, a plurality of input images to be recognized are obtained, and human image detection and clothing image detection are performed on each of the images to be recognized. Specifically, the plurality of images to be recognized are input into a pre-trained human image detection model for image detection to determine a human detection region corresponding to each of the plurality of images to be recognized. Then, the plurality of images to be recognized are input into a pre-trained target image detection model for image detection to obtain a feature detection region corresponding to each of the images to be recognized, where, in this embodiment, the feature detection region is a clothing feature detection region. The human image and the target image can be the same image to be recognized.
[0128] For each of the images to be recognized, calculate the IOU between the human detection region and the clothing detection region. When the IOU is greater than 0.6, determine the partial image within the human detection region corresponding to the image to be recognized as the human image. When the IOU is less than or equal to 0.6, re-determine the human detection region corresponding to each image to be recognized among the multiple images to be recognized, and calculate the IOU between the new human detection region and the clothing detection region. When the IOU is greater than 0.6, determine the image within the new human detection region as the human image. When the IOU is less than or equal to 0.6, do not determine the image within the new human detection region as the human image, that is, the image corresponding to the new human detection region of the image to be recognized will not be input into the pre-trained role feature extraction model for feature extraction.
[0129] Input the human image into the pre-trained role feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data, clothing feature data, and accessory feature data.
[0130] Perform M clustering operations on the multiple human images based on the feature data to obtain the M-th target clustering cluster;
[0131] Among them, the N-th clustering operation includes:
[0132] Calculate the similarity between the first clustering cluster and the second clustering cluster based on the first feature data and the second feature data, where the first feature data is the feature data corresponding to the first clustering cluster, and the second feature data is the feature data corresponding to the second clustering cluster;
[0133] When the similarity between the first clustering cluster and the second clustering cluster is less than the second preset value, merge the first clustering cluster and the second clustering cluster into one N-th clustering cluster; when the similarity between the first clustering cluster and the second clustering cluster is greater than or equal to the second preset value, separate the first clustering cluster and the second clustering cluster; both N and M are positive integers, and N is less than or equal to M.
[0134] After performing M clustering operations on the multiple human images based on the feature data, each human image is in any one of the M-th clustering clusters. And each of the M-th clustering clusters corresponds to a role category. Therefore, according to the role category corresponding to the M-th clustering cluster where the human image in each image to be recognized is located, the role category corresponding to the image to be recognized can be determined.
[0135] As Figure 4 shown, an embodiment of the present invention further provides a role recognition device 400, and the role recognition device 400 includes:
[0136] The first determination module 401 is configured to determine a human detection region corresponding to each of the multiple images to be recognized.
[0137] The second determination module 402 is configured to determine a partial image within the human detection region corresponding to the image to be recognized as a human image.
[0138] The feature extraction module 403 is configured to input the human image into a pre-trained role feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data.
[0139] The clustering processing module 404 is configured to perform clustering processing on the multiple human images based on the feature data to obtain clustering clusters; each clustering cluster includes at least one of the human images corresponding to one role category.
[0140] Optionally, the role recognition device 400 further includes:
[0141] The third determination module is configured to determine the role category corresponding to each human image based on the role category corresponding to the clustering cluster where each human image is located.
[0142] The fourth determination module is configured to determine the role category of the image to be recognized based on the role category corresponding to the human image.
[0143] Optionally, the second determination module 402 includes:
[0144] The first determination unit is configured to determine a feature detection region corresponding to each of the multiple images to be recognized; the feature detection region includes at least one of a face feature detection region, a clothing feature detection region, and an accessory feature detection region; where the image within the face feature detection region is a face feature image, the image within the clothing feature detection region is a clothing feature image, and the image within the accessory feature detection region is an accessory feature image.
[0145] The second determination unit is configured to, when the intersection over union (IOU) between the human detection region and the feature detection region is greater than a first preset value, determine a partial image within the human detection region corresponding to the image to be recognized as a human image.
[0146] Optionally, the range of the first preset value is 0.5 to 0.7.
[0147] Optionally, the first determination unit includes:
[0148] An input subunit for inputting a plurality of the images to be recognized into a pre-trained target image detection model for image detection, so as to obtain a feature detection region corresponding to each of the images to be recognized;
[0149] Wherein, the target image detection model includes at least one of the following: a face image detection model, a clothing image detection model, and an accessory image detection model.
[0150] Optionally, the clustering processing module 404 includes:
[0151] A clustering operation unit for performing M clustering operations on a plurality of the human body images based on the feature data to obtain an Mth target clustering cluster;
[0152] Wherein, the (N + 1)th clustering operation includes:
[0153] Calculating a similarity between a first clustering cluster and a second clustering cluster based on first feature data and second feature data, where the first feature data is the feature data corresponding to the first clustering cluster, and the second feature data is the feature data corresponding to the second clustering cluster; the first clustering cluster and the second clustering cluster are any two Nth clustering clusters;
[0154] In the case where the similarity between the first clustering cluster and the second clustering cluster is less than a second preset value, merging the first clustering cluster and the second clustering cluster into one Nth clustering cluster; both N and M are positive integers greater than 1, and N is less than or equal to M.
[0155] Optionally, the role recognition device 400 further includes:
[0156] An iterative training module for iteratively training a role feature extraction model to be trained by using a plurality of pre-acquired sample images to obtain the trained role feature extraction model;
[0157] Wherein, the Nth iterative training includes:
[0158] Inputting a plurality of pre-acquired sample images into the role feature extraction model to be trained for Nth role feature extraction to obtain Nth feature data;
[0159] Performing clustering processing on the plurality of sample images based on the Nth feature data to obtain at least two sample clustering clusters;
[0160] Judging whether a loss value satisfies a loss convergence condition;
[0161] In the case that the loss value does not satisfy the loss convergence condition, the parameters of the to-be-trained character feature extraction model are adjusted based on the loss value; in the case that the loss value satisfies the loss convergence condition, the currently trained to-be-trained character feature extraction model is determined as the pre-trained character feature extraction model.
[0162] Optionally, the iterative training module includes:
[0163] An acquisition unit, configured to acquire a plurality of sample images;
[0164] A preset processing unit, configured to perform preset processing on the plurality of sample images to obtain augmented sample images;
[0165] An iterative training unit, configured to perform iterative training on the to-be-trained character feature extraction model by using the plurality of sample images and the augmented sample images.
[0166] The character recognition device 400 provided by the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, details are not described herein again.
[0167] As Figure 5 shown, an embodiment of the present invention further provides an electronic device, as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0168] The memory 503 is used for storing programs;
[0169] When the processor 501 is configured to execute the programs stored in the memory 503, the following steps are implemented:
[0170] Determine a human detection area corresponding to each of the plurality of images to be recognized;
[0171] Determine a partial image within the human detection area corresponding to the image to be recognized as a human image;
[0172] Input the human image into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data;
[0173] Perform clustering processing on the plurality of human images based on the feature data to obtain clustering clusters; each clustering cluster includes at least one of the human images corresponding to a character category.
[0174] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0175] The communication interface is used for communication between the above terminal and other devices.
[0176] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0177] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0178] In another embodiment provided by the present invention, a readable storage medium is further provided. Instructions are stored in the readable storage medium, and when executed by a processor, the method described in any one of the above embodiments is implemented.
[0179] In another embodiment provided by the present invention, a program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute the method described in any one of the above embodiments.
[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0181] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0182] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.
[0183] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A role recognition method, characterized in that, Including: Determine a human detection region corresponding to each of the multiple images to be recognized; Determine a partial image within the human detection region corresponding to the image to be recognized as a human image; Input the human image into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data; Perform clustering processing on the multiple human images based on the feature data to obtain clustering clusters; each clustering cluster includes at least one of the human images corresponding to a character category, and the character category includes different characters played by the same actor, and the same character corresponding to multiple human images in consecutive time in a video appears in one clustering cluster; The performing clustering processing on the multiple human images based on the feature data to obtain clustering clusters includes: Perform M clustering operations on the multiple human images based on the feature data to obtain an M-th target clustering cluster; Wherein, the (N + 1)-th clustering operation includes: Calculate the similarity between a first clustering cluster and a second clustering cluster based on first feature data and second feature data, where the first feature data is the feature data corresponding to the first clustering cluster, and the second feature data is the feature data corresponding to the second clustering cluster; the first clustering cluster and the second clustering cluster are any two N-th clustering clusters; In the case where the similarity between the first clustering cluster and the second clustering cluster is less than a second preset value, merge the first clustering cluster and the second clustering cluster into one N-th clustering cluster; both N and M are positive integers greater than 1, and N is less than or equal to M.
2. The method according to claim 1, wherein After performing clustering processing on the multiple human images based on the feature data to obtain clustering clusters, the method further includes: Determine the character category corresponding to each human image based on the character category corresponding to the clustering cluster where each human image is located; Determine the character category of the image to be recognized based on the character category corresponding to the human image.
3. The method according to claim 1, wherein The determining a partial image within the human detection region corresponding to the image to be recognized as a human image includes: Determine a feature detection region corresponding to each of the multiple images to be recognized; the feature detection region includes at least one of a face feature detection region, a clothing feature detection region, and an accessory feature detection region; wherein, the image within the face feature detection region is a face feature image, the image within the clothing feature detection region is a clothing feature image, and the image within the accessory feature detection region is an accessory feature image; In the case where the intersection over union (IOU) between the human detection region and the feature detection region is greater than a first preset value, determine a partial image within the human detection region corresponding to the image to be recognized as a human image.
4. The method according to claim 3, characterized in that, The range of the first preset value is 0.5 to 0.
7.
5. The method according to claim 3, characterized in that The determining a feature detection region corresponding to each of the multiple images to be recognized includes: Input multiple images to be recognized into a pre-trained target image detection model for image detection, and obtain a feature detection region corresponding to each image to be recognized; Among them, the target image detection model includes at least one of the following: a face image detection model, a clothing image detection model, and an accessory image detection model.
6. The method according to claim 1, characterized in that, Before inputting the human body image into a pre-trained character feature extraction model for feature extraction to obtain feature data, the method further includes: Iteratively train a character feature extraction model to be trained using a plurality of pre-acquired sample images, and obtain the trained character feature extraction model; Among them, the Nth iterative training includes: Input a plurality of pre-acquired sample images into the character feature extraction model to be trained for the Nth character feature extraction, and obtain the Nth feature data; Perform clustering processing on the plurality of sample images based on the Nth feature data to obtain at least two sample clusters; Determine whether the loss value satisfies the loss convergence condition; In the case where the loss value does not satisfy the loss convergence condition, adjust the parameters of the character feature extraction model to be trained based on the loss value; in the case where the loss value satisfies the loss convergence condition, determine the currently trained character feature extraction model to be trained as the pre-trained character feature extraction model.
7. The method according to claim 6, wherein The iterative training of the character feature extraction model to be trained using a plurality of pre-acquired sample images includes: Obtain a plurality of sample images; Perform preset processing on the plurality of sample images to obtain augmented sample images; Iteratively train the character feature extraction model to be trained using the plurality of sample images and the augmented sample images.
8. A role recognition device, characterized in that, Includes: A first determination module, configured to determine a human body detection region corresponding to each image to be recognized in a plurality of images to be recognized; A second determination module, configured to determine a partial image within the human body detection region corresponding to the image to be recognized as a human body image; A feature extraction module, configured to input the human body image into a pre-trained character feature extraction model for feature extraction to obtain feature data, where the feature data includes face feature data and target feature data, and the target feature data includes clothing feature data and / or accessory feature data; A clustering processing module, configured to perform clustering processing on a plurality of human body images based on the feature data to obtain clusters; each cluster includes at least one human body image corresponding to a character category, and the character category includes different characters played by the same actor, and the same character corresponding to a plurality of human body images in a continuous time period in a video appears in one cluster; The clustering processing module includes: A clustering operation unit, configured to perform M clustering operations on a plurality of human body images based on the feature data to obtain the Mth target cluster; Among them, the (N + 1)th clustering operation includes: Calculate the similarity between the first cluster and the second cluster based on the first feature data and the second feature data, where the first feature data is the feature data corresponding to the first cluster, and the second feature data is the feature data corresponding to the second cluster; the first cluster and the second cluster are any two Nth clusters; In the case where the similarity between the first cluster and the second cluster is less than the second preset value, merge the first cluster and the second cluster into one Nth cluster; both N and M are positive integers greater than 1, and N is less than or equal to M.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store programs; The processor is used to implement the method steps described in any one of claims 1-7 when executing the programs stored on the memory.
10. A readable storage medium, on which a program is stored, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1-7.
Citation Information
Patent Citations
Person classification method based on video, intelligent terminal and storage medium
CN110414344A
Costume searching method and device, electronic equipment and medium
CN112905889A
Face clustering method and system based on feature normalization
CN113627476A