Image Processing Method, Video Recognition Method, Device, Equipment and Storage Medium
By extracting face features from user's avatar images and historical videos and building a target face library, the high cost and low accuracy of video originality recognition in the existing technology is solved, and low-cost and efficient face library construction and accurate video originality recognition are achieved.
Patent Information
- Application Number
- CN202111331061.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-11
AI Technical Summary
In the prior art, video original recognition methods have problems such as high storage costs, high face library construction costs and poor accuracy, and it is difficult to effectively cover various state changes of faces.
By extracting the first face feature that meets the preset requirements from the target user's avatar image, combining the second face feature that meets the requirements for the frequency of the appearance in his historical video, the candidate face feature is filtered using the benchmark face feature set to build a target face library.
It realizes low-cost and high-efficiency face library construction, which can cover a variety of face state changes, improves the accuracy of the face library, accurately recognizes the originality of videos, and reduces the probability of misidentification.
Smart Images

Figure CN113963303B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing, and in particular, to an image processing method, a video recognition method, a device, a device, and a storage medium. Background Art
[0002] With the rapid development of Internet technology, more and more applications or platforms can support users to upload audio and video content by themselves, such as short videos. With the large growth of video content on the Internet, more and more application requirements have emerged in aspects such as video content understanding and content review, among which originality recognition is an important application direction.
[0003] Currently, common implementation methods for originality recognition include methods based on video retrieval and methods based on video face recognition. Among them, the core of video retrieval is to establish a retrieval library to store the retrieval features corresponding to each original video, and then determine whether the video already exists by comparing the video to be recognized with the videos in the retrieval library, so as to determine whether the video is original. However, this method requires collecting and storing a large number of videos, has a high storage cost, and it is difficult to effectively cover off-terminal video piracy and is easily affected by video editing. The method of video face recognition requires establishing a face library in advance. For the video to be recognized, after detecting the face from the video, it is matched with the face library to determine whether the face in the video is the same as the face of the video uploader, and then determine whether the video is original. However, in this method, the face library is usually constructed manually, and the manually constructed face library is difficult to cover various state changes of the face. This construction scheme has a high cost and poor accuracy, so it needs to be improved. Summary of the Invention
[0004] Embodiments of the present invention provide an image processing method, a video recognition method, a device, a device, and a storage medium, which can optimize the existing face-based image processing and video recognition solutions.
[0005] In a first aspect, an embodiment of the present invention provides an image processing method, which includes:
[0006] Obtain an avatar image of a target user in a preset application program, and add the first face feature that meets the preset requirements to the benchmark face feature set when the first face feature is extracted from the avatar image;
[0007] Obtain the second face feature corresponding to the target face whose appearance frequency meets the preset frequency requirement from the historical videos published by the target user in the preset application program, and add the second face feature to the candidate face feature set;
[0008] Screen the face features in the candidate face feature set based on the benchmark face feature set, and determine the extended face feature set according to the screening result;
[0009] Construct a target face database corresponding to the target user based on the reference face feature set and the extended face feature set.
[0010] In a second aspect, an embodiment of the present invention provides a video recognition method, which includes:
[0011] Extract face features to be recognized from a target video uploaded by a target user.
[0012] Compare the face features to be recognized with the face features in the target face database corresponding to the target user, and identify the originality of the target video according to the comparison result, where the target face database is obtained by using the image processing method provided in the embodiment of the present invention.
[0013] In a third aspect, an embodiment of the present invention provides an image processing device, which includes:
[0014] An avatar image processing module, configured to obtain an avatar image of a target user in a preset application program, and add the first face feature to the reference face feature set when the first face feature that meets the preset requirements is extracted from the avatar image;
[0015] A historical video processing module, configured to obtain second face features corresponding to a target face with an appearance frequency that meets a preset frequency requirement from historical videos published by the target user in the preset application program, and add the second face features to the candidate face feature set;
[0016] A face feature screening module, configured to screen the face features in the candidate face feature set based on the reference face feature set, and determine an extended face feature set according to the screening result;
[0017] A face database construction module, configured to construct a target face database corresponding to the target user according to the reference face feature set and the extended face feature set.
[0018] In a fourth aspect, an embodiment of the present invention provides a video recognition device, which includes:
[0019] A face feature extraction module, configured to extract face features to be recognized from a target video uploaded by a target user;
[0020] An originality recognition module, configured to compare the face features to be recognized with the face features in the target face database corresponding to the target user, and identify the originality of the target video according to the comparison result, where the target face database is obtained by using the image processing method provided in the embodiment of the present invention.
[0021] Fifth aspect, an embodiment of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the image processing method and / or video recognition method provided by the embodiments of the present invention.
[0022] Sixth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method and / or video recognition method provided by the embodiments of the present invention.
[0023] In the image processing solution provided by the embodiments of the present invention, the first face feature that meets the preset requirements is extracted from the head image of the target user and added to the benchmark face feature set. The second face feature corresponding to the target face with an appearance frequency that meets the preset frequency requirement is obtained from the historical videos published by the target user and added to the candidate face feature set. The face features in the candidate face feature set are screened based on the benchmark face feature set to obtain an extended face feature set. Finally, a target face library corresponding to the target user is constructed according to the benchmark face feature set and the extended face feature set. By adopting the above technical solution, the face library of the user can be automatically constructed by comprehensively integrating the head image of the same user and the historical videos published by the user. The construction cost of this solution is low, the construction efficiency is high, and the constructed face library can cover more face state changes, can more comprehensively and accurately represent the face features of the user, improve the accuracy of the face library, and further enable better application effects when using this face set for applications such as video originality recognition.
[0024] In the video recognition solution provided by the embodiments of the present invention, the face feature to be recognized is extracted from the target video to be recognized uploaded by the target user, and the face feature to be recognized is compared with each face feature in the target face library corresponding to the target user obtained by using the image processing method provided by the embodiments of the present invention. According to the comparison result, the originality of the target video is recognized. Since this face library can more comprehensively and accurately represent the face features of the user, it can more accurately recognize the originality of the video, reduce the probability of misrecognition, facilitate further targeted related processing of the target video, improve the video processing efficiency, and enhance the rationality of video processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of the present invention;
[0026] Figure 2 It is a schematic flowchart of another image processing method provided by an embodiment of the present invention;
[0027] Figure 3Schematic diagram of the principle of the image processing method provided by the embodiment of the present invention;
[0028] Figure 4 Flow chart of a video recognition method provided by the embodiment of the present invention;
[0029] Figure 5 Block diagram of the structure of an image processing device provided by the embodiment of the present invention;
[0030] Figure 6 Block diagram of the structure of a video recognition device provided by the embodiment of the present invention;
[0031] Figure 7 Block diagram of the structure of a computer device provided by the embodiment of the present invention. Detailed implementation manners
[0032] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the accompanying drawings rather than all the structures. Furthermore, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0033] Figure 1 Flow chart of an image processing method provided by the embodiment of the present invention. This method can be executed by an image processing device, which can be implemented by software and / or hardware and is generally integrated in a computer device. The computer device can be a mobile device such as a mobile phone, a tablet computer, a laptop computer, and a personal digital assistant; or other devices such as a desktop computer or a server. As Figure 1 shown, this method includes:
[0034] Step 101: Obtain the avatar image of the target user in the preset application. When the first face feature that meets the preset requirements is extracted from the avatar image, add the first face feature to the reference face feature set.
[0035] In the embodiments of the present disclosure, the preset application can be an application with a video publishing function, and the specific type is not limited. For example, it can be a social application such as a short video application, or other types of applications. The preset application can be installed as a client on an electronic device, which can be the same as or different from the computer device described in the embodiments of the present invention. In different cases, the computer device can be the server device corresponding to the preset application.
[0036] Exemplarily, a user can register an account in a preset application and publish works such as videos through the registered account. Taking a specific application scenario as an example, the preset application includes a short video application. As a video author, the user can upload the short video works he shot to the short video platform, and the short video platform can send the short video works to the short video applications used by other users for playing, so that other users can watch the short video works published by the video author.
[0037] Exemplarily, after registering an account, the user can set the avatar corresponding to his account. Generally, in order to enhance the recognition of his account, the user tends to use a photo containing his own face (part of which may also include the faces of relatives, friends or partners and other people with an intimate relationship) as the avatar. In addition, the display area of the avatar image in the preset application is usually relatively small. In order to let other users clearly see his face, the proportion of the face in the entire avatar image is usually relatively high. In the embodiments of the present invention, preferentially obtaining the reference face features from the avatar image can effectively improve the acquisition efficiency of the reference face features and reduce the acquisition cost.
[0038] Exemplarily, the target user can be understood as the user who currently needs to construct a face database specifically, which can be any user or a specified user, and is not specifically limited. The target account can be understood as the account registered and used by the target user in the preset application. The obtained avatar image can include all or part of the avatar images used by the target user, that is, all or part of the avatar images set by the target user for the target account. For example, it can include the currently used avatar image, or the avatar images used in the first preset historical period (such as the most recent year), or all the avatar images used from the moment when the target account was registered to the current moment.
[0039] Exemplarily, since the user's avatar is usually a custom image, it is generally impossible to guarantee whether it contains a human face or the quality of the human face. For example, there may be natural scenery in the avatar image, the image in the avatar is a cartoon character or other non-real people, and there may be occlusion or a side face in the human face in the avatar image. In order to better obtain the features that can accurately represent the target user's face, preset requirements can be set, and the first face features that meet the preset requirements are extracted from the avatar image and added to the benchmark face feature set. Among them, the preset requirements can be set according to the actual situation. For example, the clarity of the human face meets the preset clarity requirement (such as being greater than the preset clarity threshold, and the specific calculation method of clarity is not limited), the rotation angle of the human face is less than the preset angle threshold, and there is no occlusion on the human face, etc. Among them, the rotation angle of the human face can be understood as the angle between the plane where the human face is located and the plane where the shooting lens is located. When the human face is a standard front face, the rotation angle of the human face is 0 or close to 0. The occluder may be a mask or sunglasses, etc.
[0040] Exemplarily, a preset face detection model can be used to detect whether the avatar image contains a human face. In the case of determining that a human face is contained, a preset face recognition model is used to extract the face features, and it is judged whether the face features meet the preset requirements. The face features that meet the preset requirements are recorded as the first face features. The number of the first face features can be one or more, or may be 0. In the case of extracting at least one first face feature, the extracted first face features are added to the benchmark face feature set.
[0041] Exemplarily, it is also possible to directly predict based on the avatar image whether it contains face features that meet the preset requirements. If so, the face features are extracted from the avatar image, and the extracted face features are added to the benchmark face feature set as the first face features. The advantage of such a setting is that it can reduce the data processing volume in the face detection and face recognition processes and improve the extraction efficiency of the first face features.
[0042] Exemplarily, the initial state of the benchmark face feature set can be an empty set or can contain initial benchmark face features, which is not specifically limited. The initial benchmark face features can be, for example, the face features extracted from the target image uploaded by the target user. For example, the preset application can prompt the target user to actively upload an image containing their own face as the target image to more accurately determine the benchmark face.
[0043] Step 102: Obtain the second face features corresponding to the target face whose appearance frequency meets the preset frequency requirement from the historical videos published by the target user in the preset application, and add the second face features to the candidate face feature set.
[0044] In the embodiments of the present invention, the historical videos may include all or part of the videos published by the target user in a preset application. For example, it may be the videos published within a second preset historical period (such as the most recent three months, which can be dynamically determined according to factors such as the video publishing frequency of the target user). Another example is that a screening method other than the time dimension can also be adopted. For example, according to the business characteristics of the preset application, videos recorded by the front camera can be selected, or grouped according to the recording method, and the second face features can be screened for each group respectively. For many users, they often take themselves as the shooting object, that is, the video content usually contains the face of the video publishing user. Therefore, face features can be extracted from the historical videos published by the target user to enrich the face library corresponding to the target user. In addition, the faces of other users other than the target user may appear in the historical videos, but compared with the face of the target user, the appearance frequency is generally low. Therefore, a preset frequency requirement can be set. For example, the face with the highest appearance frequency is determined as the target face; another example is that multiple faces with a higher appearance frequency (faces with a frequency greater than the preset frequency threshold, or a preset number of faces with a higher frequency) are determined as the target faces. Considering that the appearance frequency of the friends or relatives of the target user may also be high, it may be related to the style or content of the video works published by the target user. For example, some users may be used to shooting the daily life of their family members, etc. Face features are extracted for the target face, and the extracted face features can be recorded as the second face features.
[0045] Exemplarily, the face library is usually used for video recognition. The faces appearing in the video to be recognized may have various state changes and interference factors, such as posture, face rotation angle, makeup, and illumination, etc. It is difficult to effectively cover various changes when manually constructing the face library, and the cost is high. In the embodiments of the present invention, face features are automatically mined from the historical videos of the user, and through subsequent screening based on the benchmark face features, the features for representing the faces of the user in different states can be quickly and accurately screened out, improving the accuracy of the face library and being beneficial to improving the video recognition effect.
[0046] Optionally, while the candidate face feature set includes the second face features, it may also include other face features, and the specific source is not limited. For example, it may come from an avatar image or a photo actively uploaded by the user, etc.
[0047] Step 103: Screen the face features in the candidate face feature set based on the benchmark face feature set, and determine the extended face feature set according to the screening result.
[0048] Exemplarily, the facial features in the reference facial feature set can accurately represent the face of the target user, while the facial features in the candidate facial feature set are derived from videos and can cover a relatively large number of facial state changes. By using the reference facial feature set to screen the facial features in the candidate facial feature set, some facial features with poor quality or low reference value can be excluded, which is beneficial to improving the accuracy of the final facial database. Optionally, the screening method can be to retain the facial features that are close to the facial features in the reference facial feature set, and add the retained facial features to the extended facial feature set.
[0049] Step 104: Construct a target facial database corresponding to the target user according to the reference facial feature set and the extended facial feature set.
[0050] Exemplarily, the reference facial feature set and the extended facial feature set can be merged to form a target facial database corresponding to the target user. Optionally, the facial features in the reference facial feature set and the extended facial feature set are added to an initial facial database corresponding to the target user to obtain the target facial database. Among them, the initial facial database can be empty or contain initial facial features, and the source of the initial facial features is not limited. For example, it can be the facial features extracted from a target image containing the face of the target user actively uploaded by the target user, etc.
[0051] Exemplarily, the target facial database can be used to identify the originality of the video uploaded by the target user. The specific application method is not limited and optional application methods will be described below.
[0052] In the image processing method provided in the embodiments of the present invention, the first facial features that meet the preset requirements are extracted from the avatar image of the target user and added to the reference facial feature set, the second facial features corresponding to the target face with the appearance frequency meeting the preset frequency requirement are obtained from the historical videos released by the target user and added to the candidate facial feature set, the facial features in the candidate facial feature set are screened based on the reference facial feature set to obtain the extended facial feature set, and finally a target facial database corresponding to the target user is constructed according to the reference facial feature set and the extended facial feature set. By adopting the above technical solution, the facial database of the user can be automatically constructed by comprehensively integrating the avatar image of the same user and the historical videos released by the user. This solution has a low construction cost, high construction efficiency, and the constructed facial database can cover a relatively large number of facial state changes, can more comprehensively and accurately represent the facial features of the user, improve the accuracy of the facial database, and further can achieve better application effects when applying aspects such as video originality recognition using this facial feature set.
[0053] In some embodiments, before screening the face features in the candidate face feature set based on the reference face feature set, the method further includes: adding the third face features that do not meet the preset requirements extracted from the avatar image to the candidate face feature set. The advantage of this setting is that for the face features that fail to be added to the reference face feature set, there is also a certain possibility that they belong to the target user himself, but may be affected by interference factors such as occlusion or side face, and thus fail to be successfully added to the reference face feature set. However, these interference factors may also appear in the video to be recognized. Therefore, adding the third face features that do not meet the preset requirements to the candidate face feature set can help the face database cover more diverse face states, contribute to improving the anti-interference ability of video face recognition, and further enhance the application effect of the face database. Among them, the third face features may include all or some of the face features in the avatar image that do not meet the preset requirements.
[0054] In some embodiments, after obtaining the avatar image of the target user in the preset application program, the method further includes: inputting the avatar image into a preset classification model, and determining the category of the avatar image according to the output result of the preset classification model, where the labels of the training sample images corresponding to the preset classification model include available and unavailable; extracting face features from the avatar images of which the category is available to obtain the first face features that meet the preset requirements. The advantage of this setting is that the classification model can be used to quickly and accurately identify the available avatar images, improving the extraction efficiency of the first face features.
[0055] Exemplarily, the preset classification model can be a binary classification model or a multi-classification model, without specific limitation. The binary classification model has advantages such as simple training and small computational overhead, which can minimize additional cost overhead. Taking the binary classification model as an example, in the model training stage, the initial classification model can be trained using positive samples and negative samples, and then the preset classification model can be obtained. Positive samples and negative samples are collectively referred to as training samples, specifically training sample images. The training sample images are images with labels. The label of the positive sample can be available, and the label of the negative sample can be unavailable. This label can be added to the images used for training by manual annotation. When annotating, the clarity of the face in the image and the face rotation angle can be referred to. For example, an image containing a clear and unobstructed frontal face is annotated with an available label, and the rest of the images are annotated with an unavailable label. Exemplarily, a preset rule can also be used to evaluate the sample images, and then determine the corresponding labels. For example, the clarity of the face in the training sample image with the label available meets the preset clarity requirement and the face rotation angle is less than the preset angle threshold, and the clarity of the face in the training sample image with the label unavailable does not meet the preset clarity requirement or the face rotation angle is greater than or equal to the preset angle threshold. Optionally, it can also be determined whether there is an occlusion situation of the face. If so, the corresponding label is unavailable.
[0056] Exemplarily, after the avatar image is input into the preset classification model, it can be determined whether the current avatar image is an available avatar or an unavailable avatar. For an avatar image with the category of available, it can be considered that it contains face features that meet the preset requirements. At this time, the preset face recognition model is used to extract face features therefrom to obtain the first face features.
[0057] In some embodiments, the obtaining the second face features corresponding to the target face whose appearance frequency meets the preset frequency requirement from the historical videos published by the target user in the preset application program includes: performing face feature extraction on the video frames in the historical videos published by the target user in the preset application program, and adding the extracted face features as alternative face features to the alternative face feature pool; performing a preset clustering process on the alternative face features in the alternative face feature pool to obtain multiple clusters; counting the number of alternative face features included in each cluster, and determining the cluster whose cumulative number meets the preset number requirement as the high-frequency cluster, where the face to which the alternative face features in the high-frequency cluster belong is recorded as the target face whose appearance frequency meets the preset frequency requirement; obtaining the second face features from the high-frequency cluster. The advantage of such a setting is that the clustering algorithm can quickly classify the numerous face features that appear in the historical videos, facilitating the finding of the target face with a higher appearance frequency.
[0058] Exemplarily, the clustering algorithm used for the preset clustering process is not limited. For example, it can be the Density-Based Spatial Clustering of Applications with Noise (DBSCAN), spectral clustering, K-Means, etc. To ensure that the face features within each cluster correspond to the face of the same person, a relatively strict similarity threshold can be used. Two alternative face features will be connected together to form a cluster (also known as clustering) when they are highly similar as two feature points. The preset quantity requirement can be, for example, the largest quantity, the quantity being greater than a preset quantity threshold, or the ranking being greater than a preset ranking threshold when sorted from most to least in terms of quantity. After screening according to the preset quantity requirement, one or more clusters with more face features within the cluster can be obtained, denoted as high-frequency clusters, indicating that the corresponding faces appear more frequently. Subsequently, all or part of the face features are obtained from the high-frequency clusters as the second face features and added to the candidate face feature set.
[0059] In some embodiments, obtaining the second face features from the high-frequency clusters includes: for each alternative face feature in the high-frequency cluster, calculating the first similarity between the current alternative face feature and each of the other alternative face features in the same high-frequency cluster except the current alternative face feature, and calculating a first value, where the first value is the sum of the first similarities; obtaining the alternative face features in the high-frequency cluster for which the corresponding first value meets the first preset value requirement as the second face features. The advantage of this setting is that for the alternative face features in the high-frequency cluster, more representative alternative face features are further selected as the second face features for acquisition, which can reduce the computational amount of subsequent screening of the candidate face feature set based on the reference face feature set, save the device's computing and storage resources, and also improve the accuracy and precision of the face database.
[0060] Exemplarily, the calculation method of the first similarity is not limited. It can be calculating the cosine similarity or calculating the Euclidean distance, or other similarity measurement algorithms can be adopted according to the data characteristics.
[0061] In some embodiments, it further includes: in the case where no first facial feature that meets the preset requirements can be extracted from the avatar image, determining the cluster with the largest cumulative number as the target cluster, where the target cluster is included in the high-frequency clusters; adding the alternative facial features corresponding to the first value in the target cluster that meet the second preset value requirement to the reference facial feature set. Here, the second preset value requirement is generally more stringent than the second preset value requirement, for example, it can be the largest first value. The advantage of such a setting is that when the reference facial feature set cannot be determined based on the avatar image, the reference facial feature set is determined according to the face that appears most frequently in the historical video, ensuring the successful construction of the target face database.
[0062] In some embodiments, screening the facial features in the candidate facial feature set based on the reference facial feature set and determining the extended facial feature set according to the screening result includes: for each candidate facial feature in the candidate facial feature set, calculating the second similarity between the current candidate facial feature and each reference facial feature in the reference facial feature set, and in the case where there is at least one second similarity greater than the preset similarity threshold, determining the current candidate facial feature as an extended facial feature and adding it to the extended facial feature set. The advantage of such a setting is that based on the similarity, candidate facial features that are close to each reference facial feature in the reference facial feature set can be quickly screened out, which to a certain extent ensures that the screened facial features belong to the target user himself, improves the efficiency of face database construction, and can improve the accuracy and precision of the face database.
[0063] Exemplarily, the calculation method of the second similarity is not limited, and it can be the same as or different from the calculation method of the first similarity. It can be calculating the cosine similarity or calculating the Euclidean distance, and other similarity measurement algorithms can also be adopted according to the data characteristics.
[0064] Figure 2 It is a schematic flowchart of another image processing method provided by an embodiment of the present invention, which is optimized based on the above optional embodiments. Figure 3 It is a schematic diagram of the principle of the image processing method provided by an embodiment of the present invention, which can be combined with Figure 2 and Figure 3 to understand the embodiments of the present invention. As Figure 2 shown, the method may include:
[0065] Step 201, obtaining an avatar image of a target user in a preset application program.
[0066] In the embodiments of the present invention, the image processing method can be executed offline or online, and there is no specific limitation. For the offline execution case, the avatar images and historical videos of each user can be stored in a database. When it is necessary to construct a face database, the corresponding avatar image (i.e., Figure 3 the user avatar shown) can be obtained from the database associated with the preset application program for the target user.
[0067] Step 202: Input the avatar image into a preset classification model, and determine the category of the avatar image according to the output result of the preset classification model.
[0068] Exemplarily, the preset classification model can be a binary classification model for determining whether the avatar image is available, so it can be called an avatar availability model. Currently, the mainstream face recognition models usually have the best recognition effect on frontal faces that are clear and unobstructed. However, when there are interference factors such as face occlusion or profile faces, it is easier to have recognition errors. Therefore, in the embodiments of the present invention, an available avatar can be defined as an avatar containing a clear and unobstructed frontal face. Specifically, during model training, sample images with face clarity meeting the preset clarity requirements, face rotation angle less than the preset angle threshold, and no occluder on the face can be added with available labels to become positive samples, and sample images that do not meet this condition can be added with unavailable labels to become negative samples.
[0069] Step 203: Extract face features from the avatar images with the category of available to obtain the first face features that meet the preset requirements, add the first face features to the benchmark face feature set, extract face features from the avatar images with the category of unavailable to obtain the third face features, and add the third face features to the candidate face feature set.
[0070] Exemplarily, based on the judgment result of the avatar availability model, available avatars can be screened out from all the avatars used by the target user, and faces can be detected from the available avatars through a face detection model to form the benchmark face set of the target user, which can be denoted as The first face features obtained by extracting face features from the benchmark faces through a face recognition model. Considering that all the avatars of the target user may be unavailable, the benchmark face feature set may be empty. If there are no avatar images with the category of available, there is no need to process the benchmark face feature set in this step for the time being.
[0071] Exemplarily, for avatar images of the type unavailable, some interfering factors such as face occlusion or profile faces in them may also exist in the videos uploaded by users. Therefore, by utilizing these avatar images, the anti-interference ability of the face database can be improved. Specifically, face recognition can be performed on these avatar images, and the recognized faces can form an avatar face pool. Further, face features can be extracted to obtain third-party face features, which are added to the candidate face feature set for further screening in subsequent steps. Additionally, there may be no unavailable avatar images. For example, the avatars of the target user are all clear frontal faces without occlusion, and in this case, there are no third-party face features.
[0072] Step 204: Extract face features from the video frames in the historical videos published by the target user in the preset application, and add the extracted face features as alternative face features to the alternative face feature pool.
[0073] Exemplarily, corresponding historical videos can be obtained for the target user from the database associated with the preset application (i.e., Figure 3 the user-uploaded videos shown), and to improve the comprehensiveness of the face database, here it can be all the videos published by the target user, forming a historical video set, which can be denoted as For each historical video, the corresponding video frame set can be obtained by decoding respectively, which can be denoted as Use a face detection model to obtain an alternative face pool from the video frame set, which can be denoted as High-frequency faces can be understood as the faces that appear with a relatively high frequency in the videos published by the user. To find high-frequency faces from the alternative face pool, first, the face similarity in the alternative face pool can be calculated. Use a face recognition model to extract features for each face in the alternative face pool to obtain an alternative face feature pool, which can be denoted as
[0074] Step 205: Perform a preset clustering process on the alternative face features in the alternative face feature pool to obtain multiple clusters.
[0075] Exemplarily, use cosine similarity and DBSCAN clustering algorithm to divide all the alternative face features into several clusters, and each cluster contains one or more candidate face features. Exemplarily, through the features e i and face j of two faces i and e j the cosine similarity Indicates the similarity between two faces. The larger the cosine similarity, the higher the face similarity. The DBSCAN clustering algorithm is a clustering method based on density estimation. It connects similar feature points in a bottom-up manner, and all connected feature points form a cluster. The advantages of DBSCAN are fast calculation speed, and the ability to obtain clusters of any shape. It is also insensitive to outliers, so it can better meet the needs of mining high-frequency faces. In order to ensure that the facial features in each cluster correspond to the same face, a more stringent similarity threshold can be used, such as 95%.
[0076] Step 206: Count the number of candidate facial features contained in each cluster, and determine the cluster whose cumulative number meets the preset number requirement as a high-frequency cluster.
[0077] For example, after the candidate face feature pool is processed by the DBSCAN algorithm, several clusters are obtained. The more face features contained in the cluster, the higher the frequency of the corresponding face. The preset quantity requirement can be, for example, sorting by cumulative quantity, and determining the first A clusters (the value of A can be set according to actual conditions, such as 1 or 3) in the sorting as high-frequency clusters, that is, high-frequency clusters.
[0078] Step 207: For each candidate facial feature in the high-frequency cluster, calculate the first similarity between the current candidate facial feature and each candidate facial feature in the high-frequency cluster to which it belongs except the current candidate facial feature, and calculate the first value, and obtain the candidate facial feature in the high-frequency cluster whose corresponding first value meets the first preset value requirement as the second facial feature.
[0079] For example, the same cluster may contain features of multiple faces in different changing states. Some faces have poor usability due to factors such as angle, occlusion or video blur. Therefore, we can further find the most usable or high-frequency face features from the high-frequency cluster, so as to ensure that the face library built based on high-frequency faces has a high quality. To this end, we further calculate the cluster of each cluster The cosine similarity between all pairs of facial features (that is, every two facial features) in the , and the similarity matrix M∈R K×K , where M ij Represents the cosine similarity between the i-th face feature and the j-th face feature in the cluster. The second face feature is obtained by selecting the face feature with the largest or larger sum of similarities with other samples in the cluster. This process can be regarded as summing up each row of the similarity matrix M and then finding the column with the largest or larger value, that is, by Find the index p of the high-frequency face in the cluster to obtain the second face feature. The set of the second face features can be called Figure 3 The high-frequency face pool in .
[0080] Step 208: Add the second facial feature to the candidate facial feature set. When there is no avatar image with the category of available, determine the cluster with the largest cumulative quantity as the target cluster, and add the alternative facial features in the target cluster whose corresponding first value meets the requirements of the second preset value to the reference facial feature set.
[0081] Exemplarily, after obtaining the second facial feature, add the second facial feature to the candidate facial feature set as well. When there is a third facial feature, the candidate facial feature set simultaneously includes the facial features from the avatar image and the facial features of the high-frequency faces from the historical video, which can cover various facial features in multiple different states. That is, the candidate facial feature set contains Figure 3 the avatar face pool and the high-frequency face pool in
[0082] Among them, if the reference facial feature set cannot be obtained based on the avatar image, the face with the highest appearance frequency can be determined as the reference face of the user himself / herself, and the facial features with the highest or relatively high availability are selected from it and added to the reference facial feature set.
[0083] Step 209: For each candidate facial feature in the candidate facial feature set, calculate the second similarity between the current candidate facial feature and each reference facial feature in the reference facial feature set. When there is at least one second similarity greater than the preset similarity threshold, determine the current candidate facial feature as an extended facial feature and add it to the extended facial feature set.
[0084] Exemplarily, after obtaining the reference facial feature set and the candidate facial feature set, use the face library screening algorithm to screen out the facial features in the candidate facial feature set that are consistent with the reference face identity to expand the face library and improve the recognition ability of faces under various different interference factors in the video. Specifically, calculate the cosine similarity between the facial features in the candidate facial feature set and the reference facial features, and select the facial features with the cosine similarity greater than the preset similarity threshold and add them to the extended facial feature set.
[0085] Step 210: Add the reference facial feature set and the extended facial feature set to the initial face library corresponding to the target user to obtain the target face library.
[0086] Among them, the initial face database can be an empty set or already contain initial face features. Exemplarily, after obtaining the target face database, the target face database can also be updated regularly or irregularly. For example, when the target user uploads a new avatar, the update can be triggered; another example is that when the number of newly uploaded videos of the target user reaches a set value, the update can be triggered, etc. After the update is triggered, the image processing method of the embodiments of the present invention can be re-executed, or incremental updates can be performed on the newly uploaded avatar images and newly released videos. For example, for a newly uploaded avatar, it can be input into a preset classification model. If the output category is available, it can be directly added to the target face database. If the output category is unavailable, there is no need to update the target face database. Another example is that for a newly uploaded video, the face features therein can be extracted, and the second similarity between the face features and each reference face feature in the reference face feature set can be calculated. In the case that there is at least one second similarity greater than the preset similarity threshold, the face feature is added to the target face database.
[0087] The image processing method provided by the embodiments of the present invention preferentially extracts reference face features from the avatar images used by the user, uses a binary classification model to divide the avatar images into available avatars and unavailable avatars, extracts reference face features from the available avatars, constructs a candidate face feature set according to the high-frequency face features extracted from the historical videos published by the user and the face features extracted from the unavailable avatar images, and uses the reference face feature set to screen out the face features with higher similarity from the candidate face feature set as extended face features to expand the reference face feature set, so as to more comprehensively represent the face of the target user. This method avoids manually collecting the face materials of each user, reduces a large amount of labor costs, realizes the efficient automatic construction of the face database, has strong scalability of the face database, and can better adapt to the characteristics of continuous uploading of a large number of new videos by users in business scenarios such as short videos. By comprehensively using the user avatars and the videos uploaded by the user as the face database construction materials, the problem of missing face information of some users caused by using a single material can be solved, and the coverage rate at the user level can be improved. By constructing the face database with multiple different faces in the user avatars and the videos uploaded by the user, various interference factors and face state changes existing in online video face recognition can be effectively processed, thereby improving the accuracy of face recognition. Moreover, the method can be completed offline, that is, the computational cost can be in the offline face database construction stage, and for the online face recognition stage, no additional computational steps are added, so as not to increase the computational cost and processing delay of online face recognition. During the face database construction process, various mainstream face detection models and face recognition models can be used, and there is no need to collect additional training data for optimizing or adjusting the model, thereby avoiding the cost of collecting data and optimizing the model and improving the flexibility of practical applications.
[0088] Figure 4The flowchart of a video recognition method provided by an embodiment of the present invention. This method can be executed by a video recognition device, which can be implemented by software and / or hardware and is generally integrated in a computer device. The computer device can be a mobile device such as a mobile phone, a tablet computer, a laptop computer, and a personal digital assistant; it can also be other devices such as a desktop computer or a server. As Figure 4 shown, the method includes:
[0089] Step 401, extract the face features to be recognized from the target video uploaded by the target user.
[0090] Exemplarily, the target video can be an unpublished video uploaded by the target user or a published video, and there is no specific limitation. For the target video, it can be decoded to obtain the corresponding video frames, a set of faces can be detected from the video frames through a face detection model, and a face recognition model (generally the same as the face recognition model used in the face database construction process) is used to extract the face features to be recognized, and generally, there are multiple face features to be recognized.
[0091] Step 402, compare the face features to be recognized with the face features in the target face database corresponding to the target user, and identify the originality of the target video according to the comparison result.
[0092] Among them, the target face database is obtained by using any one of the image processing methods provided by the embodiments of the present invention.
[0093] Exemplarily, the face features to be recognized are matched with the face features in the target face database one by one. The matching method can be, for example, calculating the similarity between the face features to be recognized and the face features in the target face database. If there are face features to be recognized with a similarity greater than the set similarity threshold, it indicates that the target video to be recognized contains the face of the target user, and the target video can be considered an original video or has a high originality.
[0094] Exemplarily, after the originality of the target video is recognized, the target video can be processed specifically according to the originality, and the processing method can be set according to the actual business scenario. For example, if the originality is high, the probability of the target video entering the video recommendation queue can be increased, so that more users can view the target video. If the originality is low, the probability of the target video entering the video recommendation queue can be decreased, so that fewer users or even no users can view the target video.
[0095] In the embodiment of the present invention, the video recognition method extracts the face features to be recognized from the target video uploaded by the target user, compares the face features to be recognized with the face features in the target face library corresponding to the target user obtained by using the image processing method provided in the embodiment of the present invention, and identifies the originality of the target video according to the comparison result. Since the face library can more comprehensively and accurately represent the face features of the user, it can more accurately identify the originality of the video, reduce the probability of misrecognition, facilitate further targeted related processing of the target video, improve the video processing efficiency, and enhance the rationality of video processing.
[0096] Figure 5 It is a structural block diagram of an image processing device provided in an embodiment of the present invention. The device can be implemented by software and / or hardware, and is generally integrated in a computer device. It can construct a face library by executing an image processing method. As Figure 5 shown, the device includes:
[0097] The avatar image processing module 501 is used to obtain the avatar image of the target user in the preset application program, and add the first face feature to the reference face feature set when the first face feature that meets the preset requirements is extracted from the avatar image;
[0098] The historical video processing module 502 is used to obtain the second face feature corresponding to the target face with the appearance frequency meeting the preset frequency requirement from the historical videos published by the target user in the preset application program, and add the second face feature to the candidate face feature set;
[0099] The face feature screening module 503 is used to screen the face features in the candidate face feature set based on the reference face feature set, and determine the extended face feature set according to the screening result;
[0100] The face library construction module 504 is used to construct the target face library corresponding to the target user according to the reference face feature set and the extended face feature set.
[0101] The image processing device provided in the embodiment of the present invention can automatically construct the face library of the user by comprehensively integrating the avatar image of the same user and the historical videos published by the user. The scheme has a low construction cost, high construction efficiency, and the constructed face library can cover more face state changes, can more comprehensively and accurately represent the face features of the user, improve the accuracy of the face library, and thus can achieve better application effects when using the face set for applications such as video originality recognition.
[0102] Figure 6The following is a structural block diagram of a video recognition device provided by an embodiment of the present invention. This device can be implemented by software and / or hardware, and is generally integrated in a computer device. It can perform the originality recognition of a video by executing a video recognition method. As Figure 6 shown, the device includes:
[0103] A face feature extraction module 601, configured to extract face features to be recognized from a target video uploaded by a target user for recognition;
[0104] An originality recognition module 602, configured to compare the face features to be recognized with each face feature in a target face library corresponding to the target user, and recognize the originality of the target video according to the comparison result, where the target face library is obtained by using the image processing method provided by the embodiment of the present invention.
[0105] In the video recognition device provided by the embodiment of the present invention, face features to be recognized are extracted from a target video uploaded by a target user, and the face features to be recognized are compared with each face feature in a target face library corresponding to the target user obtained by using the image processing method provided by the embodiment of the present invention. The originality of the target video is recognized according to the comparison result. Since this face library can more comprehensively and accurately represent the face features of the user, it can more accurately recognize the originality of the video, reduce the probability of misrecognition, facilitate further targeted related processing of the target video, improve the video processing efficiency, and enhance the rationality of video processing.
[0106] The embodiment of the present invention provides a computer device, and the image processing device provided by the embodiment of the present invention can be integrated in this computer device. Figure 7 The following is a structural block diagram of a computer device provided by an embodiment of the present invention. The computer device 700 includes a memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702. When the processor 702 executes the computer program, it implements the image processing method and / or the video recognition method provided by the embodiment of the present invention.
[0107] The embodiment of the present invention further provides a storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute the image processing method and / or the video recognition method provided by the embodiment of the present invention.
[0108] The image processing device, video recognition device, device, and storage medium provided in the above embodiments can execute the corresponding methods provided by the embodiments of the present invention, and have the corresponding functional modules and beneficial effects of the executed methods. For technical details not described in detail in the above embodiments, reference can be made to the methods provided in any embodiment of the present invention.
[0109] Note that the above is only a preferred embodiment of the present invention. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments only. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the claims.
Claims
1. An image processing method, characterized in that, Including: Obtain the avatar image of the target user in the preset application. When the first face feature that meets the preset requirements is extracted from the avatar image, add the first face feature to the benchmark face feature set; Obtain the second face feature corresponding to the target face with an appearance frequency that meets the preset frequency requirement from the historical videos published by the target user in the preset application, and add the second face feature to the candidate face feature set; Based on the benchmark face feature set, screen the face features in the candidate face feature set, and determine the extended face feature set according to the screening result; Construct the target face library corresponding to the target user according to the benchmark face feature set and the extended face feature set; Among them, obtaining the second face feature corresponding to the target face with an appearance frequency that meets the preset frequency requirement from the historical videos published by the target user in the preset application includes: Extract face features from the video frames in the historical videos published by the target user in the preset application, and add the extracted face features to the alternative face feature pool as alternative face features; Perform preset clustering processing on the alternative face features in the alternative face feature pool to obtain multiple clusters; Count the number of alternative face features included in each cluster, and determine the cluster with the cumulative quantity meeting the preset quantity requirement as the high-frequency cluster. Among them, the face to which the alternative face features in the high-frequency cluster belong is recorded as the target face with an appearance frequency that meets the preset frequency requirement; For each alternative face feature in the high-frequency cluster, calculate the first similarity between the current alternative face feature and each alternative face feature other than the current alternative face feature in the high-frequency cluster to which it belongs, and calculate the first value, where the first value is the sum of the first similarities; Obtain the alternative face feature corresponding to the first value in the high-frequency cluster that meets the first preset value requirement as the second face feature; Among them, the method further includes: When the first face feature that meets the preset requirements cannot be extracted from the avatar image, determine the cluster with the largest cumulative quantity as the target cluster, where the high-frequency cluster includes the target cluster; Add the alternative face feature corresponding to the first value in the target cluster that meets the second preset value requirement to the benchmark face feature set.
2. The method according to claim 1, characterized in that, Before screening the face features in the candidate face feature set based on the benchmark face feature set, it further includes: Add the third face feature that does not meet the preset requirements extracted from the avatar image to the candidate face feature set.
3. The method according to claim 1, characterized in that After obtaining the avatar image of the target user in the preset application, it further includes: Input the avatar image into a preset classification model, and determine the category of the avatar image according to the output result of the preset classification model. The labels of the training sample images corresponding to the preset classification model include available and unavailable; Extract face features from the available avatar images to obtain the first face features that meet the preset requirements.
4. The method according to claim 1, wherein Screening the face features in the candidate face feature set based on the reference face feature set, and determining an extended face feature set according to the screening result, including: For each candidate face feature in the candidate face feature set, calculate the second similarity between the current candidate face feature and each reference face feature in the reference face feature set. When there is at least one second similarity greater than the preset similarity threshold, determine the current candidate face feature as an extended face feature and add it to the extended face feature set.
5. A video recognition method, characterized in that, Including: Extracting the face features to be recognized from the target video to be recognized uploaded by the target user; Comparing the face features to be recognized with each face feature in the target face library corresponding to the target user, and identifying the originality of the target video according to the comparison result, where the target face library is obtained by using the image processing method described in any one of claims 1-4.
6. An image processing apparatus, characterized in that, Including: An avatar image processing module, configured to obtain an avatar image of the target user in a preset application program, and add the first face feature to the reference face feature set when the first face feature that meets the preset requirements is extracted from the avatar image; A historical video processing module, configured to obtain the second face features corresponding to the target face with an appearance frequency meeting the preset frequency requirement from the historical videos published by the target user in the preset application program, and add the second face features to the candidate face feature set; A face feature screening module, configured to screen the face features in the candidate face feature set based on the reference face feature set, and determine an extended face feature set according to the screening result; A face library construction module, configured to construct a target face library corresponding to the target user according to the reference face feature set and the extended face feature set; Wherein, the historical video processing module is specifically configured to: Extract face features from the video frames in the historical videos published by the target user in the preset application program, and add the extracted face features as alternative face features to the alternative face feature pool; Perform a preset clustering process on the alternative face features in the alternative face feature pool to obtain multiple clusters; Count the number of alternative face features included in each cluster, and determine the cluster with the cumulative number meeting the preset number requirement as a high-frequency cluster, where the face to which the alternative face features in the high-frequency cluster belong is recorded as the target face with an appearance frequency meeting the preset frequency requirement; For each alternative face feature in the high-frequency cluster, calculate the first similarity between the current alternative face feature and each of the other alternative face features in the high-frequency cluster except the current alternative face feature, and calculate a first value, where the first value is the sum of the first similarities; Obtain the alternative face features corresponding to the first value meeting the first preset value requirement in the high-frequency cluster as the second face features; Wherein, the avatar image processing module is further configured to: When the first face feature that meets the preset requirements cannot be extracted from the avatar image, determine the cluster with the largest cumulative number as the target cluster, where the high-frequency cluster includes the target cluster; Add the alternative face features in the target cluster whose corresponding first values meet the requirements of the second preset value to the reference face feature set.
7. A video recognition device, characterized in that, Including: A face feature extraction module, configured to extract face features to be recognized from a target video uploaded by a target user for recognition; An originality recognition module, configured to compare the face features to be recognized with each face feature in the target face library corresponding to the target user, and recognize the originality of the target video according to the comparison result, where the target face library is obtained by using the image processing method described in any one of claims 1-4.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1-5 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method described in any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Age recognition method and device and storage medium
CN110321863A
Infringement video recognition method and device, electronic equipment and storage medium
CN111708988A