Image processing method and device, storage medium and computer device
By clustering and analyzing the features of facial images in images, a user relationship graph is generated, which solves the problem of low efficiency under traditional methods and achieves efficient and accurate user relationship prediction.
Patent Information
- Application Number
- CN201910690640.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-04-09
AI Technical Summary
Traditional methods of mining person relationship graphs are inefficient because they require searching a large amount of user information.
By clustering the facial images in the image, the facial images of the same user are identified, and the user relationship characteristics are determined based on the number of times the user appears in the same image. The user relationship graph is generated using the pre-trained user relationship decision model.
It improves the mining efficiency of user relationship graphs, accurately identifies user attributes and predicts user relationships without searching for user relationship information from various network sources.
Smart Images

Figure CN110414433B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, computer-readable storage medium, and computer equipment. Background Art
[0002] With the advancement of computer and network technologies, social media software has become increasingly diverse and convenient. Social media software has significantly expanded and enriched people's social lives. Among these, user family and friend relationships are crucial. To further develop and utilize social media applications, effectively mining the relationship graph has become a pressing technical challenge.
[0003] Traditional methods for mining person-relationship graphs typically require searching for user relationship information from various network sources and integrating it to generate a user relationship graph. This traditional method of mining user relationship graphs is inefficient due to the need to search a large amount of user information. Summary of the Invention
[0004] Based on this, it is necessary to provide an image processing method, apparatus, computer-readable storage medium and computer equipment to address the technical problem of low efficiency in mining user relationship graphs.
[0005] An image processing method, comprising:
[0006] Acquire an image to be processed and a face image included in the image to be processed;
[0007] performing clustering processing on the facial images, clustering facial images that meet facial similarity conditions into the same facial category;
[0008] Determine the user attribute labels corresponding to the face categories based on the face images included in each face category;
[0009] Determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed;
[0010] The user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features are input into the pre-trained user relationship decision model to obtain a user relationship graph.
[0011] An image processing device, characterized in that the device comprises:
[0012] An acquisition module, configured to acquire an image to be processed and a face image included in the image to be processed;
[0013] A face clustering module, configured to perform clustering processing on the face images, and cluster the face images that meet the face similarity condition into the same face category;
[0014] A determination module, configured to determine a user attribute label corresponding to each face category based on the face images included in each face category;
[0015] The determination module is further configured to determine the face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed;
[0016] The user relationship prediction module is used to input the user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features into the pre-trained user relationship decision model to obtain a user relationship graph.
[0017] A computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the following steps:
[0018] Acquire an image to be processed and a face image included in the image to be processed;
[0019] performing clustering processing on the facial images, clustering facial images that meet facial similarity conditions into the same facial category;
[0020] Determine the user attribute labels corresponding to the face categories based on the face images included in each face category;
[0021] Determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed;
[0022] The user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features are input into the pre-trained user relationship decision model to obtain a user relationship graph.
[0023] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:
[0024] Acquire an image to be processed and a face image included in the image to be processed;
[0025] performing clustering processing on the facial images, clustering facial images that meet facial similarity conditions into the same facial category;
[0026] Determine the user attribute labels corresponding to the face categories based on the face images included in each face category;
[0027] Determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed;
[0028] The user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features are input into the pre-trained user relationship decision model to obtain a user relationship graph.
[0029] The above-mentioned image processing method, apparatus, computer-readable storage medium, and computer equipment cluster the facial images in the image to be processed, clustering facial images that meet facial similarity conditions into the same facial category, thereby accurately identifying the facial images corresponding to different users. Based on the facial images corresponding to different users, the user attribute labels corresponding to each user are determined. Furthermore, the corresponding facial co-occurrence features are determined based on the number of times the faces of different users appear in the same image to be processed, that is, the number of photos of different users together. The pre-trained user relationship decision model can effectively integrate the user attribute labels corresponding to different users and the corresponding facial co-occurrence features to obtain a user relationship graph. In this way, the image to be processed can be directly processed to mine the attribute information of different users in the image to be processed, as well as the degree of intimacy between different users, thereby accurately and efficiently predicting the user relationship graph. This eliminates the need to search for user relationship information from various network sources, greatly improving the efficiency of mining the user relationship graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A diagram showing an application environment of an image processing method in one embodiment;
[0031] Figure 2 1 is a flow chart of an image processing method according to an embodiment;
[0032] Figure 3 A flowchart illustrating the steps of clustering facial images to cluster facial images that meet facial similarity conditions into the same facial category in one embodiment;
[0033] Figure 4 1 is a flowchart illustrating a step of determining a user attribute label corresponding to each face category based on the face images included in each face category in one embodiment;
[0034] Figure 5 A schematic diagram of a result of determining a user age label and a user gender label corresponding to a face image in one embodiment;
[0035] Figure 6 A schematic diagram of a process for inputting user attribute labels corresponding to different face categories and corresponding face co-occurrence features into a pre-trained user relationship decision model to obtain a user relationship graph in one embodiment;
[0036] Figure 7 A schematic diagram of a structure for determining user relationships through a decision tree model in one embodiment;
[0037] Figure 8 A schematic diagram of a process for performing face detection on an image to be processed and extracting a face image from the image to be processed in one embodiment;
[0038] Figure 9 1. A schematic diagram of a process effect of performing face detection on an image to be processed and extracting a face image from the image to be processed using a three-level network structure in one embodiment;
[0039] Figure 10 is a flowchart of an image processing method in a specific embodiment;
[0040] Figure 11 is a flowchart of an image processing method in another specific embodiment;
[0041] Figure 12 is a structural block diagram of an image processing device in one embodiment;
[0042] Figure 13 is a structural block diagram of an image processing device in another embodiment;
[0043] Figure 14 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0045] Figure 1 FIG. 1 is an application environment diagram of an image processing method in an embodiment. Figure 1 , the image processing method is applied to an image processing system. The image processing system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can specifically be a desktop terminal or a mobile terminal, and the mobile terminal can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster consisting of multiple servers. Both the terminal 110 and the server 120 can be used alone to execute the image processing method provided in the embodiments of the present application. The terminal 110 and the server 120 can also be used in conjunction to execute the image processing method provided in the embodiments of the present application.
[0046] It should be noted that the embodiments of the present application involve a variety of models, specifically including machine learning models and decision tree models. Among them, the machine learning model is a model that has certain capabilities after learning through samples. One machine learning model in the embodiments of the present application is a face detection model that can detect faces from images through sample learning. Another machine learning model in the embodiments of the present application is a face feature extraction model that has the ability to extract face features through sample learning. The decision tree model in the embodiments of the present application is a decision tree model that can judge the relationship between people through sample learning. Among them, the machine learning model can adopt a neural network model, such as a CNN (Convolutional Neural Networks) model, etc. Of course, the machine learning model can also adopt other types of models, which are not limited in the embodiments of the present application.
[0047] like Figure 2 As shown, in one embodiment, an image processing method is provided. This embodiment mainly uses the method applied to a computer device as an example, and the computer device can specifically be the above-mentioned Figure 1 The terminal 110 or the server 120 in FIG. Figure 2 , the image processing method specifically includes the following steps:
[0048] S202: Acquire an image to be processed and a face image included in the image to be processed.
[0049] The image to be processed is an image to be processed and analyzed. The computer device can execute the image processing method of the embodiments of the present application on the image to be processed to obtain a user relationship map. The number of images to be processed can be one or more. The face image is an image generated based on the area of the image to be processed that contains a face.
[0050] Specifically, the computer device can obtain the image to be processed from a local or other computer device, perform face detection on the image to be processed, extract the face in the image to be processed and generate a face image. It can be understood that when an image to be processed includes one face, one face image can be extracted; when an image to be processed includes multiple faces, multiple face images can be extracted. Among them, the words "multiple" or "multiple images" mentioned in the embodiments of the present application all mean "more than one" or "more than one".
[0051] In one embodiment, step S202, that is, the step of obtaining the image to be processed and the facial image included in the image to be processed, specifically includes obtaining a user image set; filtering out images including faces from the user image set as images to be processed; performing face detection on the image to be processed, and extracting the facial image from the image to be processed.
[0052] Among them, the user image set is a set composed of multiple user images, which can specifically be a set of photos uploaded by users, such as user albums, etc. Specifically, the user can collect user images through the camera of the terminal and store them in the terminal, or the user can obtain user images from other computer devices through network connection or data cable connection and store them in the terminal. When the image processing method is directly executed by the terminal, the terminal can directly obtain the locally stored user images and execute the image processing method mentioned in the embodiment of this application. When the image processing method is executed by the server, the terminal can upload the user images to the server through the locally running social media software, and the server obtains the user images uploaded by the terminal and executes the image processing method mentioned in the embodiment of this application. Among them, the social media software can specifically be software that manages the user's media data, such as album manager software. Users can manage their own media data such as photos or videos through album manager software.
[0053] Furthermore, the computer device may filter out images including faces from the user image collection as images to be processed, and perform face detection on the images to be processed using a pre-trained face detection model to extract face images from the images to be processed.
[0054] In one embodiment, the user image collection may include images of multiple categories, such as human image categories, natural landscape image categories, food image categories, or architectural image categories. The computer device may classify the user images to obtain the category corresponding to each user image, and then select the images corresponding to the human image category as images to be processed.
[0055] In one embodiment, the computer device may perform face detection on each user image in the user image collection using a pre-trained face detection model, determine that the image including the face is the image to be processed, and extract the face image from the image to be processed.
[0056] In one embodiment, a computer device may input an image to be processed into a pre-trained face detection model. The model then performs face detection on the input image and outputs the coordinates of a face detection frame used to mark the face area. A face image may be generated based on each pixel in the area determined by the coordinates of the face detection frame.
[0057] In one embodiment, a computer device may obtain a sample image and a sample face detection frame used to mark a facial region in the sample image. The computer device may use the sample face detection frame as a training label to train a face detection model based on the sample image and the training label. During model training, the computer device continuously adjusts model parameters until a training stop condition is met, thereby terminating the training and obtaining a trained face detection model. Specifically, the training stop condition may include reaching a preset number of iterations or the difference between the actual output and the training label being less than a preset difference.
[0058] In the above embodiment, images including human faces are screened out from the user image collection as images to be processed, and face detection is performed on the images to be processed, so that human face images can be extracted quickly and accurately from the images to be processed.
[0059] S204: performing clustering processing on the facial images, clustering facial images that meet facial similarity conditions into the same facial category.
[0060] Clustering is the process of dividing a collection of physical or abstract objects into multiple clusters of similar objects. A face category is the category to which a face image belongs. In this embodiment of the present application, face images corresponding to different user objects belong to different face categories, while face images corresponding to the same user object belong to the same face category.
[0061] Specifically, the computer device may extract facial features from each facial image, calculate feature similarity based on the facial features corresponding to the different facial images, and determine the degree of similarity between the two facial images based on the feature similarity. Furthermore, the computer device may cluster facial images whose similarity exceeds a preset threshold into the same facial category. The feature similarity can be calculated by calculating the Euclidean distance or cosine distance between the features, which is not specifically limited in the embodiments of this application.
[0062] In one embodiment, the computer device may assign corresponding user identifiers to different face categories. Each face category corresponds to one user identifier. The user identifier is used to uniquely identify the user and may be a number, letter, or character.
[0063] In one embodiment, step S204, i.e., performing clustering processing on the facial images to cluster facial images that meet the facial similarity condition into the same facial category, includes:
[0064] S302: Extract facial features from the facial image using a pre-trained facial feature extraction model.
[0065] Among them, the facial feature extraction model is a pre-trained model for extracting facial features. The facial feature extraction model can specifically be a lightweight deep convolutional network FaceNet model, a VGGNet (Visual Geometry Group) model, a ResNet residual neural network (Residual Neural Network, energy efficiency evaluation system), or other convolutional network models, which are not limited in the embodiments of the present application. Facial features are features that can be used to describe the face as a whole, which are extracted after the facial feature extraction model processes the facial image, and are also called facial feature vectors.
[0066] Specifically, the computer device can obtain a pre-trained facial feature extraction model, input the facial image into the facial feature extraction model, perform feature extraction on the facial image through the model parameters and network structure of the facial feature extraction model, and extract facial features from the facial image.
[0067] In one embodiment, a computer device can train a facial feature extraction model using sample images and training labels. The sample image is an image that includes a face, and the training label is the face category to which the face image belongs. Different face categories correspond to different user identifiers. The computer device can input the sample image into the facial feature extraction model, and output the probability that the sample image belongs to each face category through the output layer of the facial feature extraction model. The computer device can adjust the model parameters in a direction that reduces the difference based on the difference between the output result of the facial feature extraction model and the training label. The calculation is iterated continuously until the training stop condition is met. In this way, a facial feature extraction model can be trained based on the sample image and training labels. Furthermore, the computer device can input the face image into the trained facial feature extraction model, and extract facial features through the intermediate layer of the facial feature extraction model.
[0068] In one embodiment, step S302 specifically includes: adjusting the size of the facial image to a standard size to obtain a standard facial image; normalizing the standard facial image; inputting the normalized standard facial image into a pre-trained facial feature extraction model, and extracting facial features from the standard facial image through the facial feature extraction model.
[0069] The standard size is a pre-set, standardized size. Specifically, the computer device can resize the facial image to the standard size to obtain a standard facial image. The pixel values in the standard facial image are then normalized and input into a face detection model to extract facial features of a preset dimension, such as 128 dimensions.
[0070] In one embodiment, the computer device may resize the facial image to a standard size, such as 112*112*3, and then arrange the pixel values of each pixel in the standard facial image in the order of RGB (red, green, and blue) channels to obtain an RGB facial image. The computer device may normalize the RGB facial image and input the normalized RGB facial image into a face detection model to obtain facial features. Specifically, the normalization of the RGB facial image may include dividing the pixel value of each pixel in the RGB image by 256 to obtain the processed pixel value of each pixel.
[0071] In the above embodiment, by resizing and normalizing the facial images, facial images of various sizes and specifications can be standardized, which facilitates processing by the facial feature extraction model.
[0072] S304: Calculate similarities between different facial images based on facial features corresponding to each facial image.
[0073] Specifically, the computer device can combine multiple facial images in pairs, and calculate the feature similarity based on the facial features corresponding to each of the two facial images. The feature similarity can be considered as the similarity between the two facial images.
[0074] S306 , clustering facial images that meet facial similarity conditions into the same facial category based on similarities between different facial images.
[0075] Specifically, the computer device may cluster two facial images within each facial image into the same facial category based on the similarity between them, if the similarity exceeds a preset threshold. In other words, facial images with a similarity exceeding the threshold can be preliminarily considered to be from the same user. This process continues in a loop, clustering facial images with a similarity exceeding the threshold into the same facial category.
[0076] In one embodiment, a computer device determines a similarity matrix corresponding to all facial images based on the similarity between any two facial images. Assuming that all users are initially individual, each similarity value in the similarity matrix is compared with a preset similarity threshold. When the similarity between two facial images exceeds the threshold, the two facial images can be associated, for example, by connecting them with a virtual line segment. In the first iteration, if multiple candidate facial images with similarities greater than the threshold exist for the same facial image in other facial images, the candidate facial image with the greatest similarity can be clustered with the facial image into a single category. In other words, the identity information corresponding to each facial image is determined by the facial image with the greatest similarity among the associated facial images. At this point, the association between the facial image and the facial image with the greatest similarity can be preserved, while the associations between the facial image and the other facial images can be removed. After the first round of iterations, if a facial image corresponds to multiple face categories, it can be clustered into the face category with the largest number of face images. In other words, the face category corresponding to each face image is determined by the face category with the most identical face images. The steps after the first round of iterations are repeated until convergence, resulting in at least one face category. Each face category includes at least one face image. It should be understood that face images within the same face category can be considered to correspond to the same user.
[0077] In the above embodiment, the pre-trained facial feature extraction model accurately extracts facial features from facial images and calculates the similarity between different facial images based on the facial features corresponding to each facial image. This allows facial images that meet the facial similarity criteria to be clustered into the same face category based on the similarity between the different facial images, thereby accurately performing face recognition and clustering on the facial images to obtain a facial image corresponding to each user.
[0078] S206: Determine the user attribute label corresponding to each face category based on the face images included in each face category.
[0079] The user attribute tag is a category tag used to mark the user's attribute characteristics, which can also be understood as the user's attribute characteristics. The user attribute specifically includes at least one of the user's age, user gender, user clothing, user expression, user behavior, etc.
[0080] Specifically, for each face category, the computer device may determine the user attribute label corresponding to the face category based on the facial images included in the face category. In one embodiment, the computer device may input the facial images included in the face category into a pre-trained user attribute discrimination model, and the user attribute discrimination model may output the corresponding user attribute label.
[0081] In one embodiment, the user attribute tag includes a user age tag and a user gender tag; and determining the user attribute tag corresponding to each face category based on the face images included in each face category includes:
[0082] S402 : For each face category, select a representative face image that meets a large face condition from the face images included in the face category.
[0083] The large face condition may specifically be that the facial area in the facial image is the largest. The facial area may specifically be the area of the facial image. Specifically, for each facial category, the computer device may calculate the facial area corresponding to each facial image included in the facial image category and use the facial image with the largest facial area as the representative facial image. Alternatively, the computer device may use the facial image with the largest facial area among the facial images as the representative facial image.
[0084] S404: Input representative face images corresponding to various face categories into the user age discrimination model to obtain user age labels corresponding to various face categories.
[0085] The user age label is a label indicating the user's age category, such as infant, child, teenager, young adult, middle-aged, and elderly. Alternatively, different numerical values may be used to represent different user age labels, such as a value of "1" for "infant," a value of "2" for "child," and a value of "3" for "teenager."
[0086] Specifically, for each face category, the computer device inputs the representative face image corresponding to the face category into the pre-trained user age discrimination model to obtain the user age label corresponding to the face category.
[0087] In one embodiment, the computer device may construct a user age discrimination model based on a deep convolutional neural network, such as a three-layer convolutional neural network. The user age discrimination model is trained using sample facial images and sample age labels corresponding to the facial images, and the model parameters of the user age discrimination network are obtained using a gradient descent algorithm. When the computer device inputs a representative facial image into the trained user age discrimination model, the user age discrimination model determines the probability that the representative facial image belongs to each user age label. The user age label corresponding to the highest probability is used as the user age label corresponding to the facial category.
[0088] S406: Input representative face images corresponding to each face category into the user gender discrimination model to obtain user gender labels corresponding to each face category.
[0089] The user gender label is a label indicating the user's gender category, such as male and female. Alternatively, different values may be used to represent different user age labels, such as using a value of "1" to represent "male" and a value of "2" to represent "female."
[0090] Specifically, for each face category, the computer device inputs the representative face image corresponding to the face category into the pre-trained user gender discrimination model to obtain the user gender label corresponding to the face category.
[0091] In one embodiment, the computer device can construct a user gender discrimination model based on a deep convolutional neural network, such as a two-layer convolutional neural network. The user gender discrimination model is trained using sample facial images and sample gender labels corresponding to the facial images, and the model parameters of the user gender discrimination network are obtained using a gradient descent algorithm. When the computer device inputs the representative facial image into the trained user gender discrimination model, the user gender discrimination model can obtain the probability that the representative facial image belongs to each user gender label, and the user gender label corresponding to the maximum probability is used as the user gender label corresponding to the facial category.
[0092] refer to Figure 5 , Figure 5 FIG. 1 is a schematic diagram showing the result of determining the user age label and the user gender label corresponding to the face image in one embodiment. Figure 5 As shown, the computer device can input the facial image into the user age discrimination model and the user gender discrimination model respectively to obtain the user age label and the user gender label, for example, the user age label is "child" and the user gender label is "boy".
[0093] In the above embodiment, the facial image that meets the large face condition is used as a representative facial image in this type of face category, so that the representative facial image can be input into the user age discrimination model and the user gender discrimination model in sequence, thereby obtaining the user age label and user gender label that can represent the face category, and the user attribute label can be accurately and quickly determined.
[0094] It is understandable that user attribute labels also include user clothing labels and user eye labels, etc., which are not limited in the embodiments of the present application. Among them, user clothing labels are labels that represent user clothing categories, such as casual wear, formal wear, sportswear and other category labels. User eye labels are labels that represent user eye categories, such as loving eyes categories, or intimate eyes categories and other labels. Accordingly, the computer device can input the representative face image into the model for performing user clothing task discrimination and the model for performing user eye task discrimination in sequence, thereby obtaining the user clothing label and user eye label corresponding to the face category.
[0095] S208 , determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed.
[0096] Among them, the face co-occurrence feature is a feature where facial images corresponding to different users appear simultaneously in the same image, for example, the number of group photos of different users. Specifically, the computer device can assign a user ID to each facial category, and the user ID is used to identify the user corresponding to the facial category. For each user identified by the user ID, the computer device can count the number of group photos corresponding to the two users, that is, the number of times the facial images corresponding to the two user IDs appear in the same image to be processed, and use the number of group photos as the face co-occurrence feature corresponding to the two users.
[0097] In one embodiment, after performing face detection on the image to be processed and extracting a facial image from the image to be processed, the computer device can record the image to be processed from which the facial image originated, thereby determining the correspondence between the facial image and the image to be processed. When the facial images corresponding to two users appear in the same image to be processed, the number of group photos taken with the two users is incremented by one. Based on this, the computer device can calculate the number of group photos taken between all pairs of users.
[0098] In one embodiment, step S208, that is, determining the face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed, includes: combining the face categories in pairs to obtain at least one group of face category combinations; for each group of face category combinations, determining the face co-occurrence features corresponding to the face category combination based on the number of times face images belonging to two face categories in the face category combination appear in the same image to be processed.
[0099] Specifically, the computer device can pairwise combine all facial categories, combining two different facial categories into one facial category combination, that is, combining two user identifiers into one group. For each facial category combination, the computer device can determine the facial co-occurrence feature corresponding to that facial category combination based on the number of times the facial images corresponding to the two user identifiers appear simultaneously in the same image to be processed, that is, the number of photos taken of the users corresponding to the two user identifiers. Based on this, the number of times all users take photos together can be calculated, thereby obtaining the facial co-occurrence feature corresponding to the corresponding facial category combination.
[0100] S210: Input the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into the pre-trained user relationship decision model to obtain a user relationship graph.
[0101] The user relationship decision model is a model used to make decisions on user relationships, and can be a decision tree model or a deep learning network model, etc. The user relationship graph is a graph used to represent the relationships between different users.
[0102] Specifically, a computer device pre-trains a user relationship decision model using training samples and training labels to obtain a pre-trained user relationship decision model. User attribute labels corresponding to different face categories and corresponding face co-occurrence features are then input into the pre-trained user relationship decision model to obtain a user relationship graph. The training samples can specifically be user attribute labels and corresponding face co-occurrence features corresponding to different users, and the training labels can specifically be user relationship information between two different users.
[0103] In one embodiment, when a computer device can combine face categories in pairs to obtain at least one face category combination; for each face category combination, based on the number of times that facial images belonging to the two face categories in the face category combination appear in the same image to be processed, the face co-occurrence features corresponding to the face category combination are determined. Furthermore, step S210, that is, the step of inputting the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into a pre-trained user relationship decision model to obtain a user relationship graph, includes: for each face category combination, inputting the user attribute labels corresponding to the face categories included in the face category combination and the corresponding face co-occurrence features into a pre-trained decision tree model to obtain user relationship information corresponding to the face category combination; based on the user relationship information corresponding to each face category combination, a user relationship graph corresponding to the image to be processed is constructed.
[0104] The decision tree model is a model that classifies input data. It features fast computation, easy-to-interpret results, and strong robustness. User relationship information represents the relationships between different users, such as user A being user B's father, user C and user D being friends, and user E and user F being sisters.
[0105] In one embodiment, for each face category combination, the computer device may input the user attribute labels corresponding to the two different face categories and the corresponding facial co-occurrence features into a pre-trained decision tree model. This user relationship model then outputs the user relationship information between the two face categories corresponding to this face category combination. Accordingly, for each face category combination, the computer device can obtain the user relationship information between the two face categories using the aforementioned method. Furthermore, based on the user relationship information corresponding to each face category combination, the computer device may integrate all user relationship information to construct a user relationship graph corresponding to the image to be processed.
[0106] For example, refer to Figure 6 , Figure 6 FIG2 is a flow chart of inputting user attribute labels corresponding to different face categories and corresponding face co-occurrence features into a pre-trained user relationship decision model to obtain a user relationship graph in one embodiment. Figure 6 As shown, the computer device can input the user age labels, user gender labels, and corresponding face co-occurrence features corresponding to the two types of face categories into the decision tree model, and predict user relationships through the decision tree model to obtain a user relationship map.
[0107] refer to Figure 7 , Figure 7 FIG. 1 is a schematic diagram showing a structure of determining user relationships through a decision tree model in one embodiment. Figure 7As shown, when a computer device needs to determine the relationship between user A and user B, it may input user A's corresponding age label (A) and gender label (A), user B's corresponding age label (B) and gender label (B), and co-occurrence features of user A and user B's profile pictures, such as the number of photos taken together, into a decision tree model. The decision tree model then makes a step-by-step judgment based on the non-leaf nodes. For example, if user A's age is <20, the left branch is taken; if user A's age is >=20, the right branch is taken. This step-by-step judgment outputs user relationship information between A and B. This user relationship information between A and B may include information such as A is B's father, A is B's mother, A is B's son, or A and B are unrelated.
[0108] In the above embodiment, by sequentially inputting the user attribute labels corresponding to the two face categories included in the face category combination and the corresponding face co-occurrence features into the pre-trained decision tree model, the user relationship information corresponding to the face category combination can be obtained quickly and accurately until the user relationship information corresponding to each group of face category combinations is obtained, and based on this, a user relationship graph corresponding to the image to be processed is constructed.
[0109] In one embodiment, the computer device can determine the facial co-occurrence features between different facial categories based on the number of times facial images belonging to different facial categories appear in the same image to be processed. Based on the facial co-occurrence features, the facial co-occurrence features and the user attribute labels corresponding to the two facial categories corresponding to the facial co-occurrence features are respectively used as a set of input data. In accordance with this, the computer device can determine at least one set of input data. The computer device can input multiple sets of input data into different input channels of the user relationship decision model respectively, and process the respective input data through the intermediate layers corresponding to the multiple input channels respectively to obtain the user relationship information corresponding to each set of input data. The computer device then integrates the user relationship information corresponding to each set of input data through the fusion layer structure in the user relationship decision model to obtain a user relationship graph.
[0110] In one embodiment, the decision tree model can specifically include multiple trees, such as a model with between 2,000 and 3,000 trees. For a decision tree model with multiple trees, the input data for each decision tree is the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features. However, the output of subsequent trees is adjusted and modified based on the output of previous trees until the last tree outputs user relationship information. The computer device can use the user relationship information output by the last tree as the user relationship information corresponding to the input data.
[0111] In one embodiment, after the computer device determines the structure of the decision tree model to be constructed, it can train the decision tree model based on sample data and sample labels. Specifically, the sample data can include the user attribute labels of two users and their corresponding profile picture contribution features; the sample labels can include the user relationship information between the two users. The computer device can use a dynamic learning rate approach to perform algorithm learning, employing a cross-entropy loss function to automatically learn the decision tree model, continuously learning until a preset number of iterations is reached. In this way, the trained decision tree model can accurately discern user relationships between different users.
[0112] For example, the computer device can use the standard Xgboost (eXtreme Gradient Boosting) decision tree framework and determine the number of leaf nodes of each tree in the decision tree model and the depth of the decision tree, such as constructing a decision tree with 5, 10 or 15 leaf nodes and a depth of 6, 7, or 8. After determining the tree architecture, for example, 2,000 to 3,000 trees are used to jointly construct the decision tree model. The computer device can train the decision tree model based on the training data and training labels. During the training process, a dynamic learning speed method can be used for algorithm learning, and the initial learning rate is set to 0.001, for example. The cross entropy loss function is used to automatically learn the decision tree model, and learning is continuously stopped until the preset number of iterations is reached, and the preset number of iterations is, for example, 10,000 times.
[0113] The above-mentioned image processing method clusters the facial images in the image to be processed, and clusters the facial images that meet the facial similarity conditions into the same facial category, thereby accurately identifying the facial images corresponding to different users. Then, based on the facial images corresponding to different users, the user attribute labels corresponding to each user are determined. In addition, the corresponding facial co-occurrence features are determined based on the number of times the faces of different users appear in the same image to be processed, that is, the number of photos of different users together. The pre-trained user relationship decision model can effectively integrate the user attribute labels corresponding to different users and the corresponding facial co-occurrence features to obtain a user relationship graph. In this way, the image to be processed can be directly processed to mine the self-attribute information of different users in the image to be processed, as well as the degree of intimacy between different users, so as to accurately and efficiently predict the user relationship graph, without having to search for user relationship information from various network sources, greatly improving the efficiency of mining the user relationship graph.
[0114] In one embodiment, the steps of performing face detection on the image to be processed and extracting the face image from the image to be processed specifically include:
[0115] S802: Input the image to be processed into a face detection model, process the image to be processed through a first convolutional neural network in the face detection model, and obtain a candidate face detection frame for marking the face.
[0116] Among them, the face detection model is a model that performs face detection on the input image to determine the location of the face in the input image. The face detection model may specifically include a network structure of multiple layers, such as a first convolutional neural network, a second convolutional neural network, and a third convolutional neural network. Among them, the network structure of each convolutional neural network can be the same or different, and the parameters of each convolutional neural network are different. During model training, each convolutional neural network in the face detection model can be trained independently or jointly, etc., which is not limited in the embodiments of the present application.
[0117] Specifically, the computer device may input the image to be processed into the first convolutional neural network in the face detection model. The first convolutional neural network processes the image to be processed, extracts and analyzes features of the image, and outputs a candidate face detection frame for marking the face. The candidate face detection frame may be determined by the coordinates of two vertices: the upper left corner and the lower right corner.
[0118] S804, screening out a backup face image from the candidate face images determined by the candidate face detection frame through the second convolutional neural network in the face detection model; the probability that the backup face image is a valid face image satisfies a first probability condition.
[0119] The first probability condition can specifically be that the probability value of the facial image to be screened being a valid facial image is greater than or equal to a preset threshold. Specifically, the computer device can extract the regional image defined by the candidate facial detection frame from the image to be processed to obtain the candidate facial image. The candidate facial images are then classified using the second convolutional neural network in the face detection model to obtain the probability value of each candidate facial image being a valid facial image. When the probability value corresponding to a candidate facial image is greater than or equal to the preset threshold, or when the probability values corresponding to all candidate facial images are sorted from largest to smallest, the candidate facial image with a ranking lower than the preset ranking is selected as a backup facial image.
[0120] In one embodiment, after processing by the first convolutional neural network, multiple candidate face detection frames are output, and the areas defined by these candidate face detection frames often overlap. This means that the same face may correspond to multiple candidate face detection frames. Based on this, the computer device can use an NMS (Non-Maximum Suppression) algorithm to remove overlapping and invalid candidate face detection frames from the candidate face detection frames. Furthermore, the computer device can use the NMS algorithm to remove some invalid candidate face detection frames or merge candidate face detection frames with high overlap. Based on the processing results, the corresponding candidate face image is determined and then input into the second convolutional neural network for processing. This reduces the amount of data processed by the second convolutional neural network, eliminates invalid data, and thereby improves processing efficiency and accuracy.
[0121] S806, using the third convolutional neural network in the face detection model, screen out facial images to be clustered from the backup facial images; the probability that the screened facial images are valid facial images satisfies the second probability condition.
[0122] The second probability condition may specifically be that the probability value of the facial image to be screened being a valid facial image is greater than or equal to a preset threshold. Specifically, the computer device may input the backup facial images screened by the second convolutional neural network as input data into the third convolutional neural network. The backup facial images are classified and processed by the third convolutional neural network to obtain the probability value of each backup facial image being a valid facial image. When the probability value corresponding to a backup facial image is greater than the preset threshold, or when the probability values corresponding to all backup facial images are sorted from large to small, the backup facial image with a ranking less than the preset ranking is used as the facial image to be clustered, and the screened facial image can also be considered to be a valid facial image.
[0123] In one embodiment, the second convolutional neural network outputs multiple backup face detection frames after processing, and the areas defined by these multiple backup face detection frames often overlap. This means that the same face may correspond to multiple backup face detection frames. Based on this, the computer device can use the NMS algorithm to remove overlapping and invalid candidate face detection frames from the backup face detection frames. Furthermore, the computer device can use the NMS algorithm to remove some invalid backup face detection frames or merge backup face detection frames with high overlap. Based on the processing results, the corresponding backup face image is determined and then input into the third convolutional neural network for processing. This reduces the amount of data processed by the third convolutional neural network, eliminates invalid data, and thereby improves processing efficiency and accuracy.
[0124] In one embodiment, reference Figure 9, Figure 9 FIG. 1 is a schematic diagram showing the process of extracting a face image from an image to be processed by performing face detection on the image to be processed through a three-level network structure in one embodiment. Figure 9 As shown, the first convolutional neural network can be specifically a Small convolutional neural network, the second convolutional neural network can be specifically a Middle convolutional neural network, and the third convolutional neural network can be specifically a Big convolutional neural network. The computer device can input the image to be processed into the Small convolutional neural network to generate several candidate face detection frames, which are determined by the coordinates of the two vertices in the upper left corner and the lower right corner, and then use the classic NMS algorithm to remove invalid frames in the candidate face detection frames. Next, based on the candidate face detection frame obtained after the invalid frame is proposed, the candidate face image is determined, and the candidate face image is input into the Middle convolutional neural network for secondary filtering, and the invalid frame is removed again in combination with the NMS algorithm. Finally, the spare face image corresponding to the spare face detection frame obtained by proposing the invalid frame is input into the Big convolutional neural network to obtain the final valid face image. That is, Figure 9 The face image is marked with a valid face detection frame in the upper right corner.
[0125] In one embodiment, the computer device inputs the image to be processed into a trained face detection model to obtain a valid face image. The computer device can then modify the extracted valid face image and use it as a new training sample to retrain the face detection model. This can continuously optimize the face detection model and improve its effectiveness.
[0126] In the above embodiment, faces of different scales can be detected through a multi-level network structure, and based on the second convolutional neural network and the third convolutional neural network, the detection results are merged and filtered layer by layer to obtain a facial image that can accurately mark the face.
[0127] In one embodiment, step S210, that is, inputting the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into the pre-trained user relationship decision model, and the step of obtaining the user relationship graph specifically includes: inputting the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into the pre-trained decision tree model to obtain user relationship information between different face categories; determining the central face category in the face category; determining the user relationship information between each face category and the central face category based on the user relationship information between different face categories; and constructing a user relationship graph of the central user corresponding to the central face category based on the user relationship information between each face category and the central face category.
[0128] Specifically, the computer device can input the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into the pre-trained decision tree model to obtain user relationship information between different face categories. Among them, the user relationship information between different face categories refers to the user relationship information between two different face categories. The computer device can determine the central face category based on the face images included in each face category. There are many ways for the computer device to determine the central human category. For example, the computer device can randomly select a face category from multiple face categories as the central face category; or the computer device can use the face category selected by the user as the central face category; or the computer device can use the face category including the largest number of face images as the central face category; or other methods can be used to determine the central face category, which is not limited in the embodiments of the present application. Among them, the user corresponding to the central face category is the central user, that is, the user relationship graph to be constructed is constructed around this user.
[0129] Furthermore, the computer device can determine the user relationship information between each face category and the central face category based on the user relationship information between each face category. Based on the user relationship information between each face category and the central face category, a user relationship graph of the central user corresponding to the central face category is constructed.
[0130] In the above embodiment, the computer device may adjust the user relationship information according to the central face category, thereby constructing a user relationship graph centered on the central face category.
[0131] In one embodiment, determining the central face category in the face category specifically includes: for each face image included in each face category, calculating the area ratio between each face image and the corresponding image to be processed; calculating the number of face images included in each face category whose area ratio is greater than or equal to a preset ratio; and using the face category corresponding to the largest number as the central face category corresponding to the user relationship graph to be constructed.
[0132] Specifically, for each face category, the area ratio corresponding to each facial image in each face category is calculated. The area ratio corresponding to a facial image is the ratio between the area of the facial image and the area of the image to be processed corresponding to the facial image. The image to be processed corresponding to the facial image is the image from which the facial image is extracted, that is, the corresponding image to be processed is the source image of the facial image.
[0133] Furthermore, for each face category, the computer device may count the number of face images within that face category whose area ratio is greater than or equal to a preset ratio. The face category corresponding to the largest number is used as the central face category for the user relationship graph to be constructed. The user corresponding to the central face category is the central user; that is, the user relationship graph to be constructed is constructed around that user.
[0134] It can be understood that when the area ratio between a facial image and the corresponding image to be processed is greater than or equal to a preset threshold, the image to be processed can be considered a selfie of the user corresponding to the facial image. For multiple face categories corresponding to the image to be processed, the face category with the greatest number of selfies is considered the central face category. In other words, the user with the most selfies is considered the owner of the image to be processed, or the primary user of the image to be processed.
[0135] In the above embodiment, the area ratios between each facial image and the corresponding image to be processed are calculated, and the facial category that includes the largest number of facial images with area ratios greater than or equal to the preset ratio is used as the central facial category, so that a user relationship map centered on the central facial category can be constructed.
[0136] In one embodiment, the image processing method also includes a step of displaying a relationship map, which specifically includes: obtaining a user relationship map display instruction; and displaying a user relationship map including facial images and relationships between facial images according to the user relationship map display instruction.
[0137] Specifically, the computer device may be a terminal that displays an interactive interface. The terminal may detect a user-triggered user relationship graph display instruction within the interactive interface. For example, upon detecting a preset triggering operation, the user relationship graph display instruction may be triggered. Based on the user graph display instruction, a user relationship graph including facial images and relationships between facial images may be displayed within the interactive interface.
[0138] The preset trigger operation may be a touch operation, a cursor operation, a key operation, or a voice operation. The touch operation may be a touch click operation, a touch press operation, or a touch slide operation, and the touch operation may be a single-touch operation or a multi-touch operation; the cursor operation may be an operation of controlling the cursor to click or to press; and the key operation may be a virtual key operation or a physical key operation.
[0139] In one embodiment, a terminal is running social media software. A user can trigger a user relationship graph display instruction, such as by clicking or pressing an image import button displayed by the social media software, to trigger the user relationship graph display instruction. This instruction then imports user images from a local computer or other computer device into the social media software. Furthermore, the terminal or server can execute the image processing method described in the above embodiment to obtain a user relationship graph, which is then displayed in the form of images and text on the terminal's display interface.
[0140] In the above embodiment, by obtaining a user relationship graph display instruction and displaying a user relationship graph including facial images and the relationships between facial images, the user relationship graph can be displayed intuitively and clearly, making the corresponding user relationship information clear at a glance.
[0141] like Figure 10 As shown, in a specific embodiment, the image processing method includes the following steps:
[0142] S1002: Obtain a user image set.
[0143] S1004: Filter out images including faces from the user image collection as images to be processed.
[0144] S1006: Input the image to be processed into the face detection model, process the image to be processed through the first convolutional neural network in the face detection model, and obtain a candidate face detection frame for marking the face.
[0145] S1008, screening out a backup face image from the candidate face images determined by the candidate face detection frame through the second convolutional neural network in the face detection model; the probability that the backup face image is a valid face image satisfies a first probability condition.
[0146] S1010, using the third convolutional neural network in the face detection model, screen out facial images to be clustered from the backup facial images; the probability that the screened facial images are valid facial images satisfies a second probability condition.
[0147] S1012, adjusting the size of the facial image to a standard size to obtain a standard facial image.
[0148] S1014, performing normalization processing on the standard face image.
[0149] S1016: Input the normalized standard face image into a pre-trained face feature extraction model, and extract face features from the standard face image through the face feature extraction model.
[0150] S1018: Calculate the similarity between different facial images based on the facial features corresponding to each facial image.
[0151] S1020: Clustering facial images that meet facial similarity conditions into the same facial category based on similarities between different facial images.
[0152] S1022: For each face category, select a representative face image that meets a large face condition from the face images included in the face category.
[0153] S1024: Input representative face images corresponding to each face category into the user age discrimination model to obtain user age labels corresponding to each face category.
[0154] S1026: Input representative face images corresponding to each face category into the user gender discrimination model to obtain user gender labels corresponding to each face category.
[0155] S1028, determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed.
[0156] S1030: Input the user age labels, user gender labels, and corresponding face co-occurrence features corresponding to different face categories into a pre-trained decision tree model to obtain user relationship information between different face categories.
[0157] S1032 : For each face image included in each face category, calculate the area ratio between each face image and the corresponding image to be processed.
[0158] S1034: Calculate the number of facial images in each facial category whose area ratio is greater than or equal to a preset ratio.
[0159] S1036: Use the face category corresponding to the maximum number as the central face category corresponding to the user relationship graph to be constructed.
[0160] S1038: Determine user relationship information between each face category and the central face category based on user relationship information between different face categories.
[0161] S1040: Construct a user relationship graph of the central user corresponding to the central face category based on the user relationship information between each face category and the central face category.
[0162] S1042, obtaining a user relationship graph display instruction.
[0163] S1044: Display a user relationship graph including facial images and relationships between facial images according to the user relationship graph display instruction.
[0164] The above-mentioned image processing method clusters the facial images in the image to be processed, and clusters the facial images that meet the facial similarity conditions into the same facial category, thereby accurately identifying the facial images corresponding to different users. Then, based on the facial images corresponding to different users, the user attribute labels corresponding to each user are determined. In addition, the corresponding facial co-occurrence features are determined based on the number of times the faces of different users appear in the same image to be processed, that is, the number of photos of different users together. The pre-trained user relationship decision model can effectively integrate the user attribute labels corresponding to different users and the corresponding facial co-occurrence features to obtain a user relationship graph. In this way, the image to be processed can be directly processed to mine the self-attribute information of different users in the image to be processed, as well as the degree of intimacy between different users, so as to accurately and efficiently predict the user relationship graph, without having to search for user relationship information from various network sources, greatly improving the efficiency of mining the user relationship graph.
[0165] Figure 10 FIG. 1 is a flow chart of an image processing method in one embodiment. It should be understood that although Figure 10 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 10 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0166] In a specific embodiment, referring to Figure 11 , Figure 11 FIG. 1 is a flow chart of an image processing method in a specific embodiment. Figure 11 As shown, the computer device can obtain user images from the user's album, perform face detection, face recognition and face clustering on the user images, and thus determine the facial images corresponding to different users. Then, the computer device can perform feature analysis of different dimensions on the facial images, such as determining the age label, gender label, action label, number of group photos, etc. corresponding to the facial images. These feature data of different dimensions are then input into multiple tree models to predict user relationships. For example, Figure 11The user relationship map in the lower middle part shows the user relationships between different users. The user in the middle of the user relationship map is the owner of the album, the user in the upper left corner is the user's lover, the user in the upper right corner is the user's mother, and the users in the lower left and lower right corners are the user's children.
[0167] like Figure 12 As shown, in one embodiment, an image processing device 1200 is provided, including an acquisition module 1201, a face clustering module 1202, a determination module 1203 and a user relationship prediction module 1204.
[0168] The acquisition module 1201 is used to acquire the image to be processed and the face image included in the image to be processed.
[0169] The face clustering module 1202 is used to perform clustering processing on face images, and cluster face images that meet face similarity conditions into the same face category.
[0170] The determination module 1203 is configured to determine the user attribute label corresponding to each face category based on the face images included in each face category.
[0171] The determination module 1203 is further configured to determine the co-occurrence features of faces between different face categories based on the number of times that face images belonging to different face categories appear in the same image to be processed.
[0172] The user relationship prediction module 1204 is used to input the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into the pre-trained user relationship decision model to obtain a user relationship graph.
[0173] In one embodiment, the acquisition module 1201 includes a screening unit 12011 and a face detection unit 12012, wherein the screening unit 12011 is used to obtain a user image set; filter out images including faces from the user image set as images to be processed; and the face detection unit 12012 is used to perform face detection on the images to be processed and extract face images from the images to be processed.
[0174] In one embodiment, the face detection unit 12012 is further used to input the image to be processed into the face detection model, process the image to be processed through the first convolutional neural network in the face detection model, and obtain a candidate face detection frame for marking the face; through the second convolutional neural network in the face detection model, screen out a spare face image from the candidate face image determined by the candidate face detection frame; the probability that the spare face image belongs to a valid face image satisfies a first probability condition; through the third convolutional neural network in the face detection model, screen out a face image to be clustered from the spare face images; the probability that the screened face image belongs to a valid face image satisfies a second probability condition.
[0175] In one embodiment, the face clustering module 1202 includes an extraction unit 12021, a calculation unit 12022, and a clustering unit 12023, wherein:
[0176] The extraction unit 12021 is used to extract facial features from the facial image using a pre-trained facial feature extraction model.
[0177] The calculation unit 12022 is used to calculate the similarity between different facial images based on the facial features corresponding to each facial image.
[0178] The clustering unit 12023 is used to cluster facial images that meet facial similarity conditions into the same facial category based on the similarity between different facial images.
[0179] In one embodiment, the extraction unit 12021 is also used to adjust the size of the facial image to a standard size to obtain a standard facial image; normalize the standard facial image; input the normalized standard facial image into a pre-trained facial feature extraction model, and extract facial features from the standard facial image through the facial feature extraction model.
[0180] In one embodiment, the user attribute label includes a user age label and a user gender label; the determination module 1203 is also used to, for each face category, screen out representative face images that meet the large face condition from the face images included in the face category; the representative face images corresponding to each face category are respectively input into the user age discrimination model to obtain the user age labels corresponding to each face category; the representative face images corresponding to each face category are respectively input into the user gender discrimination model to obtain the user gender labels corresponding to each face category.
[0181] In one embodiment, the determination module 1203 is further configured to pair up face categories to obtain at least one face category combination; for each face category combination, based on the number of times facial images belonging to two face categories in the face category combination appear in the same image to be processed, determine the facial co-occurrence features corresponding to the face category combination. The user relationship prediction module 1204 is further configured to, for each face category combination, input the user attribute labels corresponding to the face categories included in the face category combination and the corresponding facial co-occurrence features into a pre-trained decision tree model to obtain user relationship information corresponding to the face category combination; and based on the user relationship information corresponding to each face category combination, construct a user relationship graph corresponding to the image to be processed.
[0182] In one embodiment, the user relationship prediction module 1204 is also used to input the user attribute labels corresponding to different face categories and the corresponding face co-occurrence features into a pre-trained decision tree model to obtain user relationship information between different face categories; determine the central face category in the face category; determine the user relationship information between each face category and the central face category based on the user relationship information between different face categories; and construct a user relationship map of the central user corresponding to the central face category based on the user relationship information between each face category and the central face category.
[0183] In one embodiment, the user relationship prediction module 1204 is also used to calculate the area ratio between each facial image and the corresponding image to be processed for each facial image included in each facial category; calculate the number of facial images included in each facial category whose area ratio is greater than or equal to a preset ratio; and use the facial category corresponding to the maximum number as the central facial category corresponding to the user relationship graph to be constructed.
[0184] refer to Figure 13 In one embodiment, the image processing device 1200 further includes a display module 1205, wherein: the acquisition module 1201 is further used to obtain a user relationship graph display instruction; the display module 1205 is used to display a user relationship graph including facial images and the relationships between facial images according to the user relationship graph display instruction.
[0185] The above-mentioned image processing device clusters the facial images in the image to be processed, and clusters the facial images that meet the facial similarity conditions into the same facial category, thereby accurately identifying the facial images corresponding to different users. Then, based on the facial images corresponding to different users, the user attribute labels corresponding to different users are determined. In addition, based on the number of times the faces of different users appear in the same image to be processed, that is, the number of photos of different users, the corresponding facial co-occurrence features are determined. The pre-trained user relationship decision model can effectively integrate the user attribute labels corresponding to different users and the corresponding facial co-occurrence features to obtain a user relationship graph. In this way, the image to be processed can be directly processed to mine the self-attribute information of different users in the image to be processed, as well as the degree of intimacy between different users, so as to accurately and efficiently predict the user relationship graph, without having to search for user relationship information from various network sources, thereby greatly improving the efficiency of mining the user relationship graph.
[0186] Figure 14 The internal structure diagram of a computer device in one embodiment is shown. The computer device can be Figure 1 The terminal 110 or the server 120 in FIG. Figure 14As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program that, when executed by the processor, causes the processor to implement the image processing method. The internal memory may also store a computer program that, when executed by the processor, causes the processor to perform the image processing method.
[0187] Those skilled in the art will understand that Figure 14 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0188] In one embodiment, the image processing apparatus provided by the present application can be implemented in the form of a computer program. The computer program can be used in Figure 14 The computer device can be operated on the computer device shown. The memory of the computer device can store various program modules that constitute the image processing device, such as, Figure 12 The acquisition module, face clustering module, determination module and user relationship prediction module shown in the figure are computer programs composed of various program modules, which enable the processor to execute the steps of the image processing method of each embodiment of the present application described in this specification.
[0189] For example, Figure 14 The computer device shown can be Figure 12 The acquisition module in the image processing apparatus shown executes step S202. The computer device may execute step S204 through the face clustering module. The computer device may execute steps S206 and S208 through the determination module. The computer device may execute step S210 through the user relationship prediction module.
[0190] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above-mentioned image processing method. The steps of the image processing method may be the steps of the image processing method in each of the above-mentioned embodiments.
[0191] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the above-mentioned image processing method. The steps of the image processing method may be the steps in the image processing methods of the above-mentioned embodiments.
[0192] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0193] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0194] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a user image set, select images including human faces from the user image set as images to be processed, perform face detection on the images to be processed, and extract human face images from the images to be processed; By using a pre-trained facial feature extraction model, facial features in the facial image are extracted, and similarities between different facial images are calculated based on the facial features of each facial image. When the similarity between two facial images is greater than a similarity threshold, the two facial images are associated. In a first round of iteration, for the same facial image, when there are multiple candidate facial images with similarities greater than a similarity threshold corresponding to it in other facial images, the candidate facial image corresponding to the maximum similarity is clustered with the facial image into one category of facial images, and the association between the facial image and the facial image corresponding to the maximum similarity is saved, and the association between the facial image and the other facial images is canceled. After the first round of iteration, when the facial image has corresponding relationships with multiple facial categories, the facial image is clustered into the facial category that includes the largest number of facial images. The steps after the first round of iteration are continuously repeated until convergence occurs, and at least one facial category is obtained. For each face category, select representative face images that meet the large face condition from the face images included in the face category, input the representative face images corresponding to each face category into the user age discrimination model to obtain the user age label corresponding to each face category, and input the representative face images corresponding to each face category into the user gender discrimination model to obtain the user gender label corresponding to each face category; Determining face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed; Inputting the user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features into a pre-trained decision tree model to obtain user relationship information between different face categories; wherein the user attribute labels include the user age label and the user gender label; For each face image included in each face category, calculate the area ratio between each face image and the corresponding image to be processed; calculate the number of face images included in each face category whose area ratio is greater than or equal to a preset ratio; use the face category corresponding to the largest number as the central face category, and determine the user relationship information between each face category and the central face category based on the user relationship information between different face categories; Based on the user relationship information between each face category and the central face category, a user relationship graph of the central user corresponding to the central face category is constructed; A user relationship graph display instruction is obtained, and according to the user relationship graph display instruction, a user relationship graph including the facial images and the relationships between the facial images is displayed.
2. The method according to claim 1, characterized in that The step of performing face detection on the image to be processed and extracting a face image from the image to be processed includes: Inputting the image to be processed into a face detection model, processing the image to be processed by a first convolutional neural network in the face detection model to obtain a candidate face detection frame for marking a face; screening a backup facial image from the candidate facial images determined by the candidate facial detection frame using a second convolutional neural network in the facial detection model; wherein a probability that the backup facial image is a valid facial image satisfies a first probability condition; The face images to be clustered are screened out from the spare face images through the third convolutional neural network in the face detection model; the probability that the screened face images are valid face images satisfies a second probability condition.
3. The method according to claim 1, characterized in that The extracting of facial features from the facial image using a pre-trained facial feature extraction model includes: Adjusting the size of the facial image to a standard size to obtain a standard facial image; performing normalization processing on the standard face image; The normalized standard face image is input into a pre-trained face feature extraction model, and the face feature extraction model is used to extract face features from the standard face image.
4. The method according to claim 1, wherein The determining of the face co-occurrence features between the different face categories based on the number of times the face images belonging to the different face categories appear in the same image to be processed includes: Combining the face categories in pairs to obtain at least one group of face category combinations; For each face category combination, determining a face co-occurrence feature corresponding to the face category combination based on the number of times face images belonging to two face categories in the face category combination appear in the same image to be processed; The user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features are input into the pre-trained decision tree model to obtain user relationship information between different face categories, including: For each face category combination, the user attribute labels corresponding to the face categories included in the face category combination and the corresponding face co-occurrence features are input into the pre-trained decision tree model to obtain user relationship information corresponding to the face category combination.
5. An image processing device, characterized in that: The device comprises: The acquisition module includes a screening unit and a face detection unit, wherein the screening unit is used to obtain a user image set; screen images including faces from the user image set as images to be processed; and the face detection unit is used to perform face detection on the images to be processed and extract face images from the images to be processed; A face clustering module is configured to extract facial features from the face image using a pre-trained facial feature extraction model, calculate similarities between different face images based on the facial features of each face image, associate the two face images when the similarity between the two face images is greater than a similarity threshold, and in a first round of iteration, for a same face image, if there are multiple candidate face images with similarities greater than a similarity threshold corresponding to the same face image in other face images, cluster the candidate face image corresponding to the maximum similarity with the face image into one category of face images, preserve the association between the face image and the face image corresponding to the maximum similarity, and cancel the association between the face image and the other face images. After the first round of iteration, if a face image has corresponding relationships with multiple face categories, cluster the face image into the face category that includes the largest number of face images, and continuously repeat the steps after the first round of iteration until convergence to obtain at least one face category. a determination module configured to, for each face category, screen out representative face images that meet the large face condition from the face images included in the face category; input the representative face images corresponding to each face category into a user age discrimination model to obtain user age labels corresponding to each face category; and input the representative face images corresponding to each face category into a user gender discrimination model to obtain user gender labels corresponding to each face category; The determination module is further configured to determine the face co-occurrence features between different face categories based on the number of times face images belonging to different face categories appear in the same image to be processed; A user relationship prediction module is configured to input the user attribute labels corresponding to the different face categories and the corresponding face co-occurrence features into a pre-trained decision tree model to obtain user relationship information between different face categories; wherein the user attribute labels include the user age label and the user gender label; for each face image included in each face category, calculate the area ratio between each face image and the corresponding image to be processed; calculate the number of face images in each face category whose area ratio is greater than or equal to a preset ratio; take the face category corresponding to the largest number as the central face category, and determine the user relationship information between each face category and the central face category based on the user relationship information between different face categories; and construct a user relationship map of the central user corresponding to the central face category based on the user relationship information between each face category and the central face category; The acquisition module is further used to obtain a user relationship graph display instruction; A display module is used to display a user relationship map including the facial images and the relationships between the facial images according to the user relationship map display instruction.
6. The device according to claim 5, characterized in that The face detection unit is further configured to: input the image to be processed into a face detection model, process the image to be processed by a first convolutional neural network in the face detection model, and obtain a candidate face detection frame for marking a face; A second convolutional neural network in the face detection model is used to screen out a backup face image from the candidate face images determined by the candidate face detection frame; the probability that the backup face image is a valid face image satisfies a first probability condition; a third convolutional neural network in the face detection model is used to screen out face images to be clustered from the backup face images; the probability that the screened face image is a valid face image satisfies a second probability condition.
7. The device according to claim 5, characterized in that The face clustering module is also used to adjust the size of the face image to a standard size to obtain a standard face image; normalize the standard face image; input the normalized standard face image into a pre-trained face feature extraction model, and extract facial features from the standard face image through the face feature extraction model.
8. The device according to claim 6, characterized in that The determination module is further configured to perform pairwise combinations of the face categories to obtain at least one face category combination; for each face category combination, determine a face co-occurrence feature corresponding to the face category combination based on the number of times face images belonging to two face categories in the face category combination appear in the same image to be processed; The user relationship prediction module is also used to input the user attribute labels corresponding to the face categories included in each face category combination and the corresponding face co-occurrence features into a pre-trained decision tree model to obtain user relationship information corresponding to the face category combination.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 4.
10. A computer device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and apparatus for presenting character relation
CN102043816A
Human face clustering method and apparatus for pictures, and storage medium
CN107909104A
Image classifying method and device, electronic equipment and computer readable storage medium
CN108021669A
Method for constructing family relationship of person for electronic album management
CN108960043A