A face classification method, device and electronic device

CN114648799BActive Publication Date: 2025-08-01HANGZHOU HIKSTORAGE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210346309.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-08-01
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

[0004]但是,受限制于拍摄条件,如拍摄到的人脸图像中的人脸为侧脸、人脸被障碍物遮挡等,人脸图像中包含的人脸信息的信息量(下文称人脸信息量)较少,导致无法准确地从人脸图像中提取人脸特征,从而无法实现对这些人脸信息量较低的人脸图像进行分类,即无法适用于人脸信息量较低的人脸图像,无法全面地对人脸图像进行分类

Benefits of technology

[0019] The face classification method, device, and electronic device provided in the embodiments of the present invention can further determine whether a target image in a first image set is classified into a second image set based on face similarity and/or similarity of description information, so as to more comprehensively classify images, thereby solving the problem that in the existing image classification process, it is impossible to perform a one-time classification on face images with a relatively low amount of face information contained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648799B_ABST
    Figure CN114648799B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a face classification method, apparatus, and electronic device. The method includes: determining a first image set, where the description information between different images in the first image set is similar, where the similarity of the description information includes the same or similar description information; determining a second image set having the same images as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold; determining whether a target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, where the target image includes any image that is different from the images in the second image set. This solution can classify images more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly to a face classification method, apparatus, and electronic device. Background Art

[0002] In some application scenarios, in order to facilitate the management and processing of face images, it is necessary to group face images according to the faces in the images, so as to divide the images containing the faces of the same person into the same group. In this article, this process is called face classification.

[0003] In the related art, face features in a face image are extracted, and by comparing the face features, face images with matching face features are determined, and these face images with matching face features are divided into the same image set.

[0004] However, limited by shooting conditions, such as the face in the captured face image being a side face, the face being blocked by an obstacle, etc., the amount of face information (hereinafter referred to as face information amount) contained in the face image is small, resulting in the inability to accurately extract face features from the face image, and thus the inability to classify these face images with a low face information amount, that is, it is not applicable to face images with a low face information amount and cannot comprehensively classify face images. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a face classification method, apparatus, and electronic device to improve the applicability of the face classification method, so as to achieve more comprehensive classification of face images. The specific technical solutions are as follows:

[0006] In the first aspect of the embodiments of the present invention, a face classification method is provided. The method includes:

[0007] Determine a first image set, where the description information between different images in the first image set is similar, where the similarity of the description information includes the same or similar description information;

[0008] Determine a second image set that has the same image as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold;

[0009] Based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, determine whether the target image in the first image set is classified into the second image set, where the target image includes any image that is different from the images in the second image set.

[0010] In a second aspect of the embodiments of the present invention, a face classification device is provided, and the device includes:

[0011] A first determination module, configured to determine a first image set, where the description information between different images in the first image set is similar, and where the similarity of the description information includes the same or similar description information;

[0012] A second determination module, configured to determine a second image set that has the same images as the first image set, and the face similarity between any two images in the second image set is greater than a first threshold;

[0013] A classification module, configured to determine whether a target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, where the target image includes any image that is different from the images in the second image set.

[0014] In a third aspect of the embodiments of the present invention, an electronic device is provided, including:

[0015] A memory, configured to store a computer program;

[0016] A processor, configured to implement the method steps described in any one of the first aspects when executing the program stored in the memory.

[0017] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, where a computer program is stored in the computer-readable storage medium, and the computer program implements the method steps described in any one of the first aspects when executed by a processor.

[0018] Advantages of the embodiments of the present invention:

[0019] The face classification method, device, and electronic device provided in the embodiments of the present invention can further determine whether a target image in a first image set is classified into a second image set based on face similarity and / or similarity of description information, so as to more comprehensively classify images, thereby solving the problem that in the existing image classification process, it is impossible to perform a one-time classification on face images with a relatively low amount of face information contained.

[0020] Of course, when implementing any product or method of the present invention, it is not necessarily required to simultaneously achieve all the above-mentioned advantages. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.

[0022] Figure 1 A flowchart of a face classification method provided by an embodiment of the present invention;

[0023] Figure 2 A flowchart of a method for partitioning a set of scene images provided by an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of image set conversion of a face classification method provided by an embodiment of the present invention;

[0025] Figure 4 A schematic structural diagram of a face classification device provided by an embodiment of the present invention;

[0026] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.

[0028] To more clearly illustrate the face classification method provided by the embodiments of the present invention, the following will exemplarily illustrate the application scenarios of the face classification method provided by the embodiments of the present invention. It can be understood that the following examples are only one possible application scenario of the face classification method provided by the embodiments of the present invention. In other possible embodiments, the face classification method provided by the embodiments of the present invention can also be applied to other possible application scenarios. The following examples do not impose any restrictions on this.

[0029] To manage the people in public places such as shopping malls, scenic spots, parks, etc., the staff pre-arrange cameras in public places. The cameras send the captured images to the server, and the server performs face detection on the received images to determine the face images that contain faces. The server performs face modeling based on the face information contained in the face images to obtain a face model, and extracts the features of the face model as the face features of the face images. The server divides the face images into multiple face image sets according to the extracted face features, where the face features of any two face images in each face image set match each other. The staff can know the people appearing in public places based on the divided face image sets, thereby effectively managing the people in public places.

[0030] For another example, obtain the captured image of the user terminal, classify and manage the image, generate a recommended album for the user, or through image classification, facilitate recommending more search results to the user during the user's image search process.

[0031] However, restricted by various conditions, for example, the captured face image is a side view, top view, or bottom view of the face, or for another example, the face in the captured face image is blocked by an obstacle, resulting in less face information in the captured face image. When the face information in the face image is less, it is difficult or even impossible to establish a face model based on the face image, so the face features cannot be accurately extracted, and subsequent face classification of the face image cannot be performed based on the face features.

[0032] Therefore, the existing solutions can only perform face classification on face images with more face information, that is, only classify some face images, with poor applicability and unable to comprehensively classify face images.

[0033] Based on this, the embodiments of the present invention provide a face classification method, which is applied to any electronic device with face classification capabilities, including but not limited to servers, user terminals (such as personal computers, mobile phones, tablets), etc.

[0034] See Figure 1 , Figure 1 As shown in the flowchart of the face classification method provided by the embodiments of the present invention, it includes:

[0035] S101, determine the first image set.

[0036] The description information between different images in the first image set is similar, where the similarity of the description information includes the same or similar description information. The description information of the image is used to describe the attributes of the shooting environment for obtaining the image. In this article, the attributes of the shooting environment may vary according to different application scenarios, such as the time when the image is taken (hereinafter referred to as the shooting time), the location where the image is taken (hereinafter referred to as the shooting location), the scene where the image is taken, the camera number of the image, etc. Specifically, the same description information may include, but is not limited to: the same shooting time, the same shooting location, the same shooting scene, the same camera number, and the similar description information may include, but is not limited to: the shooting times are relatively close (for example, the time difference between two shooting times is less than the preset time difference threshold), the shooting locations are relatively close (for example, the distance between two shooting locations is less than the preset distance threshold), the shooting scenes are relatively similar (for example, the similarity of the scene picture features on two images is greater than the preset scene similarity threshold, such as both are hospitals, schools, a certain scenic spot, a beach, etc.), the camera numbers belong to the same series of camera products and thus the differences in camera function parameters are relatively small, etc.

[0037] S102, determine a second image set that has the same images as the first image set.

[0038] The face similarity between any two images in the second image set is greater than a first threshold.

[0039] S103, based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, determine whether the target image in the first image set belongs to the second image set.

[0040] The target image includes any image that is different from the images in the second image set.

[0041] Since the description information between different images in the first image set is similar, it is considered that the shooting environments for each image in the first image set are the same or similar. And since the face similarity between any two images in the second image set is greater than the first threshold, it is considered that the same person is included in each image in the second image set.

[0042] For convenience of description hereinafter, it is assumed that the shooting environments of the images in the first image set are the same as or similar to the target shooting environment, and each image in the second image set includes the target person. Since there are identical images in the first image set and the second image set, and the images that belong to both the first image set and the second image set can be considered as the images of the target person taken in the target shooting environment (or a similar shooting environment), it is considered that the target person was once photographed in the target shooting environment (or a similar shooting environment). By reverse inference, the images taken in the target shooting environment (or a similar shooting environment) may be the images of the target person. Therefore, the target images in the first image set may include the target person. Therefore, it is possible to further determine whether the target images in the first image set are classified into the second image set based on the similarity of facial features and / or the similarity of descriptive information, so as to achieve a more comprehensive classification of images.

[0043] Exemplarily, it is assumed that person A was present at the target location at the target time and two different images were taken, denoted as the first image and the second image respectively. In the first image, the face of person A is not blocked, and in the second image, the face of person A is partially blocked.

[0044] And it is assumed that the descriptive information includes the shooting time and the shooting location. Since the first image and the second image were taken at the target time and the target location, the first image and the second image are classified into the same first image set.

[0045] Moreover, since the face of person A in the first image is not blocked, the similarity between the first image and other images including the complete face of person A is greater than the first threshold. Therefore, the first image and these images including the complete face of person A are classified into the same face image set (hereinafter denoted as the person A image set).

[0046] However, the face of person A in the second image is partially blocked. Therefore, the facial similarity between the second image and the images in the person A image set is not greater than the first threshold. Therefore, the second image does not belong to the person A image set.

[0047] Since the first image set and the person A image set include the first image, the person A image set can be regarded as the second image set. And the second image is different from the images in the person A image set. Therefore, the second image is the target image in the first image set.

[0048] Although the face of person A in the second image is partially blocked, the second image still includes a partial face of the person. Therefore, compared with the images that do not include person A, the facial similarity between the second image and the images in the person A image set is relatively high. Therefore, the second image is classified into the person A image set.

[0049] It can be seen that by using the face classification method provided in the embodiments of the present invention, the second image that was not originally classified into the image set of person A can be classified into the image set of person A, so that the image set of person A includes more images of person A, that is, more comprehensive image classification is achieved.

[0050] The following will separately describe the foregoing S101 - S103:

[0051] In S101, "similar" in this article may refer to the similarity between the two being greater than a preset similarity threshold. For example, similar description information means that the similarity between the description information is greater than the preset similarity threshold. The similarity threshold can be set according to actual needs and / or experience.

[0052] For the convenience of description hereinafter, only the example where the description information includes the shooting time, shooting location, and the scene of the captured image will be used for illustration. Since the scene of the captured image will be reflected in the captured image. For example, assuming the scene of a captured image is a beach, then the beach scene will be shown on this image. Therefore, the scene of the captured image is the scene information shown on the image (hereinafter referred to as scene information).

[0053] In other possible embodiments, the description information may also only include some of the shooting time, shooting location, and scene information, and the description information may also include other information other than the shooting time, shooting location, and scene information, such as the camera number of the captured image, the weather at the time of capturing the image, etc. This embodiment does not impose any restrictions on this.

[0054] In S102, the first threshold can be set according to actual needs and / or experience. However, if two images include the complete face of the same person and the clarity of the two images is greater than the preset clarity threshold, then the face similarity between the two images should be greater than the first threshold. And if two images include the face of the same person, but at least one of the images includes an incomplete face, then the face similarity between the two images should not be greater than the first threshold.

[0055] The magnitude of the face similarity in this article, as well as the magnitude of the similarity of the shooting time, the magnitude of the similarity of the shooting location, and the magnitude of the similarity of the scene information hereinafter, all refer to the degree of similarity represented, rather than the magnitude of the numerical value. Depending on the different forms of similarity representation, the degree of similarity represented by the similarity may be positively correlated or negatively correlated with the numerical value of the similarity.

[0056] Exemplarily, taking face similarity as an example, if face similarity is represented by the Euclidean distance between face features, the smaller the value of the face similarity between two images, the higher the similarity degree of the faces in the two images. If face similarity is represented by the cosine similarity of feature vectors, the smaller the value of the face similarity between two images, the lower the similarity degree of the faces in the two images.

[0057] In S103, it can be understood that there may be some images including the target person in the first image set, but these images do not belong to the second image set for various reasons. Therefore, these images need to be classified from the first image set to the second image set. Exemplarily, the target images in the first image set include: images with the face in an occluded state, images with the face in a profile state, images obtained by shooting the face from a top view or a bottom view.

[0058] There are the following three implementation methods for classification:

[0059] Method 1: Determine whether the target images in the first image set are classified into the second image set based on the face similarity between the target images in the first image set and the images in the second image set.

[0060] Method 2: Determine whether the target images in the first image set are classified into the second image set based on the similarity of the description information between the target images in the first image set and the images in the second image set.

[0061] Method 3: Determine whether the target images in the first image set are classified into the second image set based on the face similarity between the target images in the first image set and the images in the second image set, and the similarity of the description information between the target images in the first image set and the images in the second image set.

[0062] In a possible embodiment, the aforementioned S103 includes the following steps:

[0063] S1031, Calculate the classification value of the target image based on the face similarity between the target image in the first image set and the image in the second image set, and / or the similarity of the description information between the target image in the first image set and the image in the second image set.

[0064] There are the following three implementation methods for this step:

[0065] Method 4: Calculate the classification value of the target image based on the face similarity between the target image in the first image set and the image in the second image set.

[0066] Method 5: Calculate the classification value of the target image based on the similarity of the description information between the target image in the first image set and the images in the second image set.

[0067] Method 6: Calculate the classification value of the target image based on the face similarity between the target image in the first image set and the images in the second image set, and the similarity of the description information between the target image in the first image set and the images in the second image set.

[0068] For Method 4, the classification value is positively correlated with the face similarity, that is, the greater the face similarity, the greater the calculated classification value. For Method 5, the classification value is positively correlated with the similarity of the description information. For Method 6, the classification value is positively correlated with the face similarity and is also positively correlated with the similarity of the description information.

[0069] Exemplarily, in a possible embodiment, the classification value can be calculated according to the following formula:

[0070] Classification value = α * face similarity + (1 - α) * similarity of description information

[0071] Where α is a preset weight, and the value range of α is (0, 1).

[0072] S1032, Classify the target images with a classification value greater than the second threshold into the second image set.

[0073] For Method 4 and Method 6, if the face similarity is greater than the first threshold, the classification value calculated based on the face similarity should be greater than the second threshold.

[0074] To more clearly illustrate the face classification method provided by the present invention, taking the description information as at least including the shooting time, shooting location, and scene information as an example, the foregoing S101, that is, how to determine the first image set, will be exemplarily described below.

[0075] In this example, the foregoing S101 is as Figure 2 shown and includes:

[0076] S1011, Determine the third image set.

[0077] The shooting time and / or shooting location between any two images in the third image set are similar.

[0078] The shooting times of two images being similar means that the shooting times of the two images are the same or close, and the shooting times of two images being close means that the interval between the shooting times of the two images is less than the preset duration threshold.

[0079] The shooting locations of two images being similar means that the shooting locations of the two images are the same or close. The shooting locations of two images being close can mean that the distance between the shooting locations of the two images is less than a preset distance threshold, or it can mean that the shooting locations of the two images are in the same area. Exemplarily, in a possible embodiment, assume that the shooting locations of two images are both in the same scenic spot and the distance between them is 1000 meters. Then, in a possible embodiment, if the preset distance threshold is 500 meters, it is considered that the shooting locations of the two images are not close. However, in another possible embodiment, since the shooting locations of the two images are both in the same scenic spot, it is considered that the shooting locations of the two images are close.

[0080] S1012, obtaining a first image set based on the images in the third image set whose scene information similarity is greater than a third threshold.

[0081] That is, the similarity of the scene information of any two images in the first image set is greater than the third threshold. Exemplarily, assume that the third image set includes 4 images, denoted as Image 1 - 4 respectively, and the similarity of the scene information of each image is shown in Table 1:

[0082] Table 1

[0083] Image 1 Image 2 Image 3 Image 4 Image 1 100% 95% 95% 75% Image 2 95% 100% 95% 75% Image 3 90% 85% 100% 75% Image 4 75% 75% 75% 100%

[0084] The 100% in the second row and second column of Table 1 is the similarity of the scene information of Image 1 and the scene information of Image 1, and the 95% in the second row and third column is the similarity of the scene information of Image 1 and the scene information of Image 2, and so on.

[0085] If the third threshold is 90%, the determined third image set is {Image 1, Image 2, Image 3}

[0086] Selecting this embodiment, first determining the third image set based on the shooting time and / or shooting location, and then determining the first image set in the third image set based on the similarity of the scene information, effectively reducing the number of similarities of the scene information that need to be calculated. And compared with the similarity of the shooting time and / or shooting location, the calculation of the similarity of the scene information is more complex. Therefore, selecting this embodiment can effectively improve the efficiency of determining the first image set.

[0087] The method for determining the third image set in the foregoing S1011 can be different according to different application scenarios. Exemplarily, the K - means or a clustering processing method improved based on K - means can be used to cluster all the images in the image set to be classified according to the similarity of the shooting time and / or shooting location and the preset initial clustering center, obtaining multiple clustering clusters, and taking each clustering cluster as a third image set.

[0088] However, the accuracy of the K-means or the clustering processing method improved based on K-means is affected by the initial clustering center, and it is difficult for users to set a suitable initial clustering center according to experience. Based on this, in a possible embodiment, the foregoing S1011 includes:

[0089] S1011a, based on the similarity of shooting time and / or shooting location between the images to be classified, perform a first clustering process on the images to be classified to determine the initial clustering center.

[0090] In a possible embodiment, all the images to be classified include people. In another possible embodiment, some of the images to be classified include people, and the other part of the images to be classified are landscape images without people.

[0091] The first clustering process may vary according to different application scenarios, but it should be a clustering processing method that does not require setting the initial clustering center, such as the canopy clustering processing method. This will be exemplarily described below and will not be elaborated here.

[0092] S1011b, based on the similarity of shooting time and / or shooting location between the images to be classified, according to the initial clustering center, perform a second clustering process on the images to be classified to obtain a third image set.

[0093] The second clustering process may vary according to different application scenarios, but it should be a clustering processing method that requires setting the initial clustering center, such as K-means or the clustering processing method improved based on K-means. This will be exemplarily described below and will not be elaborated here.

[0094] Selecting this embodiment, the first clustering process can be used for preliminary clustering, providing prior knowledge for the initial clustering center required by the second clustering process, thereby improving the clustering effect of the second clustering process, that is, being able to more accurately determine the third image set from the images to be classified.

[0095] The processes of the foregoing first clustering process and the foregoing second clustering process will be exemplarily described below respectively.

[0096] For the first clustering process, the process is as follows:

[0097] Based on the similarity of shooting time and / or shooting location between the images to be classified, divide the images to be classified into multiple first clustering image sets. For each first clustering image set, select an image in the first clustering image set as the initial clustering center.

[0098] The clustering center in this article is an image that can represent a class of images, that is, the shooting time and / or shooting location between the clustering center and each image represented by the clustering center should be the same or close enough.

[0099] The way to divide the first clustered image set can be different according to different application scenarios, but it should satisfy that the similarity of shooting time and / or shooting location between any two images in the first clustered image set is greater than the fourth threshold. The fourth threshold here is used to refer to the similarity thresholds corresponding to the shooting time and shooting location respectively. Specifically, the fourth threshold for the shooting time similarity and the fourth threshold for the shooting location similarity can be the same value or different values, which can be determined according to actual needs and are not limited in the embodiments of this application.

[0100] Exemplarily, in a possible embodiment, the division of the first clustered image set is achieved through the following steps 1 to 4:

[0101] Step 1: Randomly arrange the images to be classified to obtain an image sequence.

[0102] Exemplarily, assume that there are a total of m images to be classified, denoted as Image 1 - Image m respectively, and after random arrangement, an image sequence [Image 1, Image 2, Image 3,..., Image m] is obtained.

[0103] Step 2: Randomly select an image to be classified from the image sequence to be classified as a new first clustered image set, and delete the selected image to be classified from the image sequence.

[0104] Exemplarily, assume that the randomly selected image to be classified is Image j, then Image j is used as the centroid, and Image j is deleted from the image sequence. After deleting Image j, the image sequence is [Image 1, Image 2,... Image j - 1, Image j + 1, Image j + 2,..., Image m], and at this time, there is a first clustered image set {Image j}.

[0105] Step 3: Randomly select an image to be classified from the latest image sequence, and calculate the distance between the selected image to be classified and each first clustered image set.

[0106] Among them, the distance between the image to be classified and the first clustered image set can be calculated through different distance formulas, but it should satisfy being negatively correlated with the similarity of shooting time and / or shooting location between the image to be classified and each image in the first clustered image set. Exemplarily, assume that the randomly selected image to be classified is Image i, then the higher the similarity of shooting time and / or shooting location between Image i and Image j, the smaller the distance between Image i and the first clustered image set {Image j}.

[0107] Moreover, when calculating the distance between the image to be classified and the first cluster image set, it can be based on the similarity of the shooting time and / or shooting location between the image to be classified and each image in the first cluster image set, or it can be based on the similarity of the shooting time and / or shooting location between the image to be classified and some images in the first cluster image set.

[0108] Exemplarily, in a possible embodiment, the part of the images refers to the centroid image in the first cluster image set, and the centroid image is an image whose shooting time and / or shooting location is the same as or close to the mean value of the shooting time and / or shooting location of all images in the first cluster image set.

[0109] The mean value of the shooting time and / or shooting location in this article includes three cases: the first: the mean value of the shooting time, the second: the mean value of the shooting location, and the third: the mean values corresponding to the shooting time and the shooting location respectively. Taking Image 1 and Image 2 as an example, assuming that the shooting time of Image 1 is t1 and the shooting location is p1, and the shooting time of Image 2 is t2 and the shooting location is p2, then the mean value of the shooting time of Image 1 and Image 2 is (t1 + t2) / 2, and the mean value of the shooting location is (p1 + p2) / 2. Among them, t1 and t2 can be represented in the form of UNIX timestamps (i.e., the number of seconds elapsed since January 1, 1970), and p1 and p2 can be represented in the form of spatial coordinates such as longitude and latitude coordinates and Mercator coordinates.

[0110] Since the shooting time and / or shooting location of the centroid image is the same as or close to the mean value of the shooting time and / or shooting location of all images in the first cluster image set, the shooting time and / or shooting location of the centroid image can represent the shooting time and / or shooting location of all images in the first cluster image set. That is, the similarity of the shooting time and / or shooting location between the image to be classified and the centroid image in the first cluster image set can represent the similarity of the shooting time and / or shooting location between the image to be classified and each image in the first cluster image set. Therefore, the distance between the image to be classified and the first cluster image set can be calculated only based on the similarity of the shooting time and / or shooting location between the image to be classified set and the centroid image.

[0111] Moreover, whenever an image to be classified is used as a first cluster image set, the image to be classified will become the centroid image in the first cluster image set, and a new centroid image can also be selected from the subsequent images to be classified added to the first cluster image set. How to select the centroid image will be illustrated by examples below and will not be elaborated here.

[0112] Step 4: If the minimum value of the calculated distance is less than the preset upper threshold, divide the selected image to be classified into the set of clustered images with the smallest distance from the image to be classified; if the minimum value of the calculated distance is not less than the preset upper threshold, take the selected image to be classified as a new set of the first clustered images; delete the selected image to be classified from the image sequence, and return to execute Step 3 until the image sequence is empty.

[0113] The setting of the upper threshold can be different according to different application scenarios, but it should satisfy that if the similarity of the shooting time and / or shooting location between an image and all images in a set of the first clustered images is greater than the fourth threshold, the distance between the image and the set of the first clustered images should be less than the preset upper threshold.

[0114] Exemplarily, assume that there are three sets of clustered images in total, denoted as Clustered Image Set 1 - 3, the selected image to be classified is Image i, the distance between Image i and Clustered Image Set 1 is 0.3, the distance between Image i and Clustered Image Set 2 is 0.4, the distance between Image i and Clustered Image Set 3 is 0.5, and the preset upper threshold is 0.4.

[0115] Then the set of clustered images with the smallest distance from Image i is Clustered Image Set 1. And since the distance of 0.3 between Image i and Clustered Image Set 1 is less than the preset upper threshold of 0.4, Image i is added to Clustered Image Set 1.

[0116] For the example where some of the above images are centroid images in the set of clustered images, it further includes: if the minimum value of the calculated distance is less than the preset lower threshold, take the selected image to be classified as the centroid image of the set of clustered images to be added, where the lower threshold is less than the upper threshold.

[0117] The method of selecting an image from the set of the first clustered images as the initial clustering center can be different according to different application scenarios, but the similarity of the shooting time and / or shooting location between the selected image and other images in the set of clustered images should be as high as possible.

[0118] For example, in a possible embodiment, it can be to calculate the mean value of the shooting time and / or shooting location of all images in the set of the first clustered images, and take the image in the set of the first clustered images whose shooting time and / or shooting location is closest to the mean value (for example, the difference from the mean value is less than the preset deviation value) as the initial clustering center.

[0119] Selecting this embodiment can enable the selected initial clustering centers to better reflect the differences in shooting time and / or shooting location among the images in different first clustering image sets, that is, to make the selected initial clustering centers as close as possible (even the same) to the true clustering centers, thereby improving the rate of iterating the initial clustering centers to the true clustering centers in the subsequent second clustering process and effectively improving the efficiency of the face classification method provided by the present invention.

[0120] Exemplarily, assume that a first clustering image set includes three images, denoted as Image 1 - 3 respectively. Among them, the shooting time of Image 1 is t, and the shooting location is (x, y); the shooting time of Image 2 is t + 4, and the shooting location is (x + 4, y + 4); the shooting time of Image 3 is t + 11, and the shooting location is (x + 11, y + = 11). The average value of the shooting times of Image 1 - 3 is t + 5, and the average value of the shooting locations is (x + 5, y + 5). It can be seen that the image with the shooting time and shooting location closest to the average value is Image 2. Therefore, Image 2 is used as the initial clustering center of this first clustering image set.

[0121] For the second clustering process, the flow is as follows:

[0122] Step Five: Classify all images to be classified based on the similarity of shooting time and / or shooting location between the images to be classified and each clustering center, and obtain multiple second clustering image sets.

[0123] Each clustering center is an image to be classified, and initially, the clustering center is the initial clustering center, which is determined based on the first clustering process. For details, refer to the relevant description of the first clustering process above and will not be elaborated here.

[0124] Each second clustering image set includes and only includes one clustering center, and each image in each second clustering image set satisfies that the similarity of shooting time and / or shooting location between this image and the clustering center in this second clustering image set is greater than the similarity of shooting time and / or shooting location between this image and other clustering centers.

[0125] Exemplarily, assume that there are a total of 5 clustering centers, denoted as clustering centers 1-5 respectively. If the similarity of the shooting time and / or shooting location between an image and clustering center 1 is 0.5, the similarity of the shooting time and / or shooting location between the image and clustering center 2 is 0.6, the similarity of the shooting time and / or shooting location between the image and clustering center 3 is 0.7, the similarity of the shooting time and / or shooting location between the image and clustering center 4 is 0.8, and the similarity of the shooting time and / or shooting location between the image and clustering center 5 is 0.9, then the image and clustering center 5 are classified into the same second clustering image set.

[0126] Step Six: For each second clustering image set, determine the mean value of the shooting time and / or shooting location of all images in the second clustering image set. Take the image in the second clustering image set whose shooting time and / or shooting location is closest to the corresponding mean value as the new clustering center, and return to execute Step Five until the preset convergence condition is reached. Exemplarily, the image in the second clustering image set with the shooting time closest to the obtained shooting time mean value can be taken as the new clustering center; or, the image in the second clustering image set with the shooting location closest to the obtained shooting location mean value can be taken as the new clustering center; or, the image in the second clustering image set with the shooting time closest to the obtained shooting time mean value and the shooting location closest to the obtained shooting location mean value can be taken as the new clustering center. To ensure the consistency of clustering calculation, the factors relied on in the first clustering process and the second clustering process can be kept the same. For example, both the first clustering process and the second clustering process are calculated based on the similarity of the shooting time between the images to be classified, or both the first clustering process and the second clustering process are calculated based on the similarity of the shooting location between the images to be classified, or both the first clustering process and the second clustering process are calculated based on the similarity of both the shooting location and the shooting time between the images to be classified.

[0127] Among them, the preset convergence condition can be different according to different application scenarios, including but not limited to any of the following conditions: the number of loops reaches the preset number threshold, the newly determined clustering center is the same as the clustering center when Step Five was executed last time.

[0128] Next, the calculation of the similarity of scene information will be described. In a possible embodiment, the similarity of the scenes between images can be determined by any of the following methods:

[0129] Method 7: Use a pre-trained scene information similarity calculation model to calculate the similarity between any two images in the third image set.

[0130] Exemplarily, a convolutional neural network includes a convolutional layer and a fully connected layer. Among them, the convolutional layer is used to extract the image features of the input image and input them into the fully connected layer. The fully connected layer calculates the scene similarity based on the image features input by the convolutional layer and outputs the similarity of the scene information.

[0131] As described above, the scene information of the image will be displayed on the image. Therefore, a similarity calculation model can be used to realize the mapping from the image to the similarity of the scene information.

[0132] Method 8: Use a preset feature extraction algorithm to extract the scene information features of the images in the third image set, and calculate the similarity of the scene information between any two images in the third image set based on the extracted features.

[0133] Exemplarily, use a preset feature extraction algorithm to extract the feature points in two images respectively, and determine the similarity of the scene information of the two images according to the similarity of the feature points in the two images.

[0134] Exemplarily, use Hessian-Affine to detect the feature points in two images respectively, and use SIFT as a descriptor to determine the scene similarity of the two images according to the similarity of the feature points in the two images.

[0135] Select Method 7. Since the receptive field of the convolutional neural network is relatively large, the accuracy of the obtained scene can be improved. And if Method 8 is selected, compared with the convolutional neural network, the calculation amount is less, and the efficiency of determining the similarity of the scene information can be effectively improved.

[0136] Next, an exemplary description of how to classify and obtain the second image set will be given. In one possible embodiment, the similarity of the scenes between images can be determined by any of the following methods:

[0137] Method 9: Cluster the images to be classified based on the face features of the images to be classified to obtain the second image set.

[0138] The clustering method can be different according to different application scenarios, such as the aforementioned first clustering process and second clustering process, or other clustering processes other than the aforementioned first clustering process and second clustering process can also be used.

[0139] Method 10: Calculate the face similarity between any two images in the images to be classified, and classify the images to be classified according to the face similarity.

[0140] By selecting Method 9, all the images to be classified can be divided into at least one second image set through clustering, without comparing every two images to be classified, so the efficiency is relatively high. By selecting Method 2, since it is not necessary to cluster all the images to be classified at the same time, and only the face similarity between two images to be classified needs to be calculated each time, the system resources occupied are relatively small. If the executing entity is a device with strong computing performance, such as a server, Method 9 can be selected; if the executing entity is a device with weak computing performance, such as a user terminal, Method 10 can be selected.

[0141] The execution timing of the foregoing S101-S103 can be different according to different application scenarios. Exemplarily, the foregoing S101-S103 can be executed once every preset period.

[0142] In another possible embodiment, it can also be to first respond to the user's image viewing request or the user's album generation request, and then execute the foregoing S101-S103 to divide the images into multiple second image sets according to the people included in the images, and display the second image sets after the target images are classified.

[0143] Exemplarily, the user clicks a specific control, such as the cover of the album composed of second image sets, a preset classification button, etc. The executing entity responds to the user's click operation, classifies one or more target images in the first image set into the second image set according to the foregoing S101-S103, and displays the album composed of the second image sets after the target images are classified.

[0144] In still another possible embodiment, it can also be to first execute the foregoing S101-S103 to divide the images into multiple second image sets according to the people included in the images, and then respond to the user's image viewing request or the user's album generation request, and display the second image sets after the target images are classified.

[0145] Exemplarily, the executing entity classifies one or more target images in the first image set into the second image set according to the foregoing S101-S103. The user clicks a specific control, such as a preset album display button, an image viewing request button, etc. The executing entity responds to the user's click operation and displays the album composed of the second image sets after the target images are classified.

[0146] To more clearly illustrate the face classification method provided by the embodiments of the present invention, specific examples will be described below. Refer to Figure 3 ,

[0147] S301: Divide the images in the image set to be classified into multiple second image sets according to the face similarity.

[0148] For example, it is divided into the following second image sets: the set of Person A, the set of Person B, the set of Person C, etc. Among them, the images in the set of Person A contain the face of Person A, each image in the set of Person B contains the face of Person B, and so on. Regarding the division method of the second image set, reference can be made to the foregoing relevant description, which will not be elaborated here.

[0149] S302: Divide the image set to be classified into multiple third image sets according to the similarity of shooting time and shooting location.

[0150] For example, it is divided into the following third image sets: the set of Time 1 & Location 1, the set of Time 1 & Location 2, the set of Time 2 & Location 1, etc. Among them, the shooting time of each image in the set of Time 1 & Location 1 is the same as or close to Time 1, and the shooting location is the same as or close to Location 1. The shooting time of each image in the set of Time 1 & Location 2 is the same as or close to Time 1, and the shooting location is the same as or close to Location 2, and so on. This step is equivalent to the aforementioned S1011.

[0151] S303: For each third image set, divide it into multiple first image sets according to the similarity of scene information.

[0152] For example, it is divided into the following first image sets: the set of Scene 1, the set of Scene 2, the set of Scene 3. Taking the set of Time 2 & Location 1 as an example, within the set of Scene 1 divided from the set of Time 2 & Location 1, the shooting time of each image is the same as or close to Time 2, and the shooting location is the same as or close to Location 1, and the shooting scene is the same as or close to Scene 1. Within the set of Scene 2 divided from the set of Time 2 & Location 1, the shooting time of each image is the same as or close to Time 2, and the shooting location is the same as or close to Location 1, and the shooting scene is the same as or close to Scene 2, and so on. This step is equivalent to the aforementioned S1012.

[0153] S304: For each first image set, find the images that already belong to the second image set within this scene image set. And for the found images, search for the matching images within this first image set, and add these matching images to the first image set to which the image belongs. This step is equivalent to the aforementioned S102 and S103.

[0154] Exemplarily, as Figure 4 described, assume that the set of Scene 1 includes three images that already belong to the second image set, which are respectively denoted as the first images 1 - 3, and includes five images that do not yet belong to the second image set, which are respectively denoted as the second images 1 - 5. Among them, the first images 1 - 2 contain the face of Person A, so they are divided into the set of Person A, while the first images 2 - 3 contain the face of Person B, so they are divided into the set of Person B.

[0155] If the second image 1-2 matches the first image 1, the second image 3 matches the first image 2, the second image 4 matches the first image 3, and the second image 5 does not match any of the first images 1-3, then add the second images 1-3 to the set of person A, add the second images 4-5 to the set of person B, and consider the second image 5 as a landscape image that does not contain a face (or only contains a small part of a face).

[0156] Corresponding to the foregoing face classification method, an embodiment of the present invention further provides a face classification device, as Figure 4 shown, including:

[0157] A first determination module 401, configured to determine a first image set, where the description information between different images in the first image set is similar, and where the similar description information includes the same or similar description information;

[0158] A second determination module 402, configured to determine a second image set having the same images as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold;

[0159] A classification module 403, configured to determine whether a target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, where the target image includes any image that is different from the images in the second image set.

[0160] In a possible embodiment, the classification module determines whether a target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, including:

[0161] Calculating a classification value of the target image based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set;

[0162] Classifying the target image with a classification value greater than a second threshold into the second image set;

[0163] The description information of the image at least includes shooting time, shooting location, and scene information shown in the image;

[0164] The first determination module determines a first image set, including:

[0165] Determine a third image set, where the shooting time and / or shooting location between any two images in the third image set are similar;

[0166] Based on the images in the third image set whose scene information similarity is greater than a third threshold, obtain the first image set;

[0167] The first determination module determines the third image set, including:

[0168] Based on the similarity of the shooting time and / or shooting location between the images to be classified, perform a first clustering process on the images to be classified to determine an initial clustering center;

[0169] Based on the similarity of the shooting time and / or shooting location between the images to be classified, perform a second clustering process on the images to be classified according to the initial clustering center to obtain the third image set;

[0170] The first determination module, based on the similarity of the shooting time and / or shooting location between the images to be classified, performs a first clustering process on the images to be classified to determine an initial clustering center, including:

[0171] Based on the similarity of the shooting time and / or shooting location between the images to be classified, divide the images to be classified into multiple first clustering image sets, where the similarity of the shooting time and / or shooting location between any two images in the first clustering image set is greater than a fourth threshold;

[0172] The first module, based on the similarity of the shooting time and / or shooting location between the images to be classified, performs a second clustering process on the images to be classified according to the initial clustering center to obtain the third image set, including:

[0173] Based on the similarity of the shooting time and / or shooting location between the images to be classified and each clustering center, classify all the images to be classified to obtain multiple second clustering image sets, where each clustering center is initially the initial clustering center;

[0174] For each of the second clustering image sets, determine the mean of the shooting time and / or shooting location of all the images in the second clustering image set, and use the image in the second clustering image set whose shooting time and / or shooting location is closest to the mean as a new clustering center, and return to execute the step of classifying all the images to be classified based on the similarity of the shooting time and / or shooting location between the images to be classified and each clustering center to obtain multiple second clustering image sets; until a preset convergence condition is reached, determine the second clustering image set as the third image set;

[0175] The device further includes a similarity calculation module, configured to calculate the similarity of scene information between any two images in the third image set by using a pre-trained scene information similarity calculation model; or

[0176] extract the scene information features of the images in the third image set by using a preset feature extraction algorithm, and calculate the similarity of scene information between any two images in the third image set based on the extracted features;

[0177] The device further includes a set partitioning module, configured to cluster the images to be classified based on the face features of the images to be classified, so as to obtain a second image set;

[0178] Or

[0179] calculate the face similarity between any two images in the images to be classified, and classify the images to be classified according to the face similarity to obtain a second image set;

[0180] The target images in the first image set include: images with the face in an occluded state, images with the face in a side face state, images obtained by shooting the face from above or from below;

[0181] The device further includes a start module, configured to respond to a user's image viewing request or a user's album set generation request;

[0182] The device further includes a display module, configured to display the second image set after the target images are classified.

[0183] An embodiment of the present invention further provides an electronic device, as Figure 5 shown, including:

[0184] A memory 501, configured to store a computer program;

[0185] A processor 502, configured to implement the following steps when executing the program stored in the memory 501:

[0186] Determine a first image set, where the description information between different images in the first image set is similar, where the similar description information includes the same or similar description information;

[0187] Determine a second image set having the same images as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold;

[0188] Determine whether the target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the images in the second image set, and / or the similarity of the description information between the target image in the first image set and the images in the second image set, where the target image includes any image that is different from the images in the second image set.

[0189] In addition to the memory and the processor, the electronic device provided by the embodiments of the present invention may further include other components, such as a communication bus and a communication interface.

[0190] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0191] The communication interface is used for communication between the above electronic device and other devices.

[0192] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0193] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0194] In another embodiment provided by the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned face classification method are implemented.

[0195] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which, when running on a computer, causes the computer to execute any of the face classification methods in the above embodiments.

[0196] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).

[0197] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0198] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0199] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.

Claims

1. A face classification method, characterized in that, The method includes: Determine a first image set, where the description information between different images in the first image set is similar, and where the similarity of the description information includes the same or similar description information; Determine a second image set that has the same images as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold; Based on the face similarity between a target image in the first image set and an image in the second image set, and the similarity of the description information between the target image in the first image set and the image in the second image set, determine whether the target image in the first image set is classified into the second image set, where the target image includes any image that is different from the images in the second image set; The description information of the image at least includes the shooting time, shooting location, and scene information shown in the image; The determining of the first image set includes: Determine a third image set, where the shooting time and / or shooting location between any two images in the third image set are similar; Based on the images in the third image set whose scene information similarity is greater than a third threshold, obtain the first image set; The determining of the third image set includes: Based on the similarity of the shooting time and / or shooting location between the images to be classified, perform a first clustering process on the images to be classified to determine an initial clustering center; Based on the similarity of the shooting time and / or shooting location between the images to be classified, and according to the initial clustering center, perform a second clustering process on the images to be classified to obtain the third image set.

2. The method according to claim 1, characterized in that, The determining whether the target image in the first image set is classified into the second image set based on the face similarity between the target image in the first image set and the image in the second image set, and the similarity of the description information between the target image in the first image set and the image in the second image set includes: Based on the face similarity between the target image in the first image set and the image in the second image set, and the similarity of the description information between the target image in the first image set and the image in the second image set, calculate the classification value of the target image; Classify the target image whose classification value is greater than a second threshold into the second image set.

3. The method according to claim 1, characterized in that, The performing of the first clustering process on the images to be classified based on the similarity of the shooting time and / or shooting location between the images to be classified to determine an initial clustering center includes: Based on the similarity of the shooting time and / or shooting location between the images to be classified, divide the images to be classified into multiple first clustering image sets, where the similarity of the shooting time and / or shooting location between any two images in the first clustering image set is greater than a fourth threshold; For each of the first clustering image sets, select an image in the first clustering image set as the initial clustering center; Performing a second clustering process on the images to be classified according to the initial clustering centers based on the similarity of shooting time and / or shooting location between the images to be classified, to obtain the third image set, including: Classifying all the images to be classified based on the similarity of shooting time and / or shooting location between the images to be classified and each clustering center, to obtain a plurality of second clustering image sets, where each of the clustering centers is initially the initial clustering center; For each of the second clustering image sets, determining the mean value of the shooting time and / or shooting location of all the images in the second clustering image set, and using the image in the second clustering image set whose shooting time and / or shooting location is closest to the mean value as a new clustering center, and returning to execute the step of classifying all the images to be classified based on the similarity of shooting time and / or shooting location between the images to be classified and each clustering center to obtain a plurality of second clustering image sets; until a preset convergence condition is reached, determining the second clustering image set as the third image set.

4. The method according to claim 1, characterized in that, The method further includes: Calculating the similarity of scene information between any two images in the third image set by using a pre-trained scene information similarity calculation model; or Extracting scene information features of the images in the third image set by using a preset feature extraction algorithm, and calculating the similarity of scene information between any two images in the third image set based on the extracted features.

5. The method according to claim 1, characterized in that The method further includes: Clustering the images to be classified based on the face features of the images to be classified, to obtain a second image set; Or, Calculating the face similarity between any two images in the images to be classified, and classifying the images to be classified according to the face similarity, to obtain a second image set.

6. The method according to claim 1, characterized in that The target images in the first image set include: images with the face in an occluded state, images with the face in a side face state, images obtained by shooting the face from a top view or a bottom view.

7. The method according to claim 1, wherein The method further includes: Responding to a user's image viewing request, or responding to a user's album set generation request; Displaying the second image set after classifying the target images.

8. A face classification device, characterized in that, The apparatus includes: A first determination module, configured to determine a first image set, where the description information between different images in the first image set is similar, where the similarity of the description information includes that the description information is the same or similar; A second determination module, configured to determine a second image set having the same images as the first image set, where the face similarity between any two images in the second image set is greater than a first threshold; A classification module, configured to determine whether the target images in the first image set are classified into the second image set based on the face similarity between the target images in the first image set and the images in the second image set, and the similarity of the description information between the target images in the first image set and the images in the second image set, where the target images include any image that is not the same as the images in the second image set; The description information of the image includes at least the shooting time, shooting location, and the scene information shown in the image; The determination of the first image set includes: Determining a third image set, where the shooting time and / or shooting location between any two images in the third image set are similar; Based on the images in the third image set whose scene information similarity is greater than a third threshold, obtaining the first image set; The determination of the third image set includes: Based on the similarity of the shooting time and / or shooting location between the images to be classified, performing a first clustering process on the images to be classified to determine the initial clustering centers; Based on the similarity of the shooting time and / or shooting location between the images to be classified, and according to the initial clustering centers, performing a second clustering process on the images to be classified to obtain the third image set.

9. The device according to claim 8, characterized in that, The classification module determines whether the target images in the first image set are classified into the second image set based on the face similarity between the target images in the first image set and the images in the second image set, and the similarity of the description information between the target images in the first image set and the images in the second image set, including: Based on the face similarity between the target images in the first image set and the images in the second image set, and the similarity of the description information between the target images in the first image set and the images in the second image set, calculating the classification value of the target image; Classifying the target images with the classification value greater than a second threshold into the second image set; The first determination module performs a first clustering process on the images to be classified based on the similarity of the shooting time and / or shooting location between the images to be classified to determine the initial clustering centers, including: Based on the similarity of the shooting time and / or shooting location between the images to be classified, dividing the images to be classified into multiple first clustering image sets, where the similarity of the shooting time and / or shooting location between any two images in the first clustering image sets is greater than a fourth threshold; For each of the first clustering image sets, selecting an image in the first clustering image set as the initial clustering center; The first determination module performs a second clustering process on the images to be classified based on the similarity of the shooting time and / or shooting location between the images to be classified and according to the initial clustering centers to obtain the third image set, including: Based on the similarity of the shooting time and / or shooting location between the images to be classified and each clustering center, classifying all the images to be classified to obtain multiple second clustering image sets, where each clustering center is initially the initial clustering center; For each of the second clustered image sets, determine the mean of the shooting times and / or shooting locations of all the images in the second clustered image set, and use the image in the second clustered image set whose shooting time and / or shooting location is closest to the mean as the new clustering center, and return the step of performing the classification of all the images to be classified based on the similarity of the shooting times and / or shooting locations between the images to be classified and each clustering center to obtain multiple second clustered image sets; until a preset convergence condition is reached, determine the second clustered image set as the third image set; The device further includes a similarity calculation module, configured to calculate the similarity of the scene information between any two images in the third image set by using a pre-trained scene information similarity calculation model; or extract the scene information features of the images in the third image set by using a preset feature extraction algorithm, and calculate the similarity of the scene information between any two images in the third image set based on the extracted features; The device further includes: a set partitioning module, configured to cluster the images to be classified based on the face features of the images to be classified to obtain a second image set; Or, calculate the face similarity between any two images in the images to be classified, and classify the images to be classified according to the face similarity to obtain a second image set; The target images in the first image set include: images with the face in an occluded state, images with the face in a side face state, images obtained by shooting the face from above or from below; The device further includes a start module, configured to respond to a user's image viewing request or a user's album set generation request; The device further includes a display module, configured to display the second image set after the classification of the target images.

10. An electronic device, characterized in that, Comprising: a memory for storing a computer program; a processor, configured to implement the method steps described in any one of claims 1-7 when executing the program stored in the memory.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method steps described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Archives management method, device and electronic device for urban security monitoring

    CN109145844A

  • Face recognition photo classification method and device, electronic equipment and storage medium

    CN113343920A