Face feature confidence model construction method and device, equipment and medium
By constructing a face feature confidence model, using the sample face picture data set marked with identity tags, calculate the face feature similarity, determine the positive and negative sample data, and train the machine learning model, the problem of low confidence in the face feature in the face recognition model is solved and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202311524632.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
When extracting facial features, the confidence level of existing face recognition models is not high enough, resulting in the problem of one more person. The feature quality of the feature extraction model based on deep learning is not completely in line with intuition due to various factors.
By obtaining the sample face picture data set marked with the personnel identity tag, the first face feature similarity of each pair of sample face pictures is calculated, the sample face picture group set is formed, and the negative sample data and positive sample data are determined. The training data set is constructed based on these data, and the machine learning model is trained to generate a face feature confidence model.
The generated face feature confidence model is more objective and robust, improving the recognition accuracy of face feature confidence and reducing the occurrence of one more person.
Smart Images

Figure CN120014389A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method, device, equipment and medium for constructing a facial feature confidence model. Background Art
[0002] As one of the core businesses in the current security field, portrait clustering has the problem of multiple people in one file, that is, multiple different people are classified into the same file. One of the main reasons for this phenomenon is that the facial features extracted by the face recognition model may not have sufficient confidence. There are many factors that affect the quality (confidence) of facial features, such as the camera (shooting angle, occlusion, etc.), the person's dress (wearing hats and glasses, etc.) and model training (data source is not comprehensive enough), which may lead to errors in the confidence judgment of facial features. Therefore, how to accurately judge the confidence of facial features is one of the issues worthy of attention in current model training.
[0003] In related technologies, a deep learning model based on the similarity of the best pictures in a class is used to determine the similarity of facial features. However, when annotating training data, the selection of the best picture in a class is highly subjective. For example, it is difficult to subjectively express whether a slightly blurry frontal face picture or a clearer side face picture is better. In fact, the feature extraction model based on deep learning is very robust, and its feature quality is affected by various factors and is not completely intuitive. That is, pictures that are subjectively good are not necessarily of high feature quality. Therefore, this training data annotation method has great limitations. Summary of the invention
[0004] The present invention provides a method, device, equipment and medium for constructing a facial feature confidence model, which can generate a more objective and robust facial feature confidence model based on a training data set composed of positive and negative samples, thereby improving the recognition accuracy of facial feature confidence.
[0005] According to one aspect of the present invention, a method for constructing a facial feature confidence model is provided, the method comprising:
[0006] Obtain a sample face image data set; wherein the sample face image data set is composed of sample face images marked with person identity tags;
[0007] Taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first face feature similarity between the two sample face pictures included in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold;
[0008] Determine a target sample face picture group marked with different person identity labels in the sample face picture group set, and use the sample face pictures included in the target sample face picture group as negative sample data;
[0009] Clustering the sample face pictures contained in the sample face picture group set, determining a target cluster that meets a preset condition from the cluster clusters generated by clustering, and using the sample face pictures in the target cluster except the negative sample data as positive sample data;
[0010] A training data set is formed based on the negative sample data and the positive sample data, and a preset machine learning model is trained based on the training data set to generate a facial feature confidence model.
[0011] According to another aspect of the present invention, there is provided a device for constructing a facial feature confidence model, comprising:
[0012] A sample face picture data set acquisition module is used to acquire a sample face picture data set; wherein the sample face picture data set is composed of sample face pictures marked with person identity tags;
[0013] a sample face picture group set determination module, configured to take every two sample face pictures in the sample face picture data set as a sample face picture group, calculate the first face feature similarity between the two sample face pictures included in each sample face picture group, and form a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold;
[0014] A negative sample data determination module is used to determine a target sample face picture group marked with different person identity labels in the sample face picture group set, and use the sample face pictures included in the target sample face picture group as negative sample data;
[0015] A positive sample data determination module is used to cluster the sample face pictures contained in the sample face picture group set, determine a target cluster that meets a preset condition from the cluster clusters generated by clustering, and use the sample face pictures in the target cluster except the negative sample data as positive sample data;
[0016] The facial feature confidence model generation module is used to form a training data set based on the negative sample data and the positive sample data, and to train a preset machine learning model based on the training data set to generate a facial feature confidence model.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] at least one processor; and
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for constructing a facial feature confidence model described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for constructing a facial feature confidence model described in any embodiment of the present invention when executed.
[0022] The technical solution of the embodiment of the present invention comprises the following steps: obtaining a sample face picture data set; wherein the sample face picture data set is composed of sample face pictures annotated with person identity labels; taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first face feature similarity between the two sample face pictures contained in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold; determining a target sample face picture group annotated with different person identity labels in the sample face picture group set, and taking the sample face pictures contained in the target sample face picture group as negative sample data; clustering the sample face pictures contained in the sample face picture group set, determining a target cluster that meets preset conditions from the cluster clusters generated by clustering, and taking the sample face pictures in the target cluster except the negative sample data as positive sample data; forming a training data set based on the negative sample data and the positive sample data, and training a preset machine learning model based on the training data set to generate a face feature confidence model. This technical solution can generate a more objective and robust facial feature confidence model based on a training data set consisting of positive and negative samples, thereby improving the recognition accuracy of facial feature confidence.
[0023] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 is a flow chart of a method for constructing a facial feature confidence model provided according to the first embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a similarity relationship graph provided according to Embodiment 1 of the present invention;
[0027] Figure 3 is a schematic diagram of a cluster provided according to Embodiment 1 of the present invention;
[0028] Figure 4 It is a flowchart of a method for constructing a facial feature confidence model provided in accordance with the first embodiment of the present invention;
[0029] Figure 5 is a flow chart of a method for constructing a facial feature confidence model provided in accordance with Embodiment 2 of the present invention;
[0030] Figure 6 is a schematic diagram of the structure of a device for constructing a facial feature confidence model provided according to Embodiment 3 of the present invention;
[0031] Figure 7 It is a structural schematic diagram of an electronic device for implementing a method for constructing a facial feature confidence model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] Embodiment 1
[0035] Figure 1 This is a flowchart of a method for constructing a facial feature confidence model provided in the first embodiment of the present invention. This embodiment is applicable to the case where a facial feature confidence model is constructed based on a training data set consisting of positive and negative samples. The method can be executed by a facial feature confidence model construction device. The facial feature confidence model construction device can be implemented in the form of hardware and / or software. The facial feature confidence model construction device can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:
[0036] S110, obtaining a sample face image dataset; wherein the sample face image dataset is composed of sample face images annotated with person identity tags.
[0037] In this embodiment, a sample face image dataset is constructed in advance by manual annotation. The sample face image dataset may include a large number of sample face images, each sample face image is annotated with the identity label of the person to which it belongs, and the sample face images corresponding to the same identity label represent the same person. Exemplarily, the sample face image dataset may be represented as shown in Table 1:
[0038] Table 1 Sample face image dataset
[0039] Sample face image name a.jpg b.jpg c.jpg … Personnel ID tags 1 2 2 …
[0040] As shown in Table 1, the sample face images b.jpg and c.jpg correspond to the same person, and are different from the person corresponding to the sample face image a.jpg.
[0041] S120, taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first facial feature similarity between the two sample face pictures included in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first facial feature similarity is greater than a preset similarity threshold.
[0042] In this embodiment, after obtaining the sample face picture data set, every two sample face pictures in the sample face picture data set can be used as a sample face picture group, and then the sample face pictures in each sample face picture group are respectively subjected to feature extraction, and the cosine distance or Euclidean distance can be used to calculate the similarity of the first face features between the two sample face pictures contained in each sample face picture group. Among them, feature extraction can be implemented by SIFT (scale invariant feature transform), HOG (histogram of oriented gradients) or neural network model, etc., which is not specifically limited in this embodiment. Exemplarily, the first face feature similarity can be characterized by a triple structure, wherein the triple structure can be expressed as: (unique identifier of picture A, unique identifier of picture B, similarity of the first face features between picture A and picture B). Assuming there are n sample face pictures, n*(n-1) / 2 triplets will be generated.
[0043] After determining the similarity of the first facial features, each first facial feature similarity can be compared with a preset similarity threshold, and the sample face picture groups whose first facial feature similarity is greater than the preset similarity threshold constitute a sample face picture group set. The preset similarity threshold can be set according to actual application requirements. For example, the preset similarity threshold can be set to the threshold used to determine whether two pictures belong to the same identity during face clustering. Exemplarily, each triple generated above can be filtered according to the preset similarity threshold, retaining a triple list L whose first facial feature similarity is greater than the preset similarity threshold, and constituting a sample face picture group set based on the sample face pictures in the triple list L.
[0044] In this embodiment, optionally, the similarity of the first facial features between two sample face pictures included in each sample face picture group is calculated, including: for each sample face picture group, respectively extracting facial features in the two sample face pictures included in the sample face picture group based on the face recognition model to be evaluated; and calculating the similarity of the first facial features between the two sample face pictures included in the sample face picture group based on the facial features.
[0045] Wherein, the face recognition model can be used to extract face features from face images. In this embodiment, when calculating the first face feature similarity, for each sample face image group, the face recognition model to be evaluated can be used to extract face features from two sample face images included in the sample face image group, and then the first face feature similarity between the two sample face images included in the sample face image group can be calculated using the cosine distance or the Euclidean distance based on the face features.
[0046] Through such a setting, the present scheme can calculate the similarity of the first facial features between two sample face pictures contained in the sample face picture group through the face recognition model to be evaluated, and use the set of sample face picture groups obtained by screening based on the first facial feature similarity as the training data basis of the facial feature confidence model, so as to facilitate the subsequent use of the facial feature confidence model to perform high-precision evaluation of the facial feature confidence obtained by the face recognition model.
[0047] S130, determining a target sample face picture group annotated with different person identity labels in the sample face picture group set, and using the sample face pictures included in the target sample face picture group as negative sample data.
[0048] Exemplarily, the above triple list L is traversed, and for each of the two images in the triple list, the corresponding person identity label is obtained from the sample face image dataset according to the image ID, and the triple list L in which the two images in the triple list do not belong to the same person identity label is found. low , and based on the triple list L low The sample face pictures in determine the target sample face picture group. The triple list L low The pictures in can be considered as low-confidence pictures, because the correct labels of the two faces cannot be effectively distinguished under the face recognition model.
[0049] In this embodiment, after determining the target sample face picture group annotated with different person identity labels in the sample face picture group set, the sample face pictures included in the target sample face picture group can be used as negative sample data. Optionally, using the sample face pictures included in the target sample face picture group as negative sample data includes: annotating each sample face picture in the target sample face picture group as having low confidence in facial features, and using the annotated sample face pictures in the target sample face picture group as negative sample data.
[0050] For example, the triple list L low Each triplet in is split into two confidence label data, and each confidence label data is used as negative sample data for subsequent training of the face feature confidence model. Among them, the confidence label data can be used to mark the features extracted by the face recognition model for these two pictures as having low confidence. Assume that the triplet list L low If one of the triples in is identified as (unique identifier of picture A, unique identifier of picture B, 95%), the confidence label data generated after splitting can be represented as (unique identifier of picture A, 0) and (unique identifier of picture B, 0). Among them, 0 means low confidence of facial features, that is, the picture is not credible. Because the pictures here are all low confidence pictures, the corresponding labels are all 0.
[0051] S140, clustering the sample face pictures included in the sample face picture group set, determining a target cluster that meets a preset condition from the cluster clusters generated by clustering, and using the sample face pictures in the target cluster except the negative sample data as positive sample data.
[0052] In this embodiment, after determining the negative sample data based on the sample face picture group set, it is also necessary to determine the positive sample data of the same magnitude as the negative sample data, that is, the facial feature high confidence data. First, the sample face pictures contained in the sample face picture group set are clustered, and a target cluster that meets the preset conditions is determined from the cluster clusters generated by clustering. Optionally, the sample face pictures contained in the sample face picture group set are clustered, and the target cluster that meets the preset conditions is determined from the cluster clusters generated by clustering, including: constructing a similarity relationship graph based on the first facial feature similarity corresponding to each sample face picture group in the sample face picture group set, and splitting the similarity relationship graph into cluster clusters based on a clustering algorithm; wherein each cluster cluster contains at least one sample face picture; and determining a target cluster that meets the preset conditions from the cluster clusters.
[0053] Exemplarily, firstly, based on the triple list L obtained by filtering the first facial feature similarity, a similarity relationship graph is constructed, such as Figure 2 As shown in the figure, each letter represents a sample face image. Then, the similarity relationship graph is split into multiple subgraphs (i.e., clusters) using a clustering algorithm (such as a community discovery algorithm or a graph segmentation algorithm). The splitting results are shown in the figure below. Figure 3 As shown. Each unconnected cluster is a subgraph, which may contain one or more sample face images. For each split subgraph, the categories of all the images contained in it are obtained from the sample face image dataset. If the images in a subgraph meet the preset conditions, the subgraph can be determined as the target cluster. The preset conditions can be set as: all images in the subgraph belong to the same person identity label, and all images under the person identity label in the sample face image dataset are in this subgraph.
[0054] In this embodiment, optionally, a target cluster that meets preset conditions is determined from the cluster clusters, including: for each cluster cluster, determining the target person identity label annotated with the sample face images contained in the cluster cluster; when the target person identity label is the same person identity label, and all sample face images annotated with the target person identity label in the sample face image data set are included in the cluster cluster, the cluster cluster is used as the target cluster.
[0055] Through such a setting, this solution determines the target cluster based on the above preset conditions, which helps to improve the accuracy of the target cluster.
[0056] In this embodiment, after determining the target cluster, sample face pictures other than negative sample data in the target cluster can be used as positive sample data. Optionally, using sample face pictures other than negative sample data in the target cluster as positive sample data includes: deleting sample face pictures in the target sample face picture group from the target cluster to generate a target sample face picture set; annotating each sample face picture in the target sample face picture set as a high confidence level of facial features, and using the annotated sample face pictures in the target sample face picture set as positive sample data.
[0057] Specifically, after finding the target cluster that meets the preset conditions, exclude the triple list L low The remaining pictures can be considered as high-confidence pictures. For example, the generated positive sample data can be expressed as: (unique identifier of picture C, 1), (unique identifier of picture D, 1), .... Among them, 1 indicates high confidence of facial features, that is, the picture is credible.
[0058] Through such a setting, this scheme obtains positive sample data by removing negative sample data from the target cluster, which helps to improve the accuracy of positive sample data.
[0059] S150, constructing a training data set based on the negative sample data and the positive sample data, and training a preset machine learning model based on the training data set to generate a facial feature confidence model.
[0060] In this embodiment, after determining the negative sample data and the positive sample data, the negative sample data and the positive sample data can be shuffled to form a training data set, and the preset machine learning model is trained based on the training data set to generate a facial feature confidence model. The preset machine learning model can be a convolutional neural network, specifically including an input layer, a convolutional layer, a fully connected layer, and an output layer.
[0061] Specifically, the input layer input is a training data set composed of shuffled negative sample data and positive sample data, and each sample is a picture plus a confidence label (0 or 1). Multiple convolutional layers are stacked after the input layer to extract low-level and high-level features of the picture from the RGB information of the picture. Each convolutional layer is batch normalized and activated by ReLU or LeakyReLU function after using maximum pooling. Then, multiple fully connected layers are used to map the convolutional layer to a neuron, and the sigmoid function is used to map the result to a probability value (between 0 and 1). The closer the value is to 0, the lower the feature confidence. The log-likelihood loss function is used for gradient calculation, and the weight parameters of each layer are updated using the gradient descent method. After multiple epochs of training, a facial feature confidence model can be fitted, which can infer the confidence of the facial features extracted by the face recognition model to be evaluated based on the input image information. Exemplarily, the output of the facial feature confidence model may be a floating point number between 0 and 1. The larger the number is, the higher the facial feature confidence is, that is, the better the facial feature quality is.
[0062] It should be noted that the generated facial feature confidence model is not used to judge the subjective quality of a picture, but to judge whether the facial features extracted under a specific face recognition model have sufficient recognition and discrimination. Therefore, the facial feature confidence model can well distinguish which pictures are the weak points of the current face recognition model. After finding the weak points, the weak points are eliminated, and then a better face clustering effect is achieved without changing the face recognition model.
[0063] Figure 4 The following is a flow chart of a method for constructing a facial feature confidence model provided in the first embodiment of the present invention. Figure 4 As shown in the figure, first construct a sample face image dataset, then calculate the facial feature similarity of each two images in the sample face image dataset to generate a similarity triple, and filter the similarity triple according to the preset similarity threshold to obtain a triple list L. Find the triple list L in which the two images do not belong to the same person identity label from the triple list L low , and the triple list L low The face images in are used as negative sample data. At the same time, a similarity relationship graph is constructed based on the triple list L, and the similarity relationship graph is segmented to form subgraphs (i.e., clusters). The subgraphs that completely match the annotated person identity labels (i.e., meet the preset conditions) are found, and the face images in them are used as positive sample data. Finally, the convolutional neural network is trained based on the positive sample data and the negative sample data to obtain the face feature confidence model.
[0064] The technical solution of the embodiment of the present invention comprises the following steps: obtaining a sample face picture data set; wherein the sample face picture data set is composed of sample face pictures annotated with person identity labels; taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first face feature similarity between the two sample face pictures contained in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold; determining a target sample face picture group annotated with different person identity labels in the sample face picture group set, and taking the sample face pictures contained in the target sample face picture group as negative sample data; clustering the sample face pictures contained in the sample face picture group set, determining a target cluster that meets preset conditions from the cluster clusters generated by clustering, and taking the sample face pictures in the target cluster except the negative sample data as positive sample data; forming a training data set based on the negative sample data and the positive sample data, and training a preset machine learning model based on the training data set to generate a face feature confidence model. This technical solution can generate a more objective and robust facial feature confidence model based on a training data set consisting of positive and negative samples, thereby improving the recognition accuracy of facial feature confidence.
[0065] In this embodiment, optionally, after generating the facial feature confidence model, it also includes: obtaining a first target facial image; obtaining a first facial feature confidence of the first target facial image based on the facial feature confidence model, and taking the first target facial image whose first facial feature confidence is greater than a preset confidence threshold as the second target facial image; and extracting the first facial feature of the second target facial image based on the facial recognition model to be evaluated.
[0066] In this embodiment, for each face image, while extracting face features using the face recognition model to be evaluated, the face feature confidence model can also be used to obtain face feature confidence. Before clustering face features, the face feature data can be filtered using a preset confidence threshold, and face images with face feature confidence greater than the preset confidence threshold are selected for subsequent feature extraction and clustering processing. The confidence threshold can be set according to actual application requirements.
[0067] Through such a setting, this solution filters the facial feature data based on a preset confidence threshold before clustering the facial features, and selects face images whose facial feature confidence is greater than the preset confidence threshold for clustering processing, which helps to improve the clustering accuracy of facial features.
[0068] In this embodiment, optionally, after generating the facial feature confidence model, it also includes: obtaining at least two third target facial images; obtaining second facial feature confidences of at least two third target facial images respectively based on the facial feature confidence model; for every two third target facial images of the at least two third target facial images, adjusting the second facial feature similarity between every two third target facial images based on the second facial feature confidences corresponding to every two third target facial images; and clustering the at least two third target facial images based on the adjusted second facial feature similarities.
[0069] In this embodiment, before clustering facial features, in addition to filtering the facial feature data using a preset confidence threshold, the facial feature similarity of the two images can be appropriately reduced using a preset adjustment algorithm based on the facial feature confidence of the two images, thereby improving the accuracy of the facial feature similarity between the two images.
[0070] Specifically, first obtain at least two third target face images, then obtain the second face feature confidences of at least two third target face images respectively through the face feature confidence model, and extract the second face features of at least two third target face images respectively based on the face recognition model to be evaluated. For each two third target face images in the at least two third target face images, calculate the second face feature similarity between the two third target face images based on the second face feature, and then adjust the second face feature similarity based on the second face feature confidences corresponding to the two third target face images, and finally cluster the at least two third target face images based on the adjusted second face feature similarity. It can be understood that if the second face feature confidences corresponding to the two third target face images are low, then the adjustment result of the second face feature similarity is that the second face feature similarity is reduced.
[0071] Through such a setting, this solution reduces the similarity of facial features based on the confidence of facial features before clustering the facial features, which can improve the accuracy of the similarity of facial features and further help improve the clustering accuracy of facial features.
[0072] Embodiment 2
[0073] Figure 5A flowchart of a method for constructing a facial feature confidence model provided in the second embodiment of the present invention, this embodiment is optimized based on the above embodiment. The specific optimization is: after generating the facial feature confidence model, it also includes: obtaining at least two spatiotemporal trajectory face pictures; wherein the spatiotemporal trajectory face pictures are face pictures whose collection time interval is less than a preset time length and the distance between the collection positions is greater than a preset distance threshold; for each two spatiotemporal trajectory face pictures in the at least two spatiotemporal trajectory face pictures, when the facial feature similarity between the two spatiotemporal trajectory face pictures is greater than the preset similarity threshold, the two spatiotemporal trajectory face pictures are added to the negative sample data to update the facial feature confidence model.
[0074] like Figure 5 As shown, the method of this embodiment specifically includes the following steps:
[0075] S210, obtaining a sample face image dataset; wherein the sample face image dataset is composed of sample face images annotated with person identity tags.
[0076] S220, taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first facial feature similarity between the two sample face pictures included in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first facial feature similarity is greater than a preset similarity threshold.
[0077] S230, determining a target sample face picture group annotated with different person identity labels in the sample face picture group set, and using the sample face pictures included in the target sample face picture group as negative sample data.
[0078] S240, clustering the sample face pictures included in the sample face picture group set, determining a target cluster that meets a preset condition from the cluster clusters generated by clustering, and using the sample face pictures in the target cluster except the negative sample data as positive sample data.
[0079] S250, constructing a training data set based on the negative sample data and the positive sample data, and training a preset machine learning model based on the training data set to generate a facial feature confidence model.
[0080] Among them, the specific implementation method of S210-S250 can refer to the detailed description in S110-S150, which will not be repeated here.
[0081] S260, obtaining at least two spatiotemporal trajectory face images; wherein the spatiotemporal trajectory face images are face images whose collection time interval is less than a preset time length and whose collection positions are separated by a distance greater than a preset distance threshold.
[0082] In this embodiment, after generating the facial feature confidence model, the model can also be fine-tuned through incremental training to adapt to the characteristics of images in different application scenarios. First, at least two spatiotemporal trajectory face images need to be obtained. The spatiotemporal trajectory face images can be used to characterize that the collection time interval of the face images is less than the preset time length and the distance between the collection positions is greater than the preset distance threshold. Among them, the preset time length and the preset distance threshold can be set according to actual application requirements. Different faces classified into the same file can be found using abnormal spatiotemporal trajectories of faces. Among them, abnormal spatiotemporal trajectories can be understood as the same person appearing in two very far locations in a very short period of time, which is impossible to happen, so it can be considered that the two face images are not the same person.
[0083] S270, for each two spatiotemporal trajectory face images of the at least two spatiotemporal trajectory face images, when the facial feature similarity between the two spatiotemporal trajectory face images is greater than a preset similarity threshold, the two spatiotemporal trajectory face images are added to the negative sample data to update the facial feature confidence model.
[0084] In this embodiment, after obtaining at least two spatiotemporal trajectory face images, the facial feature similarity between the two images can be calculated respectively. If the facial feature similarity is greater than the preset similarity threshold, the corresponding spatiotemporal trajectory face image can be used as the newly added negative sample data, and the original negative sample data can be expanded based on the newly added negative sample data, while the positive sample data still uses the previously used samples. The method of obtaining training samples based on abnormal spatiotemporal trajectories does not require manual annotation, and thus has the effect of unsupervised learning. Then, an incremental training data set is constructed based on the original positive sample data and the expanded negative sample data, and the facial feature confidence model is updated based on the incremental training data set. Among them, the incremental training method is similar to the previous model training method. On the basis of the existing model network architecture, the trained weight parameters are loaded, and the incremental training data set is input on this basis. In the iterative training process, the weight parameters are further adjusted using gradient descent to adapt to the distribution of the incremental data.
[0085] The technical solution of the embodiment of the present invention, after generating the facial feature confidence model, obtains at least two spatiotemporal trajectory face images; wherein the spatiotemporal trajectory face images are face images whose collection time interval is less than a preset time length and the distance between the collection positions is greater than a preset distance threshold; respectively, for each two spatiotemporal trajectory face images in the at least two spatiotemporal trajectory face images, when the facial feature similarity between the two spatiotemporal trajectory face images is greater than a preset similarity threshold, the two spatiotemporal trajectory face images are added to the negative sample data to update the facial feature confidence model. This technical solution can generate a more objective and robust facial feature confidence model based on a training data set composed of positive and negative samples, improve the recognition accuracy of facial feature confidence, and can also effectively expand the negative sample data based on abnormal spatiotemporal trajectories, and update the facial feature confidence model based on the original positive sample data and the expanded negative sample data, which helps to improve the adaptability and flexibility of the facial feature confidence model to different application scenarios.
[0086] Embodiment 3
[0087] Figure 6 This is a schematic diagram of the structure of a device for constructing a facial feature confidence model provided in the third embodiment of the present invention. The device can execute the method for constructing a facial feature confidence model provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method. Figure 6 As shown, the device comprises:
[0088] The sample face picture data set acquisition module 310 is used to acquire a sample face picture data set; wherein the sample face picture data set is composed of sample face pictures marked with person identity tags;
[0089] The sample face picture group set determination module 320 is used to take every two sample face pictures in the sample face picture data set as a sample face picture group, calculate the first face feature similarity between the two sample face pictures included in each sample face picture group, and form a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold;
[0090] A negative sample data determination module 330 is used to determine a target sample face picture group marked with different person identity labels in the sample face picture group set, and use the sample face pictures included in the target sample face picture group as negative sample data;
[0091] A positive sample data determination module 340 is used to cluster the sample face pictures included in the sample face picture group set, determine a target cluster that meets a preset condition from the cluster clusters generated by clustering, and use the sample face pictures in the target cluster except the negative sample data as positive sample data;
[0092] The facial feature confidence model generation module 350 is used to form a training data set based on the negative sample data and the positive sample data, and train a preset machine learning model based on the training data set to generate a facial feature confidence model.
[0093] Optionally, the positive sample data determination module 340 includes:
[0094] A similarity relationship graph splitting unit is used to construct a similarity relationship graph based on the similarity of the first facial features corresponding to each sample face picture group in the sample face picture group set, and split the similarity relationship graph into clusters based on a clustering algorithm; wherein each cluster contains at least one sample face picture;
[0095] The target cluster determination unit is used to determine a target cluster that meets a preset condition from the clustering clusters.
[0096] Optionally, the target cluster determination unit is used to:
[0097] For each cluster, determine the target person identity label annotated on the sample face pictures included in the cluster;
[0098] When the target person identity tag is the same person identity tag, and all sample face images in the sample face image data set marked with the target person identity tag are included in the cluster, the cluster is used as the target cluster.
[0099] Optionally, the negative sample data determination module 330 is further configured to:
[0100] Annotate each sample face picture in the target sample face picture group as having low confidence in face features, and use the annotated sample face pictures in the target sample face picture group as negative sample data;
[0101] The positive sample data determination module 340 is used to:
[0102] Deleting the sample face pictures in the target sample face picture group from the target cluster to generate a target sample face picture set;
[0103] Each sample face picture in the target sample face picture set is annotated as having a high confidence level of a face feature, and the annotated sample face pictures in the target sample face picture set are used as positive sample data.
[0104] Optionally, the device further comprises:
[0105] A spatiotemporal trajectory face image acquisition module is used to acquire at least two spatiotemporal trajectory face images after generating a face feature confidence model; wherein the spatiotemporal trajectory face images are face images whose collection time interval is less than a preset time length and whose distance between collection positions is greater than a preset distance threshold;
[0106] A facial feature confidence model updating module is used to add each two spatiotemporal trajectory face images of the at least two spatiotemporal trajectory face images to the negative sample data to update the facial feature confidence model when the facial feature similarity between the two spatiotemporal trajectory face images is greater than the preset similarity threshold.
[0107] Optionally, the device further comprises:
[0108] A first target face image acquisition module, used to acquire a first target face image after generating a face feature confidence model;
[0109] a second target face picture determination module, configured to obtain a first face feature confidence of the first target face picture based on the face feature confidence model, and use the first target face picture whose first face feature confidence is greater than a preset confidence threshold as the second target face picture;
[0110] The first facial feature extraction module is used to extract the first facial feature of the second target facial image based on the facial recognition model to be evaluated.
[0111] Optionally, the device further comprises:
[0112] A third target face image acquisition module, used to acquire at least two third target face images after generating the face feature confidence model;
[0113] A second facial feature confidence acquisition module, used to respectively acquire the second facial feature confidences of the at least two third target facial images based on the facial feature confidence model;
[0114] A second facial feature similarity adjustment module is used to adjust the second facial feature similarity between each two third target facial images of the at least two third target facial images based on the second facial feature confidences corresponding to each two third target facial images;
[0115] The third target face picture clustering module is used to cluster the at least two third target face pictures based on the adjusted second face feature similarity.
[0116] A facial feature confidence model construction device provided in an embodiment of the present invention can execute a facial feature confidence model construction method provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0117] Embodiment 4
[0118] Figure 7 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0119] like Figure 7 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0120] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0121] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a method for constructing a facial feature confidence model.
[0122] In some embodiments, the method for constructing a facial feature confidence model may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the method for constructing a facial feature confidence model described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the method for constructing a facial feature confidence model in any other appropriate manner (e.g., by means of firmware).
[0123] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0124] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0125] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0126] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0127] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0128] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0129] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0130] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for constructing a facial feature confidence model, characterized in that: The method comprises: Obtain a sample face image data set; wherein the sample face image data set is composed of sample face images marked with person identity tags; Taking every two sample face pictures in the sample face picture data set as a sample face picture group, calculating the first face feature similarity between the two sample face pictures included in each sample face picture group, and forming a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold; Determine a target sample face picture group marked with different person identity labels in the sample face picture group set, and use the sample face pictures included in the target sample face picture group as negative sample data; Clustering the sample face pictures contained in the sample face picture group set, determining a target cluster that meets a preset condition from the cluster clusters generated by clustering, and using the sample face pictures in the target cluster except the negative sample data as positive sample data; A training data set is formed based on the negative sample data and the positive sample data, and a preset machine learning model is trained based on the training data set to generate a facial feature confidence model.
2. The method according to claim 1, characterized in that: Clustering the sample face pictures included in the sample face picture group set, and determining a target cluster that meets a preset condition from the cluster clusters generated by clustering, including: Constructing a similarity relationship graph based on the similarity of the first facial features corresponding to each sample facial picture group in the sample facial picture group set, and splitting the similarity relationship graph into clusters based on a clustering algorithm; wherein each cluster contains at least one sample facial picture; A target cluster that meets a preset condition is determined from the clustering clusters.
3. The method according to claim 2, characterized in that Determining a target cluster that meets a preset condition from the clustering clusters includes: For each cluster, determine the target person identity label annotated on the sample face pictures included in the cluster; When the target person identity tag is the same person identity tag, and all sample face images in the sample face image data set marked with the target person identity tag are included in the cluster, the cluster is used as the target cluster.
4. The method according to claim 1, characterized in that The sample face pictures included in the target sample face picture group are used as negative sample data, including: Annotate each sample face picture in the target sample face picture group as having low confidence in face features, and use the annotated sample face pictures in the target sample face picture group as negative sample data; The sample face pictures in the target cluster except the negative sample data are used as positive sample data, including: Deleting the sample face pictures in the target sample face picture group from the target cluster to generate a target sample face picture set; Each sample face picture in the target sample face picture set is annotated as having a high confidence level of a face feature, and the annotated sample face pictures in the target sample face picture set are used as positive sample data.
5. The method according to claim 1, characterized in that After generating the facial feature confidence model, it also includes: Acquire at least two spatiotemporal trajectory face images; wherein the spatiotemporal trajectory face images are face images whose collection time interval is less than a preset time length and whose collection positions are separated by a distance greater than a preset distance threshold; For each two of the at least two spatiotemporal trajectory face pictures, when the facial feature similarity between the two spatiotemporal trajectory face pictures is greater than the preset similarity threshold, the two spatiotemporal trajectory face pictures are added to the negative sample data to update the facial feature confidence model.
6. The method according to claim 1, characterized in that After generating the facial feature confidence model, it also includes: Get the first target face image; Acquire a first facial feature confidence of the first target facial image based on the facial feature confidence model, and use the first target facial image whose first facial feature confidence is greater than a preset confidence threshold as the second target facial image; Extracting the first facial feature of the second target facial image based on the facial recognition model to be evaluated.
7. The method according to claim 1, characterized in that After generating the facial feature confidence model, it also includes: Obtain at least two third target face images; Based on the facial feature confidence model, respectively obtain the second facial feature confidences of the at least two third target facial images; For each two third target face pictures of the at least two third target face pictures, adjusting the second facial feature similarity between the each two third target face pictures based on the second facial feature confidences corresponding to the each two third target face pictures; The at least two third target face images are clustered based on the adjusted second facial feature similarity.
8. A device for constructing a facial feature confidence model, characterized in that: The device comprises: A sample face picture data set acquisition module is used to acquire a sample face picture data set; wherein the sample face picture data set is composed of sample face pictures marked with person identity tags; a sample face picture group set determination module, configured to take every two sample face pictures in the sample face picture data set as a sample face picture group, calculate the first face feature similarity between the two sample face pictures included in each sample face picture group, and form a sample face picture group set based on the sample face picture groups whose first face feature similarity is greater than a preset similarity threshold; A negative sample data determination module is used to determine a target sample face picture group marked with different person identity labels in the sample face picture group set, and use the sample face pictures included in the target sample face picture group as negative sample data; A positive sample data determination module is used to cluster the sample face pictures contained in the sample face picture group set, determine a target cluster that meets a preset condition from the cluster clusters generated by clustering, and use the sample face pictures in the target cluster except the negative sample data as positive sample data; The facial feature confidence model generation module is used to form a training data set based on the negative sample data and the positive sample data, and to train a preset machine learning model based on the training data set to generate a facial feature confidence model.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for constructing a facial feature confidence model described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for constructing a facial feature confidence model described in any one of claims 1 to 7 when executed.