A method and apparatus for archiving face pictures, and a computer-readable storage medium

By using preset clusters and twin networks to match face images in the intelligent video surveillance system, the problem of inefficient clustering of massive portrait images is solved, and more accurate clustering and improved video surveillance efficiency are achieved.

CN114969412BActive Publication Date: 2025-06-10ZHEJIANG DAHUA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210489733.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2025-06-10
Estimated Expiration
2042-05-06

AI Technical Summary

Technical Problem

With the popularization of intelligent video surveillance equipment, the accumulation of massive portrait images has led to a continuous growth of face image libraries. The amount of re-clustering of all samples is large and the clustering time is long, which affects the efficiency of video surveillance.

Method used

Using preset multiple class clusters, at least one representative picture is extracted from each class cluster, and the representative picture is matched with the target face picture using a pre-trained twin network to determine whether the matching value is greater than or equal to the preset threshold. If so, archive it to the class cluster with the largest matching value, and if otherwise, a new class cluster is created.

Benefits of technology

The matching degree of face pictures output through the twin network is obtained, and the characteristics of higher dimensions are obtained, making the clustering results more accurate, reducing clustering time and improving video surveillance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969412B_ABST
    Figure CN114969412B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for archiving face pictures, as well as a computer-readable storage medium. The method for archiving face pictures includes: obtaining a plurality of preset clusters, and extracting at least one representative picture from each cluster; respectively inputting the representative pictures of each cluster and a target face picture into a pre-trained siamese network to obtain the matching values of each cluster and the target face picture; determining whether the matching values of the plurality of clusters are greater than or equal to a first preset threshold; if so, archiving the target face picture into the cluster with the largest matching value; if not, creating a new cluster based on the target face picture. Through the above method, the face picture archiving apparatus can utilize the siamese network to output the matching degree of face pictures, can obtain higher-dimensional features of the pictures, and make the clustering results more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent video surveillance, and particularly to a method and apparatus for archiving face images, as well as a computer-readable storage medium. Background Art

[0002] With the widespread popularity of intelligent video surveillance devices, a huge amount of portrait images are accumulated every day. In the public security system, it is often possible to find photos belonging to the same person from a large number of photos, and then combine the spatio-temporal information of these photos to complete the trajectory tracking of a certain person. Therefore, how to cluster and archive photos belonging to the same person from a large number of photos is a very important task.

[0003] Since the video surveillance device captures images at any time, the number of the portrait image library increases with time. Whenever a new image is stored in the library, a clustering process needs to be performed. With the rapid increase in the number of images, the computational complexity of re-clustering all samples becomes larger and larger, and the clustering time is relatively long, resulting in low efficiency of video surveillance. Summary of the Invention

[0004] This application provides a method and apparatus for archiving face images, as well as a computer-readable storage medium.

[0005] This application provides a method for archiving face images, and the method for archiving face images includes:

[0006] Obtain a plurality of preset clusters, and extract at least one representative image from each cluster;

[0007] Input the representative image of each cluster and the target face image into a pre-trained siamese network respectively to obtain the matching value between each cluster and the target face image;

[0008] Determine whether the matching values of the multiple clusters are greater than or equal to a first preset threshold;

[0009] If so, archive the target face image into the cluster with the largest matching value;

[0010] If not, establish a new cluster based on the target face image.

[0011] Wherein, the extracting at least one representative image from each cluster includes:

[0012] Obtain the number of images in each cluster;

[0013] When the number of images in the cluster is 1, use the image in the cluster as the representative image;

[0014] When the number of pictures in the cluster is greater than 1, two pictures with the smallest time and / or space gap from the target face picture in the cluster are taken as the representative pictures.

[0015] Among them, taking two pictures with the smallest time and / or space gap from the target face picture in the cluster as the representative pictures includes:

[0016] Select a picture in the cluster that is closest to the target face picture in terms of acquisition time as the first representative picture;

[0017] Select a picture in the cluster that is closest to the target face picture in terms of acquisition space as the second representative picture.

[0018] Among them, taking two pictures with the smallest time and / or space gap from the target face picture in the cluster as the representative pictures includes:

[0019] Select two pictures collected by the same acquisition device as the target face picture in the cluster as the first representative picture and the second representative picture.

[0020] Among them, when the number of pictures in the cluster is greater than 1, the representative pictures include the first representative picture and the second representative picture;

[0021] Respectively inputting the representative pictures of each cluster and the target face picture into a pre-trained siamese network to obtain the matching value between each cluster and the target face picture includes:

[0022] Input the first representative picture and the target face picture into the siamese network to obtain the first matching value between the first representative picture and the target face picture;

[0023] Input the second representative picture and the target face picture into the siamese network to obtain the second matching value between the second representative picture and the target face picture;

[0024] Judge whether the first matching value and the second matching value are both greater than or less than a second preset threshold;

[0025] If so, take the average value of the first matching value and the second matching value as the matching value between the cluster and the target face picture.

[0026] Among them, the face picture filing method further includes:

[0027] In the case where one of the first matching value and the second matching value is greater than the second preset threshold and the other is less than the first preset threshold, input the first representative picture and the second representative picture into the siamese network to obtain a third matching value between the first representative picture and the second representative picture;

[0028] Determine whether the third matching value is greater than or equal to a third preset threshold;

[0029] If so, use the average value of the first matching value and the second matching value as the matching value between the cluster and the target face picture;

[0030] If not, take out the first representative picture and the second representative picture from the cluster and re - file them.

[0031] Wherein, the face picture filing method further includes:

[0032] Obtain a plurality of preset training clusters;

[0033] Select two training pictures from one training cluster among the plurality of training clusters to form a positive sample;

[0034] Select one training picture from each of two training clusters among the plurality of training clusters to form a negative sample;

[0035] Use the positive sample and / or the negative sample to form a training set, and use the training set to train the siamese network.

[0036] This application also provides a face picture filing device, which includes an acquisition module, a matching module, and a filing module; wherein,

[0037] The acquisition module is used to obtain a plurality of preset clusters and extract at least one representative picture from each cluster;

[0038] The matching module is used to input the representative picture of each cluster and the target face picture into a pre - trained siamese network respectively to obtain the matching value between each cluster and the target face picture;

[0039] The filing module is used to determine whether the matching values of the plurality of clusters are greater than or equal to a first preset threshold; if so, file the target face picture into the cluster with the largest matching value; if not, establish a new cluster based on the target face picture.

[0040] This application also provides another face picture filing device, which includes a processor and a memory. Program data is stored in the memory, and the processor is used to execute the program data to implement the face picture filing method as described above.

[0041] The present application also provides a computer-readable storage medium for storing program data, which, when executed by a processor, is used to implement the above-mentioned face picture filing method.

[0042] The beneficial effects of the present application are as follows: The face picture filing device obtains a plurality of preset clusters, and extracts at least one representative picture from each cluster; respectively inputs the representative pictures of each cluster and the target face picture into a pre-trained siamese network to obtain the matching values of each cluster and the target face picture; determines whether the matching values of the multiple clusters are greater than or equal to a first preset threshold; if so, files the target face picture into the cluster with the largest matching value; if not, creates a new cluster based on the target face picture. Through the above method, the face picture filing device can use the siamese network to output the matching degree of face pictures, can obtain higher-dimensional features of the pictures, and make the clustering results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0044] Figure 1 is a schematic flowchart of an embodiment of the face picture filing method provided by the present application;

[0045] Figure 2 is Figure 1 a specific flowchart of the face picture filing method shown;

[0046] Figure 3 is Figure 1 a specific flowchart of step S12 of the face picture filing method shown;

[0047] Figure 4 is a schematic flowchart of another embodiment of the face picture filing method provided by the present application;

[0048] Figure 5 is Figure 4 a specific flowchart of the face picture filing method shown;

[0049] Figure 6 is a schematic framework diagram of an embodiment of the siamese network provided by the present application;

[0050] Figure 7 is a schematic structural diagram of an embodiment of the face picture filing device provided by the present application;

[0051] Figure 8 is a schematic structural diagram of another embodiment of the face picture archiving device provided by this application;

[0052] Figure 9 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by this application. Specific embodiments

[0053] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0054] Please refer to Figure 1 and Figure 2 , Figure 1 is a schematic flowchart of an embodiment of the face picture archiving method provided by this application, Figure 2 is Figure 1 the specific flowchart of the face picture archiving method shown.

[0055] Among them, the face picture archiving method of this application is applied to a face picture archiving device. Among them, the face picture archiving device of this application can be a server or a system in which the server and the terminal device cooperate with each other. Correspondingly, each part included in the face picture archiving device, such as each unit, subunit, module, and submodule, can be all set in the server or can be respectively set in the server and the terminal device.

[0056] Furthermore, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or can be implemented as a single software or software module, which is not specifically limited herein. In some possible implementation manners, the face picture archiving method of the embodiments of this application can be implemented by a processor calling computer-readable instructions stored in a memory.

[0057] Specifically, as Figure 1 shown, the face picture archiving method of the embodiments of this application specifically includes the following steps:

[0058] Step S11: Obtain a preset plurality of clusters, and extract at least one representative picture from each cluster.

[0059] In the embodiment of the present application, the face image filing device obtains a plurality of formed clusters, and each cluster represents all face images of a real person. Among them, the preset plurality of clusters can be obtained by manually dividing a large number of face images, and clustering the face images belonging to the same real person into one cluster.

[0060] In the embodiment of the present application, in order to avoid selecting all face images to match with the target face image to be filed, the face image filing device selects at least one representative image from each cluster, and uses the matching value between the representative image and the target face image as the matching value between the cluster where the representative image is located and the target face image, thereby reducing the number of matches and improving the matching efficiency.

[0061] On the premise of sufficient processing resources, the more representative images are selected from each cluster, the more accurate the matching value between each cluster and the target face image is. Therefore, the face image filing device selects a corresponding number of representative images to match with the target face image according to the number of images in each cluster.

[0062] Specifically, if the number of images in a cluster is 1, then the cluster can only use this image as the representative image, that is, the cluster only selects one representative image. In the subsequent matching process, directly use the matching value between this representative image and the target face image as the matching value between this cluster and the target face image.

[0063] If the number of images in a cluster is greater than 1, then the cluster can select two or more images as representative images. The embodiment of the present application takes the way of selecting two representative images from each cluster as a specific example: In order to further improve the relevance between the representative images selected from each cluster and the target face image, the face image filing device can select the images in each cluster that are closest to the target face image in the time dimension and / or space dimension as the representative images. Among them, the relevance in the time dimension can be the time gap between the acquisitions of two images, and the relevance in the space dimension can be the acquisition position gap between two images. In the embodiment of the present application, the position coordinates of the acquisition device for acquiring images, such as the position coordinates of a bayonet camera, can be used as the acquisition position of the images. Through the combination of the time dimension and the space dimension, the ways of selecting representative images in the embodiment of the present application may include but are not limited to the following three cases:

[0064] First, select the two images in the cluster that are closest to the acquisition time of the target face image as the first representative image and the second representative image respectively.

[0065] Second, select the two images in the cluster that are closest to the acquisition position of the target face image as the first representative image and the second representative image respectively. For example, select the images acquired by the same bayonet camera as the target face image in the cluster as the representative images.

[0066] Third, select a picture in the cluster that is closest in acquisition time to the target face picture as the first representative picture, and select a picture that is closest in acquisition location to the target face picture as the second representative picture.

[0067] For example, assume that the face picture archiving device has stored K clusters, and each cluster has N i pictures, where 1 ≤ i ≤ K, and N i takes an integer value not less than 1. If N i = 1, that is, there is only one picture in this cluster, which is I i . If N i ≥ 2, select two face pictures from this cluster as the first representative picture and the second representative picture respectively. Among them, the first representative picture is the picture that is closest in acquisition time to the target face picture, that is The second representative picture is the picture that is closest in acquisition location to the target face picture, that is

[0068] Step S12: Input the representative picture of each cluster and the target face picture into a pre-trained siamese network respectively to obtain the matching value between each cluster and the target face picture.

[0069] In the embodiment of the present application, the face picture archiving device inputs the representative picture of each cluster and the target face picture into a pre-trained siamese network, and the siamese network predicts the matching value between the target face picture and the representative picture.

[0070] Among them, the siamese network is a deep learning model suitable for solving matching problems. The siamese network takes two samples as inputs and outputs their representations embedded in a high-dimensional space for comparing the similarity of the two samples. Generally speaking, the two neural networks of the siamese network have exactly the same parameters. If designed as two different neural networks, it is called a pseudo-siamese network.

[0071] For the cluster with N i = 1, since there is only one representative picture, the face picture archiving device inputs this representative picture and the target face picture into the siamese network to obtain the predicted matching value between this representative picture and the target face picture, which can be used as the matching value between this cluster and the target face picture.

[0072] For the cluster with N i ≥ 2, the face picture archiving device will select the first representative picture and the second representative picture Match with the target face image respectively, and comprehensively obtain the matching value between this cluster and the target face image based on the matching values of the two representative images and the target face image. The calculation method can directly calculate the average value of the matching values of the two representative images as the matching value of the cluster, or use the minimum or maximum matching value of the two representative images as the matching value of the cluster.

[0073] In the embodiments of the present application, Figure 3 The shown embodiment provides a specific method for calculating the matching value of a cluster, where Figure 3 is Figure 1 a specific process schematic diagram of step S12 of the face image filing method shown.

[0074] Please continue to refer to Figure 2 and Figure 3 , the face image filing method in the embodiments of the present application specifically further includes the following steps:

[0075] Step S121: Input the first representative image and the target face image into the siamese network to obtain the first matching value between the first representative image and the target face image.

[0076] Obtain the first representative image through step S121 and the first matching value S of the target face image GT .

[0077] Step S122: Input the second representative image and the target face image into the siamese network to obtain the second matching value between the second representative image and the target face image.

[0078] Obtain the second representative image through step S122 and the second matching value S of the target face image GS .

[0079] Step S123: Determine whether the first matching value and the second matching value are both greater than or less than the second preset threshold.

[0080] In the embodiments of the present application, as Figure 2 shown, when the first matching value S GT and the second matching value S GS are both greater than or less than the second preset threshold, then enter step S124.

[0081] When the first matching value S GT and the second matching value S GSWhen one of them is greater than the second preset threshold and the other is less than the second preset threshold, it indicates that there is a situation where the representative pictures selected from the same cluster do not belong to the same real person, that is, a misfiling situation. Among them, there may be multiple face acquisition pictures belonging to the same person in a certain cluster formed after face clustering. If there are face pictures belonging to other real people among them, it is called misfiling.

[0082] At this time, the face picture filing device needs to perform misfiling correction on these two representative pictures, that is, when filing the target face picture, dynamically remove the misfiled pictures under this cluster and put them into the set to be filed, which is called misfiling correction. For details, please continue to refer to the following steps:

[0083] The face picture filing device obtains the first representative picture through the siamese network and the second representative picture of the matching value, and then judges whether the matching value of the first representative picture and the second representative picture is greater than or equal to the third preset threshold. If so, it means that there is no misfiling situation in this cluster, and the average value of the first matching value S GT and the second matching value S GS can be used as the matching value of this cluster and the target face picture. If not, it means that there is a misfiling situation in this cluster, and the first representative picture and the second representative picture are removed from this cluster, and together with the target face picture, they are added to the set of pictures to be filed for re-filing.

[0084] Step S124: Take the average value of the first matching value and the second matching value as the matching value of the cluster and the target face picture.

[0085] In the embodiment of the present application, the face picture filing device takes the average value of the first matching value S GT and the second matching value S GS as the matching value of the target face picture and the cluster where the representative picture is located. Among them, the matching value reflects the similarity between the target face picture and the cluster. The higher the matching value, the higher the possibility that the target face picture and the face pictures in the cluster belong to the same real person.

[0086] Step S13: Judge whether the matching values of multiple clusters are greater than or equal to the first preset threshold.

[0087] In the embodiment of the present application, the face picture filing device judges whether the matching value of each cluster is greater than or equal to the first preset threshold. If so, the cluster is divided into the candidate cluster and step S14 is entered; if not, the cluster is not divided into the candidate cluster and step S15 is entered.

[0088] Step S14: Archive the target face image to the cluster with the largest matching value.

[0089] In an embodiment of the present application, after the face image archiving device completes the determination of the matching values of all clusters, it checks whether the candidate cluster is empty. If not, it uses the cluster with the largest matching value in the candidate cluster as the destination cluster for the target face image.

[0090] Step S15: Create a new cluster based on the target face image.

[0091] In an embodiment of the present application, after the face image archiving device completes the determination of the matching values of all clusters, it checks whether the candidate cluster is empty. If so, the target face image alone forms a new cluster.

[0092] In an embodiment of the present application, the face image archiving device obtains a plurality of preset clusters, extracts at least one representative image from each cluster; respectively inputs the representative image of each cluster and the target face image into a pre-trained siamese network to obtain the matching value between each cluster and the target face image; determines whether the matching values of the plurality of clusters are greater than or equal to a first preset threshold; if so, archives the target face image to the cluster with the largest matching value; if not, creates a new cluster based on the target face image. In the above manner, the face image archiving device can utilize the siamese network to output the matching degree of face images, can obtain higher-dimensional features of the images, and make the clustering result more accurate.

[0093] Next, continue to introduce Figure 1 and Figure 2 the training process of the siamese network shown. Please continue to refer to Figure 4 and Figure 5 , Figure 4 is a schematic flowchart of another embodiment of the face image archiving method provided by the present application, Figure 5 is Figure 4 the specific flowchart of the face image archiving method shown.

[0094] Specifically, as Figure 4 shown, the face image archiving method of the embodiment of the present application specifically includes the following steps:

[0095] Step S21: Obtain a plurality of preset training clusters.

[0096] In an embodiment of the present application, the face image archiving device obtains a plurality of training clusters input by the staff and / or obtains a plurality of training clusters provided by the network training set. Each training cluster includes at least one face image, and each training cluster represents a set of face images of a real person.

[0097] During the training process of the Siamese network in the embodiments of this application, semi-supervised learning is adopted. When training the Siamese network, some paired labeled samples are required. If the two input pictures are from the same person, it is a positive sample; if the two input pictures are of different people, it is a negative sample.

[0098] Since it is too costly to manually label all the collected face images, semi-supervised learning methods can be used. First, manually distinguish a part of the face images and label this part of the samples. If the number of samples in a certain cluster is too small, data augmentation for training can be performed through operations such as image flipping, cropping, changing brightness and chromaticity. In addition, outside the dataset of this design, some public datasets can be added as the training set to increase the robustness of the model.

[0099] Step S22: Select two training pictures from one training cluster among multiple training clusters to form a positive sample.

[0100] Step S23: Select one training picture from each of two training clusters among multiple training clusters to form a negative sample.

[0101] Step S24: Use the positive samples and / or negative samples to form a training set, and use the training set to train the Siamese network.

[0102] In the embodiments of this application, the face picture archiving device constructs a corresponding training set according to a preset multiple training clusters. The specific process is as follows: Extract one face picture from any two clusters to form a negative sample, that is, the face pictures of the negative sample come from different real people; extract two face pictures from the same cluster to form a positive sample, that is, the face pictures of the positive sample come from the same real person. Thus, a large amount of training data can be constructed and input into the Siamese network for training to learn the ability to distinguish whether two inputs are similar.

[0103] The face picture archiving device inputs the training set into the Siamese network. As Figure 6 shown, the face picture archiving device respectively passes the two inputs (input1, input2) of a pair of samples into the two neural networks (network1, network2) of the Siamese network. These two neural networks share parameters. They map the inputs to a new space to form the representations of the inputs in the new space. The training process updates the weight matrix parameters of the two neural networks through the backpropagation of the loss function. In this process, the Siamese network learns the ability to distinguish the similarity of two pictures and outputs a value representing the matching degree of the two inputs. Among them, the loss function is:

[0104]

[0105] Among them, m is a set threshold value, Y is the label of a pair of samples, Y = 1 indicates that the two input pictures are similar or matched, and Y = 0 indicates that the two input pictures do not match. Indicates the Euclidean distance of the features X of two input pictures 1 , X 2 , and P represents the number of feature bits of the sample.

[0106] Specifically, when Y = 1, that is, when the pair of samples is similar, the above loss function can be simplified as:

[0107]

[0108] At this time, for the originally similar sample pairs, if the Euclidean distance in the feature space is large, it means that the current Siamese network is not good, so the loss is increased.

[0109] When Y = 0, that is, when the pair of samples is not similar, the above loss function can be simplified as:

[0110]

[0111] At this time, the smaller the Euclidean distance in its feature space, the greater the loss, which also meets the requirements of the face image filing method of this application.

[0112] Figure 6 The W shown in is the shared weight matrix parameter. After each iterative training, the Siamese network calculates the distance between the weight matrix parameters of the two neural networks, and then calculates a new shared weight matrix parameter Ew.

[0113] Such as Figure 5 shown, after the training of the Siamese network is completed, the filing problem of the target face image can be transformed into the prediction problem of the Siamese network, that is, using the trained Siamese network to implement Figure 1 and Figure 2 shown face image filing method.

[0114] In the embodiments of this application, the loss function design of the Siamese network can well express the matching degree of paired samples and can also be well used to train the model for feature extraction. Therefore, using the Siamese network to match two face pictures can obtain higher-dimensional features of the pictures and make the clustering results more accurate; when clustering and filing the target pictures, at the same time, the representative pictures in the formed clusters are matched and verified. Every time a new picture is put into the library, a dynamic adjustment and verification are carried out. Therefore, the process completes the misfiling correction and gradually optimizes the face clustering results; based on semi-supervised learning, that is, only a small amount of data needs to be manually labeled and then automatically expanded; and adding a public face dataset as the training samples of the Siamese network, so a large number of paired samples can be formed without a large amount of manual labeling, and the robustness of the model can be enhanced.

[0115] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0116] To implement the face image archiving method of the above embodiments, the present application also proposes a face image archiving device. For details, please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an embodiment of the face image archiving device provided by the present application.

[0117] The face image archiving device 300 of the embodiments of the present application includes an acquisition module 31, a matching module 32, and an archiving module 33.

[0118] Among them, the acquisition module 31 is used to acquire a plurality of preset clusters and extract at least one representative image from each cluster.

[0119] The matching module 32 is used to respectively input the representative images of each cluster and the target face image into a pre-trained siamese network to obtain the matching value of each cluster and the target face image.

[0120] The archiving module 33 is used to determine whether the matching values of the plurality of clusters are greater than or equal to a first preset threshold; if so, archive the target face image to the cluster with the largest matching value; if not, establish a new cluster based on the target face image.

[0121] To implement the face image archiving method of the above embodiments, the present application also proposes a face image archiving device. For details, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of another embodiment of the face image archiving device provided by the present application.

[0122] The face image archiving device 400 of the embodiments of the present application includes a memory 41 and a processor 42. Among them, the memory 41 and the processor 42 are coupled.

[0123] The memory 41 is used to store program data, and the processor 42 is used to execute the program data to implement the face image archiving method described in the above embodiments.

[0124] In this embodiment, the processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 42 can also be any conventional processor, etc.

[0125] To implement the face picture archiving method of the above embodiment, the present application also provides a computer-readable storage medium, such as Figure 9 shown, the computer-readable storage medium 500 is used to store program data 51. When the program data 51 is executed by the processor, it is used to implement the face picture archiving method as described in the above embodiment.

[0126] The present application also provides a computer program product. Among them, the above computer program product includes a computer program, and the above computer program can be operated to make a computer execute the face picture archiving method as described in the embodiments of the present application. The computer program product can be a software installation package.

[0127] When the face picture archiving method described in the above embodiments of the present application exists in the form of a software functional unit and is sold or used as an independent product, it can be stored in a device, such as a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0128] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.

Claims

1. A method for archiving face pictures, characterized in that, the method for archiving face pictures includes: Obtaining a plurality of preset clusters, and extracting at least one representative picture from each cluster; Respectively inputting the representative pictures of each cluster and the target face picture into a pre-trained siamese network to obtain the matching values of each cluster and the target face picture; Judging whether the matching values of the multiple clusters are greater than or equal to a first preset threshold; If so, archiving the target face picture into the cluster with the largest matching value; If not, establishing a new cluster based on the target face picture; When the number of pictures in the cluster is greater than 1, the representative pictures include a first representative picture and a second representative picture; Inputting the first representative picture and the target face picture into the siamese network to obtain the first matching value of the first representative picture and the target face picture; inputting the second representative picture and the target face picture into the siamese network to obtain the second matching value of the second representative picture and the target face picture; Judging whether the first matching value and the second matching value are both greater than or less than a second preset threshold; When one of the first matching value and the second matching value is greater than the second preset threshold and the other is less than the first preset threshold, inputting the first representative picture and the second representative picture into the siamese network to obtain the third matching value of the first representative picture and the second representative picture; Judging whether the third matching value is greater than or equal to a third preset threshold; If so, taking the average value of the first matching value and the second matching value as the matching value of the cluster and the target face picture; If not, taking out the first representative picture and the second representative picture from the cluster and re-archiving them.

2. The method for archiving face pictures according to claim 1, characterized in that, the extracting at least one representative picture from each cluster includes: Obtaining the number of pictures in each cluster; When the number of pictures in the cluster is 1, taking the picture in the cluster as the representative picture; When the number of pictures in the cluster is greater than 1, taking the two pictures in the cluster with the smallest time and / or space gap from the target face picture as the representative pictures.

3. The method for archiving face pictures according to claim 2, characterized in that, the taking the two pictures in the cluster with the smallest time and / or space gap from the target face picture as the representative pictures includes: Selecting a picture in the cluster with the closest acquisition time to the target face picture as the first representative picture; Selecting a picture in the cluster with the closest acquisition space to the target face picture as the second representative picture.

4. The method for archiving face pictures according to claim 2, characterized in that, the taking the two pictures in the cluster with the smallest time and / or space gap from the target face picture as the representative pictures includes: Selecting two pictures in the cluster collected by the same acquisition device as the first representative picture and the second representative picture.

5. The face image filing method according to any one of claims 2 to 4, characterized in that when the number of images in the cluster is greater than 1, the representative image includes a first representative image and a second representative image; the step of respectively inputting the representative image of each cluster and the target face image into a pre-trained siamese network to obtain the matching value between each cluster and the target face image includes: inputting the first representative image and the target face image into the siamese network to obtain a first matching value between the first representative image and the target face image; inputting the second representative image and the target face image into the siamese network to obtain a second matching value between the second representative image and the target face image; judging whether the first matching value and the second matching value are both greater than or less than a second preset threshold; if so, taking the average value of the first matching value and the second matching value as the matching value between the cluster and the target face image.

6. The face image filing method according to claim 1, characterized in that the face image filing method further includes: obtaining a plurality of preset training clusters; selecting two training images from one training cluster among the plurality of training clusters to form a positive sample; selecting one training image from each of two training clusters among the plurality of training clusters to form a negative sample; forming a training set from the positive sample and / or the negative sample, and using the training set to train the siamese network.

7. A face image filing device, characterized in that the face image filing device includes an acquisition module, a matching module and a filing module; wherein, the acquisition module is configured to obtain a plurality of preset clusters and extract at least one representative image from each cluster; the matching module is configured to respectively input the representative image of each cluster and the target face image into a pre-trained siamese network to obtain the matching value between each cluster and the target face image; the filing module is configured to judge whether the matching values of the plurality of clusters are greater than or equal to a first preset threshold; if so, filing the target face image into the cluster with the largest matching value; if not, establishing a new cluster based on the target face image; the matching module is further configured to, when the number of images in the cluster is greater than 1, the representative image includes a first representative image and a second representative image; inputting the first representative image and the target face image into the siamese network to obtain a first matching value between the first representative image and the target face image; inputting the second representative image and the target face image into the siamese network to obtain a second matching value between the second representative image and the target face image; The filing module is further configured to determine whether the first matching value and the second matching value are both greater than or less than a second preset threshold; in the case where one of the first matching value and the second matching value is greater than the second preset threshold and the other is less than the first preset threshold, input the first representative picture and the second representative picture into the siamese network to obtain a third matching value between the first representative picture and the second representative picture; determine whether the third matching value is greater than or equal to a third preset threshold; if so, use the average value of the first matching value and the second matching value as the matching value between the cluster and the target face picture; if not, take out the first representative picture and the second representative picture from the cluster and re-file them.

8. A face picture filing device, characterized in that, the face picture filing device includes a processor and a memory, the memory stores program data, and the processor is configured to execute the program data to implement the face picture filing method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, the computer-readable storage medium is used to store program data, and when the program data is executed by a processor, it is used to implement the face picture filing method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Archiving method and device

    CN109783672A