Face Retrieval Method, Device and Storage Medium

Through the combination of the deep learning feature extraction model and the Faiss retrieval algorithm, the problem of missing information dimensions in face retrieval by traditional methods and deep learning methods is solved, and fast and accurate face retrieval is achieved.

CN113392672BActive Publication Date: 2025-07-08BEIJING ZHONGKE JINDEZHU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010167991.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-11
Publication Date
2025-07-08
Estimated Expiration
2040-03-11

AI Technical Summary

Technical Problem

In the prior art, traditional facial feature extraction methods cannot obtain better facial features. Although deep learning methods can obtain rich facial features with high dimensions, dimension reduction or hash encoding is required to improve the search rate when establishing feature indexes, resulting in the lack of information dimensions and low accuracy.

Method used

A rich first feature vector is generated using a deep learning-based feature extraction model, and a Faiss search algorithm is used to directly establish a face database index to avoid dimensionality reduction or hash encoding, and search through cosine similarity.

Benefits of technology

Fast and accurate face retrieval is achieved, ensuring the fast retrieval speed and does not cause the missing information dimensions, greatly improving the search accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113392672B_ABST
    Figure CN113392672B_ABST
Patent Text Reader

Abstract

The present application discloses a face retrieval method, device, and storage medium. Among them, the method includes: obtaining a to-be-retrieved image containing the face of a target object; using a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and performing a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of face retrieval technology, and in particular, to a face retrieval method, device, and storage medium. Background Art

[0002] The rise of AI has accelerated the development of face recognition. The applications of face recognition have brought more and more convenience to people, such as face brushing payment, face brushing access control, etc. Face recognition can also be seen everywhere in the fields of security, finance, etc. While face recognition is developing, a large amount of face data has been accumulated. These massive face data also bring new challenges to face recognition. For a given face, how to quickly and accurately find the top N pictures with the highest similarity from hundreds of millions of face databases has become an urgent issue to be solved, which is the origin of the development of face retrieval.

[0003] Face retrieval can be roughly divided into two stages: the first stage is to extract the features of face images, and the second stage is to establish a feature index for these face features. In the first stage, existing methods include those based on traditional face feature extraction, such as using the Local Binary Pattern (LBP) operator to extract features from face images after wavelet transform; there are also methods based on deep learning for face feature extraction. In the second stage, most methods are to reduce the dimension of the face features extracted in the first stage or perform hash coding, and then save them to the feature library as the final index. However, the traditional way of extracting features cannot obtain good face features, so the final accuracy is not good. While the deep learning method can obtain relatively high-dimensional and rich face features, in the second stage, in order to improve the retrieval rate, reducing the dimension of the face features or performing hash coding often leads to the loss of information dimensions and also results in poor final accuracy.

[0004] Aiming at the technical problem in the above-mentioned existing technology that the traditional way of extracting features cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, it is necessary to reduce the dimension of the face features or perform hash coding, resulting in the loss of information dimensions and low accuracy, there is currently no effective solution. Summary of the Invention

[0005] Embodiments of the present disclosure provide a face retrieval method, device, and storage medium, so as to at least solve the technical problem in the existing technology that the traditional way of extracting features cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, it is necessary to reduce the dimension of the face features or perform hash coding, resulting in the loss of information dimensions and low accuracy.

[0006] According to one aspect of the embodiments of the present disclosure, a face retrieval method is provided, including: obtaining a to-be-retrieved image including a face of a target object; using a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and performing a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.

[0007] According to another aspect of the embodiments of the present disclosure, a storage medium is further provided. The storage medium includes a stored program, where, when the program runs, the method described in any one of the above is executed by a processor.

[0008] According to another aspect of the embodiments of the present disclosure, a face retrieval device is further provided, including: an obtaining module, configured to obtain a to-be-retrieved image including a face of a target object; a feature extraction module, configured to use a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; a determining module, configured to determine a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and a retrieval module, configured to perform a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.

[0009] According to another aspect of the embodiments of the present disclosure, a face retrieval device is further provided, including: a processor; and a memory, connected to the processor and configured to provide instructions for the processor to perform the following processing steps: obtaining a to-be-retrieved image including a face of a target object; using a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and performing a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.

[0010] In the embodiments of the present disclosure, in the feature extraction stage of face retrieval, a feature extraction model based on deep learning is adopted to extract face features from the image to be retrieved, obtaining a first feature vector with rich feature information, so that the accuracy of the retrieval result is relatively high. In the face retrieval stage, according to the first feature vector, a target feature vector is determined from multiple second feature vectors in a preset feature database. Since in the previous index establishment stage, in this embodiment, the second feature vectors are not dimensionally reduced or hash-coded, but the Faiss retrieval algorithm with very fast retrieval speed is directly used to determine the second feature vectors as the indexes of face images in the face database. Therefore, according to the target feature vector, retrieval can be performed in the face database to obtain a retrieval result corresponding to the image to be retrieved. Since the Faiss retrieval algorithm is directly used to determine the second feature vectors as the indexes of face images in the face database, it not only ensures fast retrieval speed, but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy. It achieves the technical effects of fast retrieval speed and high accuracy in face retrieval. Furthermore, it solves the technical problem in the prior art that the traditional feature extraction method cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, the face features need to be dimensionally reduced or hash-coded, resulting in the loss of information dimensions and low accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present disclosure and form a part of this application. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure. In the drawings:

[0012] Figure 1 is a hardware structure block diagram of a computing device for implementing the method according to Embodiment 1 of the present disclosure;

[0013] Figure 2 is a flowchart of the face retrieval method according to the first aspect of Embodiment 1 of the present disclosure;

[0014] Figure 3 is an overall flowchart of the face retrieval method according to the first aspect of Embodiment 1 of the present disclosure;

[0015] Figure 4 is a schematic diagram of the network structure of the feature extraction model according to the first aspect of Embodiment 1 of the present disclosure;

[0016] Figure 5 is a schematic diagram of the structure of the bottleneck unit according to the first aspect of Embodiment 1 of the present disclosure;

[0017] Figure 6 is a schematic diagram of the face retrieval device according to Embodiment 2 of the present disclosure; and

[0018] Figure 7 is a schematic diagram of the face retrieval device according to Embodiment 3 of the present disclosure. Detailed implementation manners

[0019] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] First, some nouns or terms that appear during the description of the embodiments of the present disclosure are applicable to the following explanations:

[0022] Face retrieval: For a given face image, the process of finding one or N faces that are most similar to it in a large face database is called face retrieval;

[0023] Faiss retrieval algorithm: The core algorithm in the Faiss library for clustering and similarity search open-sourced by the Facebook AI team. Faiss provides efficient similarity search and clustering for dense vectors, supports the search of vectors at the billion level, and is currently the most mature approximate nearest neighbor search library;

[0024] feature map: Feature map, which refers to the result in the deep learning convolution process;

[0025] GAP: Abbreviation for Global average Pooling, and its Chinese name is global average pooling; and

[0026] Bottleneck layer: It refers to the bottleneck layer in deep learning.

[0027] Embodiment 1

[0028] According to this embodiment, an embodiment of a face retrieval method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0029] The method embodiment provided in this embodiment can be executed in a mobile terminal, a computer terminal, a server, or a similar computing device. Figure 1 A hardware structure block diagram of a computing device for implementing the face retrieval method is shown. As Figure 1 shown, the computing device may include one or more processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory for storing data, and a transmission device for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computing device may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0030] It should be noted that the above one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computing device. As involved in the embodiments of the present disclosure, the data processing circuit is a processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0031] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the face retrieval method in the embodiments of the present disclosure. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the face retrieval method of the above application program. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories can be connected to the computing device through a network. Examples of the above networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0032] The transmission device is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computing device. In one instance, the transmission device includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0033] The display can be, for example, a touch-screen liquid crystal display (LCD), and the liquid crystal display enables a user to interact with the user interface of the computing device.

[0034] It should be noted here that, in some alternative embodiments, the above Figure 1 illustrated computing device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above computing device.

[0035] Under the above operating environment, according to the first aspect of this embodiment, a face retrieval method is provided. Figure 2 The flowchart of the method is shown. Refer to Figure 2 shown, the method includes:

[0036] S202: Obtain a to-be-retrieved image including a face of a target object;

[0037] S204: Use a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image;

[0038] S206: Determine a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and

[0039] S208: Retrieve in the face database according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images established in the face database according to the Faiss retrieval algorithm.

[0040] As described in the background art above, face retrieval can be roughly divided into two stages: the first stage is to extract the features of face images, and the second stage is to establish feature indexes for these face features. In the first stage, existing methods include traditional face feature extraction methods, such as using the Local Binary Pattern (LBP) operator to extract features from face images after wavelet transform; there are also deep learning-based methods for face feature extraction. In the second stage, most methods are to perform dimensionality reduction or hash coding on the face features extracted in the first stage, and then save them to the feature library as the final index. However, the traditional way of extracting features cannot obtain good face features, so the final accuracy is not good. The deep learning method can obtain relatively high-dimensional and rich face features, but in the second stage, in order to improve the retrieval speed, dimensionality reduction processing or hash coding of face features often leads to information dimensionality loss and also results in poor final accuracy.

[0041] To address the problems existing in the above background art, combined with Figure 2 as shown, this embodiment provides a face retrieval method with fast retrieval speed and high accuracy. Specifically, referring to Figure 2 as shown, in this embodiment, first, an image to be retrieved containing the face of a target object is obtained. Then, a feature extraction model is used to generate a first feature vector corresponding to the face in the image to be retrieved. Among them, the feature extraction model proposed in this embodiment is, for example but not limited to, the lightweight feature extraction network MobileFaceNet model, and is trained based on millions of pieces of training data. Therefore, the feature information of the first feature vector generated by using the feature extraction model is rich.

[0042] Further, in this embodiment, two databases with a one-to-one mapping relationship are involved. One is a face database for storing a large number of face images, and the other is a feature database for storing the face feature vectors corresponding to each face image in the face database. Specifically: First, a large number of face images are obtained from the public dataset MS-1M and stored in the face database. Then, the above-mentioned feature extraction model is used to extract features from each face image in the face database, obtaining a plurality of feature vectors (corresponding to the second feature vectors in the claims) and storing them in the feature database. Moreover, in the process of establishing a feature index for the plurality of second feature vectors in this embodiment, instead of using the existing methods of dimensionality reduction processing or hash coding, the open-source Faiss retrieval algorithm is adopted to establish an index for each of the second feature vectors, such that the plurality of second feature vectors are respectively the indexes of multiple face images in the face database.

[0043] Therefore, after the first feature vector is extracted using the feature extraction model in this embodiment, it is necessary to determine the target feature vector from the plurality of second feature vectors in the preset feature database (corresponding to the above-mentioned feature database) according to the first feature vector. Then, based on the characteristic that each of the second feature vectors is the index of each face image in the face database, this embodiment can perform a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved. For example, the face image with the highest similarity is retrieved from the face database and the result is returned.

[0044] Thus, in this way, in this embodiment, in the feature extraction stage of face retrieval, a feature extraction model based on deep learning is adopted to extract face features from the image to be retrieved, obtaining a first feature vector with rich feature information, so that the accuracy of the retrieval result is relatively high. In the face retrieval stage, according to the first feature vector, a target feature vector is determined from multiple second feature vectors in the preset feature database. Since in the previous index establishment stage, this embodiment does not perform dimensionality reduction or hashing encoding on the second feature vector, but directly determines the second feature vector as the index of the face image in the face database by using the Faiss retrieval algorithm with very fast retrieval speed. Therefore, the retrieval can be performed in the face database according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved. Since the Faiss retrieval algorithm is used to directly determine the second feature vector as the index of the face image in the face database, it not only ensures a fast retrieval speed but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy. The technical effects of fast retrieval speed and high accuracy in face retrieval are achieved. Furthermore, it solves the technical problem in the prior art that the traditional way of extracting features cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, it is necessary to perform dimensionality reduction processing or hashing encoding on the face features, resulting in the loss of information dimensions and low accuracy.

[0045] Optionally, it further includes establishing indexes of multiple face images in the face database in the following manner: using the feature extraction model to generate multiple second feature vectors respectively corresponding to the faces in the multiple face images; and according to the Faiss retrieval algorithm, respectively determining the multiple second feature vectors as the indexes of the multiple face images in the face database.

[0046] Specifically, in the process of establishing indexes of multiple face images in the face database, the above-mentioned feature extraction model can be used to generate multiple second feature vectors corresponding to the faces in the multiple face images respectively. In this embodiment, the feature extraction model used is, for example but not limited to, the lightweight feature extraction network MobileFaceNet model. Since the original MobileFaceNet model outputs a 128-dimensional feature vector, in order to obtain more rich-dimensional features in this embodiment, the output of the MobileFaceNet model is specifically changed to a 512-dimensional feature vector. Then, the Faiss retrieval algorithm is used to establish an index for it, and all 512-dimensional information is used as the feature index without dimensionality reduction or hash coding, and no other processing is performed. In this way, high-dimensional and rich second feature vectors can be obtained, and based on the fast retrieval performance of the Faiss retrieval algorithm, the multiple second feature vectors are respectively determined as the indexes of the multiple face images in the face database, which not only ensures a fast retrieval speed but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy.

[0047] Optionally, it further includes: determining the first feature vector as the index of the image to be retrieved in the face database according to the Faiss retrieval algorithm; and storing the first feature vector and the image to be retrieved in the feature database and the face database respectively.

[0048] Specifically, in order to continuously update and expand the feature database and the face database, in this embodiment, according to the Faiss retrieval algorithm, the first feature vector can also be determined as the index of the image to be retrieved in the face database, and then the first feature vector and the image to be retrieved are respectively stored in the feature database and the face database.

[0049] Optionally, the operation of determining the target feature vector from multiple second feature vectors in the preset feature database according to the first feature vector includes: comparing the first feature vector with the multiple second feature vectors respectively; and determining the target feature vector from the multiple second feature vectors according to the comparison result.

[0050] Specifically, in the process of determining the target feature vector, in this embodiment, the first feature vector is first compared with the multiple second feature vectors one by one, and then the target feature vector is determined from the multiple second feature vectors according to the comparison result. That is, the second feature vector with the best comparison result is determined as the target feature vector. In this way, the accuracy of the determined target feature vector is guaranteed, laying a foundation for retrieving the most similar face image from the face database in the next step.

[0051] Optionally, the operation of comparing the first feature vector with multiple second feature vectors respectively includes: calculating the cosine similarity between the first feature vector and the multiple second feature vectors respectively; and the operation of determining the target feature vector from the multiple second feature vectors according to the comparison result includes: determining the feature vector with the highest cosine similarity with the first feature vector among the multiple second feature vectors as the target feature vector.

[0052] Specifically, in the process of comparing the first feature vector with multiple second feature vectors respectively, in this embodiment, the cosine similarity is used as the evaluation criterion to calculate the cosine similarity between the first feature vector and the multiple second feature vectors respectively. Then, the second feature vector with the highest cosine similarity is determined as the target feature vector. In this way, the target feature vector most similar to the first feature vector can be determined from the multiple second feature vectors.

[0053] Optionally, before the operation of generating the first feature vector corresponding to the face in the image to be retrieved by using the feature extraction model, it further includes: performing face detection on the image to be retrieved, determining a face image region containing a face from the image to be retrieved; and performing face alignment processing on the face image region.

[0054] Specifically, Figure 3 An exemplary overall process schematic diagram of the face retrieval method described in this embodiment is shown. Refer to Figure 2 As shown, for a given face image, it often contains a background image, and these background images are not needed during face retrieval and will affect the retrieval accuracy. Therefore, the background needs to be removed before feature extraction, that is, performing face detection on the image to be retrieved and determining a face image region containing a face from the image to be retrieved. In this way, a face image region containing only a face can be obtained. Further, since a higher accuracy can be obtained when the aligned face is fed into the feature extraction model, and the angle of the face in the detected face image region may not be very straight, it is necessary to align it, that is, perform face alignment processing on the face image region.

[0055] Optionally, the operation of performing face detection on the image to be retrieved and determining a face image region containing a face from the image to be retrieved includes: performing face detection on the image to be retrieved; judging whether the image to be retrieved is a face image according to the detection result of the face detection; and in the case where it is determined that the image to be retrieved is a face image, cropping out a face image region containing a face from the image to be retrieved.

[0056] Specifically, refer to Figure 3As shown, during the face detection process, first, the detection result of face detection is used to determine whether the image to be retrieved is a face image. If no face is detected, the following operations are not performed. If a face is detected, the face region is intercepted to obtain a face image region that only contains the face. In this way, not only the workload of face detection is reduced, but also it effectively guarantees that the determined face image region only contains the face.

[0057] Optionally, after the operation of intercepting the face image region containing the face from the image to be retrieved, it further includes: converting the resolution of the face image region into the resolution of an image suitable for feature extraction by the feature extraction model.

[0058] Specifically, since the feature extraction model can often only perform feature extraction on images within a certain resolution range, in order to ensure that the feature extraction model can effectively perform feature extraction on the image to be retrieved, it is necessary to convert the resolution of the face image region in the retrieved image into the resolution of an image suitable for feature extraction by the feature extraction model. For example, the face image region is uniformly scaled to an image of size 112*112.

[0059] In addition, it should be supplemented that the feature extraction model used in this embodiment is a lightweight feature extraction network, the MobileFaceNet model, and in order to obtain more rich-dimensional features, the output of the MobileFaceNet model is specifically changed from 128 dimensions to a feature vector of 512 dimensions. Figure 4 An example shows the schematic diagram of the network structure of the feature extraction model used in this embodiment, where c is the number of channels, n is the number of repetitions, and s is the stride. Refer to Figure 3 and Figure 4 As shown, for the aligned face image of 112*112, it is fed into the feature extraction network. First, it will go through a 3*3 convolution to increase the number of channels from 3 to 64 and halve the resolution of the image. The final output is a feature map of 56*56*64. Then, it goes through a convolutional layer with a convolutional kernel size of 3*3 and a stride of 1, followed by 5 bottleneck layers, and outputs a feature map of 7*7*128. Then, a 1*1 convolution is used to increase the number of channels to 512 dimensions. Finally, a GAP layer is used for global pooling to become a feature map of 1*1*512 as the final output.

[0060] Among them, the bottleneck structure is as Figure 5As shown in the figure, the input image goes through three convolutional layers respectively, that is, first through a 1*1 convolution, then through a 3*3 convolution, and finally through a 1*1 convolution. In addition, in order to prevent the loss of information during the convolution process, a direct connection operation is added in the bottleneck structure.

[0061] In addition, the Green DeepPupil MS-1M dataset after screening and duplicate removal and the private dataset can be used to train the feature extraction model. Among them, the screened public dataset MS-1M has a total of 242,116 images of 16,132 individuals, and these images are all Asian face images. In addition, the private dataset uses the internally accumulated face liveness data, with a total of 5 million face images of 1 million individuals. Before training the feature extraction network, the training sets are all uniformly scaled to a size of 112*112.

[0062] In summary, the face retrieval method proposed in this disclosure can produce the following beneficial effects:

[0063] 1) Fast retrieval speed. Since this embodiment does not adopt the traditional database retrieval method, but uses the open-source retrieval algorithm faiss launched by Facebook, the faiss retrieval speed is very fast. In the retrieval stage, the open-source faiss retrieval algorithm is used to build an index for the 512-dimensional feature vectors output by the feature extraction model, without performing dimensionality reduction or hash coding, and using cosine similarity as the evaluation criterion. According to experimental data, when the face database has 1 million faces, it only takes 0.2s to retrieve once using the CPU and only 20ms using the GPU, which can achieve real-time retrieval.

[0064] 2) High retrieval accuracy. This embodiment uses the relatively advanced deep learning model MobileFaceNet model, which is trained on the public dataset and the private dataset. In addition, this disclosure removes the last layer of the MobileFaceNet network, increasing the dimension of the output of the original model from 128 to 512, thus obtaining higher-dimensional face features and higher retrieval accuracy. In addition, all conv layers in the MobileFaceNet model are replaced with depthwise separable convolutions, which can reduce the network parameters by about 1 / 9.

[0065] In addition, referring to Figure 1 As shown in the figure, according to the second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, where, when the program runs, the method described in any one of the above is executed by a processor.

[0066] Thus, according to this embodiment, in the feature extraction stage of face retrieval, a feature extraction model based on deep learning is adopted to extract face features from the image to be retrieved, obtaining a first feature vector with rich feature information, which makes the accuracy of the retrieval result relatively high. In the face retrieval stage, according to the first feature vector, a target feature vector is determined from multiple second feature vectors in the preset feature database. Since in the previous index establishment stage, this embodiment does not perform dimensionality reduction or hash encoding on the second feature vector, but directly uses the Faiss retrieval algorithm with very fast retrieval speed to determine the second feature vector as the index of the face image in the face database. Therefore, the retrieval can be performed in the face database according to the target feature vector to obtain the retrieval result corresponding to the image to be retrieved. Since the Faiss retrieval algorithm is directly used to determine the second feature vector as the index of the face image in the face database, it not only ensures a fast retrieval speed but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy. It achieves the technical effects of fast retrieval speed and high accuracy in face retrieval. Furthermore, it solves the technical problem in the prior art that the traditional way of extracting features cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, dimensionality reduction processing or hash encoding needs to be performed on the face features, resulting in the loss of information dimensions and low accuracy.

[0067] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in multiple embodiments of the present invention.

[0069] Embodiment 2

[0070] Figure 6The face retrieval device 600 according to the present embodiment is shown. This device 600 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 6 As shown, the device 600 includes: an acquisition module 610, configured to acquire a to-be-retrieved image including the face of a target object; a feature extraction module 620, configured to generate a first feature vector corresponding to the face in the to-be-retrieved image by using a feature extraction model; a determination module 630, configured to determine a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively the feature vectors of multiple face images in a preset face database; and a retrieval module 640, configured to perform a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively the indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.

[0071] Optionally, it further includes a first index establishment module, configured to establish indexes of multiple face images in the face database in the following manner: generate multiple second feature vectors respectively corresponding to the faces in the multiple face images by using a feature extraction model; and determine the multiple second feature vectors as the indexes of the multiple face images in the face database according to the Faiss retrieval algorithm.

[0072] Optionally, it further includes: a second index establishment module, configured to determine the first feature vector as the index of the to-be-retrieved image in the face database according to the Faiss retrieval algorithm; and a storage module, configured to store the first feature vector and the to-be-retrieved image into the feature database and the face database respectively.

[0073] Optionally, the determination module 630 includes: a comparison sub-module, configured to compare the first feature vector with the multiple second feature vectors respectively; and a determination sub-module, configured to determine a target feature vector from the multiple second feature vectors according to the comparison result.

[0074] Optionally, the comparison sub-module includes: a calculation unit, configured to calculate the cosine similarity between the first feature vector and the multiple second feature vectors respectively; and the determination sub-module includes: a determination unit, configured to determine the feature vector with the highest cosine similarity with the first feature vector among the multiple second feature vectors as the target feature vector.

[0075] Optionally, it further includes: a face image region determination module, configured to perform face detection on the to-be-retrieved image before the operation of the feature extraction module 620 generating a first feature vector corresponding to the face in the to-be-retrieved image by using a feature extraction model, and determine a face image region including a face from the to-be-retrieved image; and perform face alignment processing on the face image region.

[0076] Optionally, the face image region determination module includes a retrieval sub-module for performing face detection on the image to be retrieved; a determination sub-module for determining whether the image to be retrieved is a face image according to the detection result of the face detection; and a cropping sub-module for cropping out a face image region containing a face from the image to be retrieved when it is determined that the image to be retrieved is a face image.

[0077] Therefore, according to this embodiment, in the feature extraction stage of face retrieval, a feature extraction model based on deep learning is adopted to extract face features from the image to be retrieved, and a first feature vector with rich feature information is obtained, so that the accuracy of the retrieval result is relatively high. In the face retrieval stage, according to the first feature vector, a target feature vector is determined from multiple second feature vectors in a preset feature database. Since in the previous index establishment stage, this embodiment does not perform dimensionality reduction or hash coding on the second feature vector, but directly determines the second feature vector as the index of the face image in the face database by using the Faiss retrieval algorithm with very fast retrieval speed. Therefore, the face database can be retrieved according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved. Since the Faiss retrieval algorithm is used to directly determine the second feature vector as the index of the face image in the face database, it not only ensures a fast retrieval speed but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy. The technical effects of fast retrieval speed and high accuracy of face retrieval are achieved. Furthermore, it solves the technical problem in the prior art that the traditional feature extraction method cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of establishing a feature index for face features, in order to improve the retrieval rate, it is necessary to perform dimensionality reduction processing or hash coding on the face features, resulting in the loss of information dimensions and low accuracy.

[0078] Embodiment 3

[0079] Figure 7 Fig. shows a face retrieval device 700 according to this embodiment, and the device 700 corresponds to the method described in the first aspect of Embodiment 1. Refer to Figure 7As shown, the device 700 includes: a processor 710; and a memory 720, connected to the processor 710, for providing instructions for the processor 710 to process the following processing steps: obtaining a to-be-retrieved image including a face of a target object; using a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and retrieving in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of multiple face images in the face database established according to the Faiss retrieval algorithm.

[0080] Optionally, the memory 720 is further used to provide instructions for the processor 710 to process the following processing steps: establishing indexes of multiple face images in the face database in the following manner: using a feature extraction model to generate multiple second feature vectors respectively corresponding to the faces in the multiple face images; and according to the Faiss retrieval algorithm, respectively determining the multiple second feature vectors as indexes of the multiple face images in the face database.

[0081] Optionally, the memory 720 is further used to provide instructions for the processor 710 to process the following processing steps: according to the Faiss retrieval algorithm, determining the first feature vector as an index of the to-be-retrieved image in the face database; and storing the first feature vector and the to-be-retrieved image into the feature database and the face database respectively.

[0082] Optionally, the operation of determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector includes: comparing the first feature vector with the multiple second feature vectors respectively; and determining the target feature vector from the multiple second feature vectors according to the comparison result.

[0083] Optionally, the operation of comparing the first feature vector with the multiple second feature vectors respectively includes: calculating the cosine similarity between the first feature vector and the multiple second feature vectors respectively; and the operation of determining the target feature vector from the multiple second feature vectors according to the comparison result includes: determining the feature vector with the highest cosine similarity between the multiple second feature vectors and the first feature vector as the target feature vector.

[0084] Optionally, the memory 720 is further configured to provide instructions for the processor 710 to perform the following processing steps: before generating the first feature vector corresponding to the face in the image to be retrieved by using the feature extraction model, performing face detection on the image to be retrieved, determining a face image region containing a face from the image to be retrieved; and performing face alignment processing on the face image region.

[0085] Optionally, the operation of performing face detection on the image to be retrieved and determining a face image region containing a face from the image to be retrieved includes: performing face detection on the image to be retrieved; determining whether the image to be retrieved is a face image according to the detection result of the face detection; and when it is determined that the image to be retrieved is a face image, cropping out the face image region containing the face from the image to be retrieved.

[0086] Thus, according to this embodiment, in the feature extraction stage of face retrieval, a feature extraction model based on deep learning is adopted to extract face features from the image to be retrieved, and a first feature vector with rich feature information is obtained, so that the accuracy of the retrieval result is relatively high. In the face retrieval stage, according to the first feature vector, a target feature vector is determined from multiple second feature vectors in the preset feature database. Since in the previous index building stage, this embodiment does not perform dimensionality reduction or hash coding on the second feature vector, but directly uses the Faiss retrieval algorithm with very fast retrieval speed to determine the second feature vector as the index of the face image in the face database. Therefore, the retrieval can be performed in the face database according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved. Since the Faiss retrieval algorithm is directly used to determine the second feature vector as the index of the face image in the face database, it not only ensures a fast retrieval speed but also does not cause the loss of information dimensions, greatly improving the retrieval accuracy. It achieves the technical effects of fast retrieval speed and high accuracy in face retrieval. Furthermore, it solves the technical problem in the prior art that the traditional way of extracting features cannot obtain good face features, while the deep learning method can obtain relatively high-dimensional and rich face features, but in the process of building a feature index for face features, in order to improve the retrieval rate, dimensionality reduction processing or hash coding needs to be performed on the face features, resulting in the loss of information dimensions and low accuracy.

[0087] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0088] In the above embodiments of the present invention, the descriptions of multiple embodiments each have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0089] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0090] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0091] In addition, in various embodiments of the present invention, each functional unit can be integrated in a processing unit, or multiple units can exist separately physically, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0092] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, and other media that can store program codes.

[0093] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A face retrieval method, characterized in that, Including: Obtain a to-be-retrieved image including the face of a target object; Utilize a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image; According to the first feature vector, determine a target feature vector from multiple second feature vectors in a preset feature database, where the multiple second feature vectors are respectively feature vectors of multiple face images in a preset face database; and According to the target feature vector, conduct a retrieval in the face database to obtain a retrieval result corresponding to the to-be-retrieved image, where the multiple second feature vectors in the feature database are respectively indexes of the multiple face images in the face database established according to the Faiss retrieval algorithm; It further includes establishing indexes of the multiple face images in the face database in the following manner: Utilize the feature extraction model to generate the multiple second feature vectors respectively corresponding to the faces in the multiple face images; and According to the Faiss retrieval algorithm, respectively determine the multiple second feature vectors as indexes of the multiple face images in the face database; It further includes: According to the Faiss retrieval algorithm, determine the first feature vector as an index of the to-be-retrieved image in the face database; and Respectively store the first feature vector and the to-be-retrieved image into the feature database and the face database; The operation of determining a target feature vector from multiple second feature vectors in a preset feature database according to the first feature vector includes: Compare the first feature vector with the multiple second feature vectors respectively; and According to the result of the comparison, determine the target feature vector from the multiple second feature vectors; The operation of comparing the first feature vector with the multiple second feature vectors respectively includes: respectively calculating the cosine similarity between the first feature vector and the multiple second feature vectors; and The operation of determining the target feature vector from the multiple second feature vectors according to the result of the comparison includes: determining the feature vector with the highest cosine similarity with the first feature vector among the multiple second feature vectors as the target feature vector; Before the operation of utilizing a feature extraction model to generate a first feature vector corresponding to the face in the to-be-retrieved image, it further includes: Conduct face detection on the to-be-retrieved image, and determine a face image region including a face from the to-be-retrieved image; and Conduct face alignment processing on the face image region; The operation of conducting face detection on the to-be-retrieved image and determining a face image region including a face from the to-be-retrieved image includes: Conduct face detection on the to-be-retrieved image; According to the detection result of the face detection, determine whether the to-be-retrieved image is a face image; and In the case of determining that the to-be-retrieved image is a face image, intercept a face image region including a face from the to-be-retrieved image.

2. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program runs, the method according to claim 1 is executed by a processor.

3. A face retrieval device, characterized in that, Including: An acquisition module, configured to acquire a to-be-retrieved image including a face of a target object; A feature extraction module, configured to generate a first feature vector corresponding to the face in the to-be-retrieved image by using a feature extraction model; A determination module, configured to determine a target feature vector from a plurality of second feature vectors in a preset feature database according to the first feature vector, where the plurality of second feature vectors are respectively feature vectors of a plurality of face images in a preset face database; And A retrieval module, configured to perform a retrieval in the face database according to the target feature vector to obtain a retrieval result corresponding to the to-be-retrieved image, where the plurality of second feature vectors in the feature database are respectively indexes of the plurality of face images in the face database established according to the Faiss retrieval algorithm; The feature extraction module is further configured to: Establish indexes of the plurality of face images in the face database in the following manner: generate the plurality of second feature vectors respectively corresponding to the faces in the plurality of face images by using the feature extraction model; and determine the plurality of second feature vectors as indexes of the plurality of face images in the face database according to the Faiss retrieval algorithm; The feature extraction module is further configured to: Determine the first feature vector as an index of the to-be-retrieved image in the face database according to the Faiss retrieval algorithm; and store the first feature vector and the to-be-retrieved image into the feature database and the face database respectively; The determination module is specifically configured to: compare the first feature vector with the plurality of second feature vectors respectively; and determine the target feature vector from the plurality of second feature vectors according to the comparison result; The determination module is further specifically configured to: The operation of comparing the first feature vector with the plurality of second feature vectors respectively includes: calculating the cosine similarity between the first feature vector and the plurality of second feature vectors respectively; and the operation of determining the target feature vector from the plurality of second feature vectors according to the comparison result includes: determining the feature vector with the highest cosine similarity with the first feature vector among the plurality of second feature vectors as the target feature vector; The acquisition module is further configured to: Perform face detection on the to-be-retrieved image, determine a face image region including a face from the to-be-retrieved image; and perform face alignment processing on the face image region; The acquisition module is further specifically configured to: Perform face detection on the to-be-retrieved image; Judge whether the to-be-retrieved image is a face image according to the detection result of the face detection; and When it is determined that the to-be-retrieved image is a face image, intercept a face image region including a face from the to-be-retrieved image.

4. A face retrieval device, characterized in that, Including: A processor; And A memory, connected to the processor, configured to provide instructions for the processor to perform the following processing steps: Acquire a to-be-retrieved image including a face of a target object; Generate a first feature vector corresponding to the face in the image to be retrieved by using a feature extraction model; Determine a target feature vector from a plurality of second feature vectors in a preset feature database according to the first feature vector, where the plurality of second feature vectors are respectively feature vectors of a plurality of face images in a preset face database; and Retrieve in the face database according to the target feature vector to obtain a retrieval result corresponding to the image to be retrieved, where the plurality of second feature vectors in the feature database are respectively indexes of the plurality of face images in the face database established according to the Faiss retrieval algorithm; The memory is further configured to provide instructions for the processor to perform the following processing steps to establish indexes of the plurality of face images in the face database: generate the plurality of second feature vectors respectively corresponding to the faces in the plurality of face images by using the feature extraction model; and determine the plurality of second feature vectors as indexes of the plurality of face images in the face database according to the Faiss retrieval algorithm; Determine the first feature vector as an index of the image to be retrieved in the face database according to the Faiss retrieval algorithm; and store the first feature vector and the image to be retrieved into the feature database and the face database respectively; The determining module is specifically configured to: compare the first feature vector with the plurality of second feature vectors respectively; and determine the target feature vector from the plurality of second feature vectors according to the comparison result; The operation of comparing the first feature vector with the plurality of second feature vectors respectively includes: calculating the cosine similarity between the first feature vector and the plurality of second feature vectors respectively; and the operation of determining the target feature vector from the plurality of second feature vectors according to the comparison result includes: determining the feature vector with the highest cosine similarity between the first feature vector and the plurality of second feature vectors as the target feature vector; Perform face detection on the image to be retrieved, and determine a face image region containing a face from the image to be retrieved; and perform face alignment processing on the face image region; Perform face detection on the image to be retrieved; Judge whether the image to be retrieved is a face image according to the detection result of the face detection; and In the case of determining that the image to be retrieved is a face image, extract a face image region containing a face from the image to be retrieved.

Citation Information

Patent Citations

  • Target retrieval method and device, computer readable storage medium and computer device

    CN110866491A