Facial recognition model training method, device and electronic device

By adjusting the distance relationship between face images using the five-tuple loss function, the problem of low recognition accuracy of face recognition model training results in the prior art is solved, especially when the face image quality is uneven, higher recognition accuracy is achieved.

CN114333015BActive Publication Date: 2025-05-20ISA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111640073.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-20
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The existing model recognition accuracy of the face recognition model training results based on the ternary loss function is low, especially when the face image quality is uneven.

Method used

The initial face recognition model is trained using a five-tuple loss function. By adjusting the distance relationship between different face images, the distance between the face image to be trained is closer to the face image with high quality, and the distance between the face image with low quality is further away.

Benefits of technology

The recognition accuracy of the face recognition model is improved, especially when the face image quality is uneven, high-quality face images can be used more effectively, thereby improving the overall performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333015B_ABST
    Figure CN114333015B_ABST
Patent Text Reader

Abstract

The present application provides a training method, device and electronic device for a face recognition model, which relates to the field of image recognition technology and alleviates the technical problem of low model recognition accuracy after the existing face recognition model is trained. The method includes: obtaining training samples of face images; determining the face image quality of the face image of each object; training the initial face recognition model based on the training samples using a five-tuple loss function to reduce the distance between the first image and the second image and the third image and increase the distance between the first image and the fourth image and the fifth image, and to make the distance between the first image and the second image smaller than the distance between the first image and the third image, and obtain the face recognition model training result; the five-tuple includes the first image, the second image and the third image of the first object, and the fourth image and the fifth image of the second object; the face image quality corresponding to the second image is higher than the face image quality corresponding to the third image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a method, device, and electronic device for training a face recognition model. Background Art

[0002] Currently, the current face recognition algorithm based on metric learning mainly uses the Triple loss function, also known as the triplet loss function. For example, the three elements in it are Anchor, Negative, and Positive. Among them, Anchor is a face image randomly selected from the training dataset, Positive is a face image of the same person as Anchor, and Negative is a face image of a different person from Anchor. However, the model recognition accuracy of the face recognition model obtained after training the face recognition model using the triplet loss function is relatively low. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, device, and electronic device for training a face recognition model to alleviate the technical problem that the model recognition accuracy of the existing face recognition model after training is relatively low.

[0004] In a first aspect, an embodiment of the present application provides a method for training a face recognition model, the method comprising:

[0005] Obtain training samples of face images; wherein, the face images include face images of multiple different objects;

[0006] Determine the face image quality of the face images of each of the objects;

[0007] Train an initial face recognition model using a loss function of a quintuple based on the training samples, so that the distance between the first image and the second image and the third image decreases, and the distance between the first image and the fourth image and the fifth image increases, and the distance between the first image and the second image is less than the distance between the first image and the third image, to obtain a face recognition model training result; wherein, the quintuple includes the first image, the second image, and the third image of the first object, and the fourth image and the fifth image of the second object; the first image is a face image to be trained, and the face image quality corresponding to the second image is higher than the face image quality corresponding to the third image.

[0008] In a possible implementation, the step of determining the face image quality of the face images of each of the objects includes:

[0009] Detect the face image quality of the face images using a specified face image quality detection model to obtain an image quality detection result;

[0010] Determine the face image quality for each of the objects according to the image quality detection result.

[0011] In a possible implementation, the step of determining the face image quality for each of the objects according to the image quality detection result includes:

[0012] Perform face image quality ranking for the same object based on the face image quality to obtain the face image quality ranking result corresponding to each object.

[0013] In a possible implementation, the loss function of the five-tuple includes:

[0014]

[0015] Wherein, is used to represent the first image, the is used to represent the second image; is used to represent the third image; is used to represent the fourth image; is used to represent the first image; is used to represent the fifth image; the α1 is used to represent the interval between the distance between the second image and the first image and the distance between the first image and the fourth image; the α2 is used to represent the interval between the distance between the third image and the first image and the distance between the first image and the fifth image.

[0016] In a possible implementation, the is used to represent that the distance between the second image and the first image becomes the smallest, while the distance between the first image and the fourth image becomes larger;

[0017] The is used to represent that the distance between the third image and the first image becomes the smallest, while the distance between the first image and the fifth image becomes larger;

[0018] The is used to represent that the distance between the second image and the first image is less than the distance between the third image and the first image.

[0019] In a possible implementation, it further includes:

[0020] Clean the face images to remove the face images in which the face image quality is lower than the preset image quality.

[0021] In a possible implementation, the step of obtaining training samples of face images includes:

[0022] Collect face images and group the face images based on different objects to obtain the face images corresponding to different objects respectively.

[0023] In a second aspect, a training device for a face recognition model is provided. The device includes:

[0024] An acquisition module for acquiring training samples of face images; wherein, the face images include face images of multiple different objects;

[0025] A determination module for determining the face image quality of the face images of each object;

[0026] A training module for training an initial face recognition model based on the training samples using a loss function of a five-tuple, so as to reduce the distance between a first image and a second image and a third image, and increase the distance between the first image and a fourth image and a fifth image, and make the distance between the first image and the second image less than the distance between the first image and the third image, to obtain a face recognition model training result; wherein, the five-tuple includes the first image, the second image and the third image of a first object, and the fourth image and the fifth image of a second object; the first image is a face image to be trained, and the face image quality corresponding to the second image is higher than the face image quality corresponding to the third image.

[0027] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the method described in the first aspect above is implemented.

[0028] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to run the method described in the first aspect above.

[0029] The embodiments of the present application bring the following beneficial effects:

[0030] A training method, device, and electronic device for a face recognition model provided by an embodiment of the present application can obtain training samples of face images; wherein, the face images include face images of multiple different objects; determine the face image quality of the face images of each object; train an initial face recognition model based on the training samples using a loss function of a five-tuple, so that the distance between the first image and the second image and the third image decreases, and the distance between the first image and the fourth image and the fifth image increases, and make the distance between the first image and the second image less than the distance between the first image and the third image, to obtain a face recognition model training result; wherein, the five-tuple includes the first image, the second image, and the third image of the first object, and the fourth image and the fifth image of the second object; the first image is the face image to be trained, and the face image quality corresponding to the second image is higher than the face image quality corresponding to the third image. In this solution, through the loss function of the five-tuple, not only can the distance between the face image to be trained and the face images of the same object be minimized, and the distance between the face image to be trained and the face images of different objects be maximized, but also the distance between the face image to be trained and the face image with higher quality among the face images of the same object can be made closer than the distance between the face image with lower quality. Even if the face image quality is uneven, during the process of training the face recognition model using the above five-tuple loss function, the face feature vector can distinguish the quality of the image, and can make the face feature vector as close as possible to the face image with higher quality, instead of the default image quality collected by the triplet loss function being the same. Furthermore, it can ensure that the face feature vector is close to the face image with higher quality during the training process. Therefore, by using the five-tuple loss function to train the face recognition model, the situation of uneven face quality can be more effectively utilized, that is, in the actual situation, the face images in the collected pictures are affected by factors such as illumination, distance, occlusion, and shooting angle, resulting in uneven face image quality, thereby improving the recognition accuracy of the face recognition model after training.

[0031] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1Schematic flowchart of the method for training a face recognition model provided by an embodiment of the present application;

[0034] Figure 2 Another schematic flowchart of the method for training a face recognition model provided by an embodiment of the present application;

[0035] Figure 3 An example of a five-tuple in the method for training a face recognition model provided by an embodiment of the present application;

[0036] Figure 4 Schematic structural diagram of a training device for a face recognition model provided by an embodiment of the present application;

[0037] Figure 5 Schematic structural diagram of an electronic device provided by an embodiment of the present application is shown. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0039] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0040] Currently, the current face recognition algorithm based on metric learning mainly uses the Triple loss function, also known as the triplet loss function. For example, the three elements in it are Anchor, Negative, and Positive. Among them, Anchor is a face image randomly selected from the training dataset, Positive is a face image of the same person as Anchor, and Negative is a face image of a different person from Anchor. After learning through the triplet loss function, the distance between Positive and Anchor is minimized, while the distance between Positive and Negative is maximized. The specific formula of the triplet loss is as follows: In the above formula is the Euclidean distance represents the Euclidean distance between the Positive element and Anchor. It represents the Euclidean distance between Negative and Anchor. α, also known as margin, refers to the minimum interval between the distance between Positive and Anchor and the distance between Negative and Anchor. The smaller the margin value is set, the easier it is for the loss to approach 0, but it is difficult to distinguish similar images. The larger the margin value is set, the more difficult it is for the loss value to approach 0, and it may even cause the network not to converge, but it can more confidently distinguish relatively similar images. After learning the triplet loss function, the distance between Positive and Anchor becomes smaller, while the distance from Negative becomes larger, that is, the face vectors of the training samples approach the face vectors of the same face and move away from the face vectors of different faces.

[0041] However, in reality, the face images of the same person are affected by conditions such as distance, lighting, occlusion, and shooting angle, and the image quality is uneven. When training a face recognition model at this time, the face feature vectors should be made as close as possible to the Positive images with high image quality. If the triplet loss function is still used, during the training process of the face recognition model, the face feature vectors do not distinguish the quality of the images, assuming that the image quality of the collected images is the same, and it cannot ensure that the face feature vectors are close to the high-quality Positive images during the training process. The triplet loss function assumes that all face qualities are the same, but in the real-world scenario, the face qualities are different. For example, in the real-world scenario, the face quality of a frontal shot is better than that of a side shot. Although both the frontal and side shots are of the same person, the face vectors trained by the face recognition model should be more inclined to the face images of the frontal shot, resulting in the triplet loss function ignoring the face quality problem. Therefore, the recognition accuracy of the face recognition model obtained after training the face recognition model using the triplet loss function is relatively low.

[0042] Based on this, the embodiments of the present application provide a method, device, and electronic device for training a face recognition model, which can alleviate the technical problem of relatively low recognition accuracy of the existing face recognition model after training.

[0043] The embodiments of the present invention will be further introduced below with reference to the accompanying drawings.

[0044] Figure 1 It is a schematic flowchart of a method for training a face recognition model provided by an embodiment of the present application. Among them, this method is applied to a computer device. As Figure 1 shown, this method includes:

[0045] Step S110, obtaining training samples of face images.

[0046] Among them, the facial image contains facial images of multiple different objects.

[0047] In this step, exemplarily, such as Figure 2 shown, the facial image can be obtained by collecting face pictures. The collection of face pictures can be the face capture photos under a real camera, or the capture photos of a bayonet camera and the capture photos of a face camera.

[0048] Step S120, determine the facial image quality of the facial image of each object.

[0049] In this step, exemplarily, such as Figure 2 shown, a face quality model can be used to score the facial image. The higher the score, the higher the face quality.

[0050] Step S130, train the initial facial recognition model using the loss function of the five-tuple based on the training samples, so that the distance between the first image and the second image, the third image decreases, and the distance between the first image and the fourth image, the fifth image increases, and make the distance between the first image and the second image less than the distance between the first image and the third image, to obtain the training result of the facial recognition model.

[0051] Among them, the five-tuple includes the first image, the second image and the third image of the first object, and the fourth image and the fifth image of the second object; the first image is the facial image to be trained, and the facial image quality corresponding to the second image is higher than the facial image quality corresponding to the third image.

[0052] In this step, exemplarily, such as Figure 2 shown, use the five-tuple loss function to train the facial recognition model. When training, randomly select a face picture each time, two pictures of people belonging to the same person as the previous face, and two pictures of different people.

[0053] Regarding the five-tuple loss function, it should be noted that the triplet loss function is one Anchor, one Negative, and one Positive, while the five-tuple loss function is one Anchor, two Negatives, and two Positives. As Figure 3 shown, where Anchor is a face picture randomly selected from the training dataset, two Positives are two face pictures belonging to the same person as Anchor. Two Negatives are each one face picture belonging to two different people from Anchor.

[0054] For example, such as Figure 3As shown, the quintuple loss includes an Anchor face image to be trained, two Positive1 and Positive2 images of the same person as the image to be trained but with different image qualities, where the face quality of Positive1 is greater than that of Positive2, and two Negative1 and Negative2 images of different people from the image to be trained.

[0055] In the embodiments of the present application, through the loss function of the quintuple, it can not only minimize the distance between the face image to be trained and the face images of the same object, and increase the distance between the face image to be trained and the face images of different objects, but also make the distance between the face image to be trained and the face image with higher quality among the face images of the same object closer than the distance between the face image with lower quality. Even if the face image qualities are uneven, during the process of training the face recognition model using the above quintuple loss function, the face feature vector can distinguish the image quality, and can make the face feature vector as close as possible to the image with higher face image quality, rather than the default image quality collected by the triple loss function being the same. Furthermore, it can ensure that the face feature vector is close to the face image with higher quality during the training process. Therefore, the process of training the face recognition model using the quintuple loss function can make more effective use of the situation where the face qualities are uneven, that is, in the actual situation, the face images in the collected pictures are affected by factors such as illumination, distance, occlusion, and shooting angle, resulting in uneven face image qualities.

[0056] The above steps will be introduced in detail below.

[0057] In some embodiments, the above step S120 may include the following steps:

[0058] Step a), using a specified face image quality detection model to detect the face image quality of the face image, and obtaining an image quality detection result;

[0059] Step b), determining the face image quality for each object according to the image quality detection result.

[0060] For example, a face quality model can be used to score all the pictures in the same group. The higher the score, the higher the face quality. By using a specified face image quality detection model to detect the face image quality of the face image, the accuracy of the face image quality result can be higher.

[0061] Based on the above steps a) and b), the above step b) may include the following steps:

[0062] Step c), performing face image quality sorting for the same object based on the face image quality, and obtaining a face image quality sorting result corresponding to each object.

[0063] In practical applications, exemplarily, a face quality model can be used to score all faces, and the face scores of the same person can be sorted from high to low. By sorting the face image quality for the same object, the face image quality of the face images under the same object can be made more explicit and easier to distinguish the quality level.

[0064] In some embodiments, the loss function of the five-tuple includes:

[0065]

[0066] Wherein, is used to represent the first image, is used to represent the second image; is used to represent the third image; is used to represent the fourth image; is used to represent the first image; is used to represent the fifth image; α1 is used to represent the interval between the distance between the second image and the first image and the distance between the first image and the fourth image; α2 is used to represent the interval between the distance between the third image and the first image and the distance between the first image and the fifth image.

[0067] In the specific formula of the above five-tuple loss function, α is also called margin, which refers to the minimum interval between the distance between the second image and the first image and the distance between the first image and the fourth image. The smaller the margin value is set, the easier it is for the loss to approach 0, but it is difficult to distinguish similar images; the larger the margin value is set, the more difficult it is for the loss value to approach 0, and even the network may not converge, but it can more surely distinguish relatively similar images.

[0068] Exemplarily, as Figure 3 described, the five-tuple loss includes an Anchor of a face picture to be trained, two pictures Positive1 and Positive2 of the same person as the picture to be trained with different picture qualities, where the face quality of Positive1 is greater than that of Positive2, and two pictures Negative1 and Negative2 of different people from the picture to be trained.

[0069] In the embodiments of the present application, through the formula in the above loss function of the five-tuple, the five-tuple loss function with more accurate data can be efficiently determined, so that the model result trained using the five-tuple loss function is more accurate.

[0070] Based on this, is used to represent that the distance between the second image and the first image becomes smaller, while the distance between the first image and the fourth image becomes larger; Used to indicate that the distance between the third image and the first image is minimized, while the distance between the first image and the fifth image increases; Used to indicate that the distance between the second image and the first image is less than the distance between the third image and the first image.

[0071] For example, as Figure 3 shown, the first part of the formula indicates that the distance between Positive1 and Anchor is minimized, while the distance between Positive1 and Negative1 increases; the second part of the formula indicates that the distance between Positive2 and Anchor is minimized, while the distance between Positive2 and Negative2 increases; the third part of the formula indicates that the distance between Positive1 and Anchor is less than the distance between Positive2 and Anchor.

[0072] Through the five - tuple loss function in the embodiments of the present application, it can not only minimize the distance between the face image to be trained and the face images of the same object, and increase the distance between the face image to be trained and the face images of different objects, but also make the distance between the face image to be trained and the face image with higher quality among the face images of the same object closer to the distance between the face image with lower quality, as Figure 3 shown. Even if the quality of face images is uneven, during the training of the face recognition model, the face feature vectors can distinguish the quality of the images, and can make the face feature vectors as close as possible to the face images with higher quality, and then can make the face feature vectors close to the face images with higher quality during the training process.

[0073] In some embodiments, the method may further include the following steps:

[0074] Step d), cleaning the face images to remove the face images with face image quality lower than the preset image quality.

[0075] In practical applications, due to the influence of factors such as illumination, occlusion, and capture angle, the quality of captured face pictures is different. The captured photos can be preliminarily cleaned to remove blurred faces, avoiding the influence of images with too low quality and improving the utilization efficiency of training samples.

[0076] In some embodiments, the above - mentioned step S110 may include the following steps:

[0077] Step e), collecting face images and grouping the face images based on different objects to obtain face images corresponding to different objects respectively.

[0078] For example, for the collection of face images, it can be a face capture photo under a real camera, or a capture photo from a bayonet camera and a capture photo from a face camera. After collecting face images in reality, group the face images, and group the images of the same person into one group. Grouping the facial images based on different objects makes the use of facial images more targeted, so as to use facial images for each object more efficiently.

[0079] Figure 4 A structural schematic diagram of a training device for a face recognition model is provided. As Figure 4 shown, the training device 400 for the face recognition model includes:

[0080] An acquisition module 401, configured to acquire training samples of face images; wherein, the face images include face images of multiple different objects;

[0081] A determination module 402, configured to determine the face image quality of the face images of each object;

[0082] A training module 403, configured to train an initial face recognition model based on the training samples using a loss function of a five-tuple, so that the distance between the first image and the second image and the third image decreases, and the distance between the first image and the fourth image and the fifth image increases, and make the distance between the first image and the second image less than the distance between the first image and the third image, to obtain a face recognition model training result; wherein, the five-tuple includes the first image, the second image and the third image of the first object, and the fourth image and the fifth image of the second object; the first image is the face image to be trained, and the face image quality corresponding to the second image is higher than the face image quality corresponding to the third image.

[0083] In some embodiments, the determination module is specifically configured to:

[0084] Use a specified face image quality detection model to detect the face image quality of the face images, and obtain an image quality detection result;

[0085] Determine the face image quality for each object according to the image quality detection result.

[0086] In some embodiments, the determination module is further configured to:

[0087] Perform face image quality ranking for the same object based on the face image quality, and obtain a face image quality ranking result corresponding to each object.

[0088] In some embodiments, the loss function of the five-tuple includes:

[0089]

[0090] Among them, is used to represent the first image, the is used to represent the second image; is used to represent the third image; is used to represent the fourth image; is used to represent the first image; is used to represent the fifth image; the α1 is used to represent the interval between the distance between the second image and the first image and the distance between the first image and the fourth image; the α2 is used to represent the interval between the distance between the third image and the first image and the distance between the first image and the fifth image.

[0091] In some embodiments, the is used to represent that the distance between the second image and the first image becomes the smallest, while the distance between the first image and the fourth image becomes larger;

[0092] The is used to represent that the distance between the third image and the first image becomes the smallest, while the distance between the first image and the fifth image becomes larger;

[0093] The is used to represent that the distance between the second image and the first image is less than the distance between the third image and the first image.

[0094] In some embodiments, the device further includes:

[0095] A cleaning module, configured to clean the face image to remove the face image with the face image quality lower than the preset image quality in the face image.

[0096] In some embodiments, the obtaining module is specifically configured to:

[0097] Collect face images, and group the face images based on different objects to obtain the face images respectively corresponding to different objects.

[0098] The training device of the face recognition model provided by the embodiments of the present application has the same technical features as the face recognition model training method provided by the above embodiments, so it can also solve the same technical problems and achieve the same technical effects.

[0099] An electronic device provided by the embodiments of the present application, such as Figure 5As shown, the electronic device 500 includes a processor 502 and a memory 501. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the steps of the method provided in the foregoing embodiments are implemented.

[0100] See Figure 5 , the electronic device further includes: a bus 503 and a communication interface 504. The processor 502, the communication interface 504, and the memory 501 are connected through the bus 503. The processor 502 is configured to execute an executable module stored in the memory 501, such as a computer program.

[0101] Among them, the memory 501 may include a high-speed random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 504 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0102] The bus 503 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only a bidirectional arrow is used in

[0103] but it does not mean that there is only one bus or one type of bus.

[0104] The processor 502 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 502 or instructions in the form of software. The above-mentioned processor 502 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 501, and the processor 502 reads the information in the memory 501 and combines its hardware to complete the steps of the above method.

[0105] Corresponding to the above-mentioned face recognition model training method, an embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to run the steps of the above-mentioned face recognition model training method.

[0106] The face recognition model training device provided by the embodiments of the present application may be specific hardware on a device or software or firmware installed on the device, etc. For the device provided by the embodiments of the present application, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can all refer to the corresponding processes in the above method embodiments, and will not be repeated here.

[0107] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0108] For another example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0109] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] In addition, the various functional units in the embodiments provided in the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0111] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the face recognition model training method described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM for short), random access memories (RAM for short), magnetic disks, or optical discs.

[0112] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.

[0113] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solution of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for training a face recognition model, characterized in that: The method comprises: Acquire training samples of facial images; wherein the facial images include facial images of multiple different objects; determining a facial image quality of the facial image of each of said subjects; Based on the training samples, an initial face recognition model is trained using a quintuple loss function to reduce the distance between the first image and the second image and the third image and increase the distance between the first image and the fourth image and the fifth image, and to make the distance between the first image and the second image smaller than the distance between the first image and the third image, so as to obtain a face recognition model training result; wherein the quintuple includes the first image, the second image and the third image of a first object, and the fourth image and the fifth image of a second object; the first image is a face image to be trained, and the quality of the face image corresponding to the second image is higher than the quality of the face image corresponding to the third image.

2. The method according to claim 1, characterized in that The step of determining the facial image quality of each of the facial images of the object comprises: Using a specified facial image quality detection model to detect the facial image quality of the facial image, and obtain an image quality detection result; The facial image quality for each of the objects is determined according to the image quality detection result.

3. The method according to claim 2, characterized in that The step of determining the facial image quality for each of the objects according to the image quality detection result comprises: The facial image qualities of the same objects are sorted based on the facial image qualities to obtain a facial image quality sorting result corresponding to each of the objects.

4. The method according to claim 1, characterized in that: The loss function of the quintuple includes: in, for representing the first image, the for representing the second image; for representing the third image; for representing the fourth image; is used to represent the fifth image; the α 1 is used to represent the distance between the second image and the first image, and the interval between the distance between the first image and the fourth image; the α 2 Used to represent the interval between the distance between the third image and the first image, and the distance between the first image and the fifth image.

5. The method according to claim 4, characterized in that Said used to indicate that the distance between the second image and the first image becomes smallest, while the distance between the first image and the fourth image becomes larger; Said used to indicate that the distance between the third image and the first image becomes smallest, while the distance between the first image and the fifth image becomes larger; Said Used to indicate that the distance between the second image and the first image is smaller than the distance between the third image and the first image.

6. The method according to claim 1, characterized in that Also includes: The facial images are cleaned to remove facial images whose facial image quality is lower than a preset image quality.

7. The method according to claim 1, characterized in that The step of obtaining a training sample of a facial image comprises: Facial images are collected, and the facial images are grouped based on different objects to obtain the facial images corresponding to different objects.

8. A training device for a face recognition model, characterized in that: The device comprises: An acquisition module, used to acquire training samples of facial images; wherein the facial images include facial images of multiple different objects; a determination module, configured to determine the facial image quality of the facial image of each of the objects; A training module is used to train an initial face recognition model based on the training samples using a five-tuple loss function, so that the distance between the first image and the second image, the third image is reduced and the distance between the first image and the fourth image, the fifth image is increased, and the distance between the first image and the second image is smaller than the distance between the first image and the third image, so as to obtain a face recognition model training result; wherein the five-tuple includes the first image, the second image and the third image of the first object, and the fourth image and the fifth image of the second object; the first image is a face image to be trained, and the quality of the face image corresponding to the second image is higher than the quality of the face image corresponding to the third image.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Identity authentication method and apparatus

    CN108491805A

  • Face recognition method and device, electronic equipment and storage medium

    CN112052789A