Face database construction and updating method and device based on mobile terminal and storage medium
By uploading dynamic facial videos and photos from mobile devices, a 3D feature face is generated and simulated, solving the problems of low recognition accuracy and untimely updates in facial recognition systems in crowded places, and achieving high-precision and real-time updated recognition results.
Patent Information
- Application Number
- CN202211434081.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-11-16
AI Technical Summary
In existing technologies, the static data collection method of facial recognition systems in crowded places such as residential areas, campuses and offices results in low recognition accuracy, cannot adapt to factors such as face occlusion, angle changes and expression changes, and cannot be updated in real time.
By uploading dynamic videos and photos of faces to mobile devices, a 3D feature face is generated. The face is then simulated by combining information such as light intensity, angle, occlusion status, skin color, and age. The face database is updated in real time to improve recognition accuracy and comprehensiveness.
It achieves high-precision face recognition under different lighting, angle, occlusion and skin color conditions, adapts to changes in the face over time, and improves the accuracy and real-time performance of the recognition system.
Smart Images

Figure CN115830672B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to face recognition technology, in particular to a face database building and updating method and device based on a mobile terminal and a storage medium. BACKGROUND
[0002] Since the face recognition technology was introduced in the 1950s, it has mainly experienced three recognition research stages, from the initial research based on the geometric structure features of the face, to the face appearance modeling, to the current research and development of spatial models and factor recognition. Face recognition technology has begun to be widely applied and developed in large areas, such as face verification for bank payment, face unlocking for mobile phone applications, and face access control in communities. It has basically involved all aspects of human daily life.
[0003] In fact, face recognition access control systems have been basically applied and installed in personnel gathering places such as cities, campuses, and office locations. However, due to the gathering of personnel in these places, the collection of face data is a huge workload that cannot be completed in a short period of time, and the face data cannot be updated in real time due to the flow of personnel in these places. Therefore, the face data collection in these personnel gathering places is mostly completed by static collection methods such as photos to build a face data database. This static data database building method also causes the face recognition process to fail due to factors such as occlusion, angle change, and expression change, thereby causing inconvenience.
[0004] The authorized number CN102254154B discloses a face identity authentication method based on three-dimensional model reconstruction. A three-dimensional face model is established based on the support of the three-dimensional deformation-based nonlinear optimization method of an identity card photo, and a three-dimensional face is generated based on a series of algorithms and two-dimensional photos. On the one hand, the three-dimensional face model generated by three-dimensional transformation itself has certain simulation deviation, and on the other hand, the face will change based on age, occlusion, skin color, and various comprehensive factors. Only by generating a three-dimensional face model based on an identity photo, it is impossible to adapt to the changes of the face over time and the influence of factors such as occlusion, thereby causing recognition failure. SUMMARY
[0005] The present application provides a face database building and updating method and device based on a mobile terminal and a storage medium to solve the technical problem of insufficient recognition accuracy and comprehensiveness caused by low photo quality when building a face database using a mobile terminal, which addresses the defects of two-dimensional face recognition distortion and the inability to meet the face recognition needs of concentrated areas such as communities, campuses, and office locations in the prior art.
[0006] In a first aspect, the embodiments of the present application provide a face database building and updating method based on a mobile terminal, the method comprising:
[0007] uploading a dynamic video and face photos of a face through the mobile terminal, wherein the face photos comprise at least a reference photo and a verification photo;
[0008] extracting three-dimensional geometric features of the face from the dynamic video, generating a corresponding three-dimensional feature face according to the three-dimensional geometric features and the reference photo, and storing the three-dimensional feature face into a face database;
[0009] simulating the three-dimensional feature face according to light intensity, angle, occlusion state, skin color and age information of the verification photo, and comparing the verification photo with the simulated three-dimensional feature face;
[0010] when the similarity between the verification photo and the simulated three-dimensional feature face is less than a preset first similarity threshold, generating a new three-dimensional feature face according to the dynamic video and the verification photo, and replacing the three-dimensional feature face stored in the face database;
[0011] when the similarity between the verification photo and the simulated three-dimensional feature face is greater than or equal to the preset first similarity threshold and less than a preset second similarity threshold, generating a new three-dimensional feature face according to the reference photo, the verification photo and the dynamic video, and replacing the three-dimensional feature face stored in the face database; the second similarity threshold is greater than the first similarity threshold;
[0012] when the similarity between the verification photo and the simulated three-dimensional feature face is greater than or equal to the preset second similarity threshold, returning the original three-dimensional feature face.
[0013] Specifically, the embodiment adopts the way of building a three-dimensional feature face to improve the accuracy of face recognition. The face database in places such as communities, campuses and office parks is usually built by uploading photos from mobile terminals, so the face data stored in the database is low-dimensional two-dimensional photo information. When the face changes with light intensity in the recognition process, is blocked by glasses, masks, hair, has different tilt angles, has different skin colors and ages, the two-dimensional photo stored in the database will not be suitable for sudden situations and long-term use, such as changes in women's hairstyles, makeup, wearing glasses or not, wearing a mask or not, and the growth of children. Compared with the real three-dimensional face, the two-dimensional photo itself has a certain degree of distortion, even if the photo is maintained and updated, there will still be problems of recognition accuracy and comprehensiveness, such as unrecognizable and incorrect recognition. The embodiment generates a three-dimensional feature face by using dynamic video and reference photos of the face, and stores the three-dimensional feature face in the face database. The three-dimensional feature face can greatly improve the accuracy of face recognition, and the three-dimensional feature face can obtain the light intensity, angle, blocking state, skin color and age information of the current user when the user passes through the verification device, so as to simulate the three-dimensional feature face in the three-dimensional space according to the current situation of the user, thereby enhancing the consistency of the three-dimensional feature face and the actual face, and improving the accuracy of face recognition. It should be understood that if the photo information obtained by the user during recognition includes {light intensity is 3500Lux, face tilt angle is 35°, hair and wearing a mask, skin color No: 15, age 15 years old}, the processing unit simulates the three-dimensional feature face according to the above information, so that the three-dimensional feature face is in the condition of {light intensity is 3500Lux, face tilt angle is 35°, hair and wearing a mask, skin color No: 15, age 15 years old}, so as to simulate the current face state of the user and improve the accuracy of face recognition. Therefore, users who want to enter or exit the community, campus or office park can directly upload their face information through the mobile terminal to build the face database. The dynamic video includes at least two of the turning head action, blinking action, nodding action and mouth opening action, and the reference photo includes at least one of the front photo and side photo.
[0014] In addition, the application provides a method for updating the face database. Although the method for building the face database can solve the quality problem of the face database of the mobile terminal, the face database cannot be updated as time goes by and the face changes. Therefore, the embodiment further provides a real-time maintenance method for the face database. That is, the embodiment presets a first similarity threshold, which is the minimum acceptable similarity, and a second similarity threshold, which is a distinguishing similarity between similar and extremely similar faces. Extremely similar faces are basically the same as the current user's face.
[0015] Therefore, when the similarity between the corrected photo and the simulated three-dimensional feature face is less than the first preset similarity threshold, a new three-dimensional feature face is generated according to the dynamic video and the corrected photo, and the three-dimensional feature face stored in the face database is replaced. That is, when the corrected photo is obviously changed compared with the three-dimensional feature face in the face database, a new three-dimensional feature face is obtained from the corrected photo and the stored dynamic video, so that the face database is updated in real time, and the three-dimensional feature face in the face database is always in the latest state.
[0016] When the similarity between the corrected photo and the simulated three-dimensional feature face is greater than or equal to the first preset similarity threshold and less than or equal to the second preset similarity threshold, a new three-dimensional feature face is generated according to the reference photo, the corrected photo and the dynamic video, and the three-dimensional feature face stored in the face database is replaced. The second similarity threshold is greater than the first similarity threshold. In fact, the face shape changes slowly with the growth of age, the influence of sunlight and the like, and is not changed suddenly. In order to adapt to the slow change process of the face database, when the current face and the three-dimensional feature face in the face database are within the acceptable similarity but not extremely similar, the three-dimensional feature face in the face database is updated gradually based on the corrected photo, the reference photo and the dynamic video, so as to adapt to the change of the face shape caused by the age change of the user, such as the gradual growth of a child or the aging of a middle-aged person.
[0017] When the similarity between the corrected photo and the simulated three-dimensional feature face is greater than the second preset similarity threshold, the original three-dimensional feature face is returned. This process is used when the current face and the three-dimensional feature face in the face database are extremely similar, and the three-dimensional feature face in the face database is not updated.
[0018] In a first aspect, when the similarity between the corrected photo and the simulated three-dimensional feature face is less than the first preset similarity threshold more than the preset acceptable threshold, the user is prompted to update the dynamic video and the reference photo to generate a new three-dimensional feature face.
[0019] In a further possible implementation method of the first aspect, when comparing the similarity between the proof photo and the simulated three-dimensional feature face, two-dimensional reference geometric features of the face in the proof photo are extracted; the simulated three-dimensional feature face is processed in low dimension and normalization, and two-dimensional reference geometric features of the processed three-dimensional feature face are extracted;
[0020] For the eight feature parts of forehead, eye, ear, cheek, nose, mouth, chin and facial shadow, eight weight values of the first to eighth are set;
[0021] By comparing the similarity of the two-dimensional reference geometric features and the two-dimensional reference geometric features in the above eight feature parts, the similarity between the proof photo and the simulated three-dimensional feature face is calculated according to the eight weight values of the first to eighth.
[0022] Specifically, limited by the recognition duration and shooting accuracy of the access control to the user to be identified, in order to adapt to the planar camera, in order to improve the consistency of the comparison, after the simulation of the three-dimensional feature face, the simulated three-dimensional feature face is low-dimensional and normalized to make the three-dimensional feature face consistent with the format of the comparison photo, thereby improving the accuracy of the comparison. In addition, based on the advantages of the three-dimensional feature face, eight weight values corresponding to the first to eighth are set for the eight feature parts of the forehead, eye, ear, cheek, nose, mouth, jaw and facial shadow. In fact, the influence degree of different parts on the recognition result is different in the face recognition process, so the weight formula is set: S1+S2+S3+S4+S5+S6+S7+S8=1. Among them, S1 is the first weight value, indicating the influence degree of the forehead on the recognition result; S2 is the second weight value, indicating the influence degree of the eye on the recognition result; S3 is the third weight value, indicating the influence degree of the ear on the recognition result; S4 is the fourth weight value, indicating the influence degree of the cheek on the recognition result; S5 is the fifth weight value, indicating the influence degree of the nose on the recognition result; S6 is the sixth weight value, indicating the influence degree of the mouth on the recognition result; S7 is the seventh weight value, indicating the influence degree of the jaw on the recognition result; S8 is the eighth weight value, indicating the influence degree of the facial shadow on the recognition result, so the similarity D=d1*S1+d2*S2+d3*S3+d4*S4+d5*S5+d6*S6+d7*S7+d8*S8. Among them, D represents the similarity of the comparison photo and the simulated three-dimensional feature face, d1 is the similarity of the comparison photo at the forehead part and the simulated three-dimensional feature face, d2 is the similarity of the comparison photo at the eye and the simulated three-dimensional feature face, d3 is the similarity of the comparison photo at the ear and the simulated three-dimensional feature face, d4 is the similarity of the comparison photo at the cheek part and the simulated three-dimensional feature face, d5 is the similarity of the comparison photo at the nose and the simulated three-dimensional feature face, d6 is the similarity of the comparison photo at the mouth and the simulated three-dimensional feature face, d7 is the similarity of the comparison photo at the jaw part and the simulated three-dimensional feature face, and d8 is the similarity of the comparison photo at the facial shadow part and the simulated three-dimensional feature face.
[0023] In still another implementation method of the first aspect, before the three-dimensional feature face is simulated according to the light intensity, angle, occlusion state, skin color and age information of the comparison photo, the quality of the comparison photo is verified according to the light intensity, angle, occlusion state, skin color and age of the comparison photo, comprising:
[0024] According to the light intensity, angle, shielding state, skin color and age of the proofing photo, corresponding light intensity coefficient, angle coefficient, shielding state coefficient, skin color coefficient and age coefficient are set, and the photo quality value of the proofing photo is calculated through the above coefficients;
[0025] When the photo quality value is less than the preset quality threshold, the simulation simulation and subsequent comparison and updating operation of the three-dimensional feature face are terminated;
[0026] When the photo quality value is greater than or equal to the preset quality threshold, the simulation simulation and subsequent comparison and updating operation of the three-dimensional feature face are completed.
[0027] Specifically, the three-dimensional feature face can be updated in real time through the proofing photo, but in order to ensure the quality of the updated three-dimensional feature face, the corresponding coefficients are calculated and obtained according to the light intensity, angle, shielding state, skin color and age of the proofing photo, and the photo quality value is obtained through these coefficients. In fact, only when the light intensity is moderate, a good photo quality can be obtained, and when the light intensity is smaller or larger, the light intensity coefficient will be smaller; similarly, the angle coefficient of the front photo is the largest, and as the left or right angle gradually increases, the angle coefficient gradually decreases; the shielding state coefficient of the unshielded photo is the largest, and as the shielding rate increases, the shielding state coefficient gradually decreases; when the skin color is moderate, the skin color coefficient is the largest, and as the skin color becomes whiter or darker, the skin color coefficient gradually decreases; the age coefficient of the dynamic video and the reference photo is the largest, and as time goes by, the age coefficient gradually decreases. The photo quality value is calculated and obtained through the setting of the above coefficients. For example, the reference photo is {light intensity is 3500Lux, face tilt angle is 0°, no shielding, skin color No: 15, age is 15 years old}, and the corresponding coefficients are {0.9, 1, 1, 0.95, 1}, then the initial photo quality value is Q 初 =0.9*1*1*0.95*1=0.855, if the preset quality threshold Q 预 =0.7, then the initial photo quality value Q 初 >Q 预 , the reference photo is completed, and the three-dimensional feature face is generated; when the information of the proofing photo 1 is {light intensity is 3500Lux, face tilt angle is 35°, hair, skin color No: 20, age is 15 years old}, the corresponding coefficients are {1, 0.9, 0.9, 0.9, 1}, then Q 校1 =1*0.9*0.9*0.9*1=0.729, which is greater than the preset quality threshold Q 预If the information of the calibration photo 2 is {light intensity is 4000 Lux, face tilt angle is 20°, no occlusion, skin color No: 20, age 22 years old}, the corresponding coefficients are {1, 0.95, 1, 0.9, 0.8}, and Q 校1 = 1*0.95*1*0.9*0.8 = 0.684, which is less than the preset quality threshold Q 预 Therefore, the calibration photo 2 cannot be used for simulation and subsequent comparison and updating of the three-dimensional feature face in the face database. Through the embodiment, the quality of the reference photo can be controlled, thereby ensuring the quality of the three-dimensional feature face in the input face database. Meanwhile, the embodiment can also control the quality of the calibration photo, thereby ensuring the updating quality of the three-dimensional feature face in the face database, and further improving the precision of the three-dimensional feature face and the recognition precision and comprehensiveness.
[0028] Further, if the photo quality continuously is less than the preset quality threshold, the user is reminded to update the dynamic video and the face photo.
[0029] In still another implementation method of the first aspect, when generating a new three-dimensional feature face according to the reference photo, the calibration photo and the dynamic video, according to the photo quality value, the reference weight and the calibration weight of the reference photo and the calibration photo for the new three-dimensional feature face are set, and the new three-dimensional feature face is calculated and obtained by using the reference weight and the calibration weight.
[0030] Specifically, in order to cope with the change of the face shape of the user over time, in the case of meeting the minimum similarity and not meeting the extremely high similarity, the reference weight and the calibration weight are determined according to the photo quality value of the reference photo and the calibration photo, thereby adapting to the updating process of the three-dimensional feature face over time.
[0031] In still another implementation method of the first aspect, according to the influence degree of the light intensity coefficient, the angle coefficient, the occlusion state coefficient, the skin color coefficient and the age coefficient on the comparison similarity of the calibration photo and the simulated three-dimensional feature face, the first to eighth, a total of eight weight values are adjusted.
[0032] Specifically, taking {light intensity of 3500 Lux, face tilt angle of 0°, no occlusion, skin color No: 20, age 15} as an example, if the information of the calibration photo of the current user is {light intensity of 2000 Lux, face tilt angle of 35°, wearing a mask, skin color No: 20, age 15}, obviously, if the first value is the eighth, a total of eight weight values still maintain the preset weight, which will cause the similarity to decrease, therefore, when the light intensity of the calibration photo is low, the weight value of the facial shadow should be increased, and when wearing a mask, the weight values of the mouth, cheek, ear and lower jaw decrease, and even become 0, and the weight values of other obtainable parts increase accordingly, so as to improve the accuracy of similarity comparison in different scenarios.
[0033] In still another implementation method of the first aspect, the calibration photo at least includes a photo uploaded by the user through the mobile terminal autonomously or a photo obtained by the user successfully completing face recognition once.
[0034] Specifically, the user can upload a calibration photo according to actual needs, such as the user having makeup, wearing a mask, and having hair, etc. When the user has poor recognition effect in these states, the calibration photo in the corresponding state can be improved, so as to improve the face recognition effect. In addition, in order to improve the similarity between the three-dimensional feature face in the face database and the user, after the user successfully verifies each time, the photo at the time of verification is used to update the three-dimensional feature face of the face database in real time, so as to ensure that the accuracy of the three-dimensional feature face in the face database is always in a highly similar state.
[0035] In still another implementation method of the first aspect, the three-dimensional geometric features and the reference photo are input into a face space training model, and a corresponding three-dimensional feature face is obtained through the face space training model.
[0036] The face space training model is a model obtained by training and learning according to a plurality of three-dimensional geometric feature samples, corresponding reference photo samples and actual faces.
[0037] Specifically, the three-dimensional feature face is obtained through the face space training model, so as to ensure the similarity between the three-dimensional feature face and the actual face of the user.
[0038] In the second aspect, the embodiment of the present application provides a face database building and updating method based on a mobile terminal. In the case that the access control camera can capture user face action video and face photo, the captured face action video and face photo are input into the face space training model for training, so as to realize the updating of the face space training model.
[0039] In a further implementation method of the second aspect, the face space training model can be arranged on a server, and the face action videos and face photos collected by the access control devices are input into the face space training model for training, and the trained face space training model is distributed to the access control terminals of the communities, campuses and office parks.
[0040] In a third aspect, a face database building and updating device based on a mobile terminal includes:
[0041] a transceiving unit configured to acquire dynamic videos and face photos of faces uploaded through the mobile terminal, and input the generated three-dimensional feature faces into the face database
[0042] a generating unit configured to generate corresponding three-dimensional feature faces according to the three-dimensional geometric features and the reference photos and / or the collated photos;
[0043] a processing unit configured to compare the similarity between the collated photos and the three-dimensional feature faces after the simulation, and perform the following operations: when the similarity between the collated photos and the three-dimensional feature faces after the simulation is less than a preset first similarity threshold, generate a new three-dimensional feature face according to the dynamic videos and the collated photos, and replace the three-dimensional feature face stored in the face database; when the similarity between the collated photos and the three-dimensional feature faces after the simulation is greater than or equal to the preset first similarity threshold and less than or equal to a preset second similarity threshold, generate a new three-dimensional feature face according to the reference photos, the collated photos and the dynamic videos, and replace the three-dimensional feature face stored in the face database; and when the similarity between the collated photos is greater than the preset second similarity threshold, return to the original three-dimensional feature face.
[0044] Specifically, the device described above realizes the method described in the first aspect or any possible implementation manner of the first aspect.
[0045] In a fourth aspect, an embodiment of the present application provides a terminal including at least one processor, a communication interface and a memory, the communication interface is configured to send and / or receive data, the memory is configured to store a computer program, and the at least one processor is configured to call the computer program stored in the at least one memory, and the terminal can execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0046] It should be noted that the processor included in the terminal described in the fifth aspect can be a processor specially used for executing the methods (referred to as a special-purpose processor for convenience), or a processor that executes the methods by calling computer programs, such as a general-purpose processor. Optionally, the at least one processor can include both a special-purpose processor and a general-purpose processor.
[0047] Optionally, the computer program can exist on the above-mentioned memory. Exemplarily, the memory can be a non-transitory memory, such as a Read Only Memory (ROM), which can be integrated on the same device with the processor, or can be separately arranged on different devices, and the type of the memory and the arrangement mode of the memory and the processor are not limited in the embodiments of the present application.
[0048] In a possible implementation, the at least one memory is located outside the terminal.
[0049] In another possible implementation, the at least one memory is located inside the terminal.
[0050] In yet another possible implementation, part of the at least one memory is located inside the terminal, and the other part of the at least one memory is located outside the terminal.
[0051] In the present application, the processor and the memory can also be integrated into one device, that is, the processor and the memory can also be integrated together.
[0052] In a fifth aspect, an embodiment of the present application provides a server, which comprises a processor, a memory and a communication interface; the memory stores a computer program; when the processor executes the computer program, the communication interface is used for sending and / or receiving data, and the server can execute the method described in the foregoing second aspect or any possible implementation manner of the second aspect.
[0053] It should be noted that the processor included in the server described in the foregoing fifth aspect can be a processor specially used for executing the method (referred to as a special-purpose processor for the sake of convenience of distinction), or can be a processor executing the method by calling the computer program, such as a general-purpose processor. Optionally, the at least one processor can include both the special-purpose processor and the general-purpose processor.
[0054] Optionally, the computer program can exist on the above-mentioned memory. Exemplarily, the memory can be a non-transitory memory, such as a Read Only Memory (ROM), which can be integrated on the same device with the processor, or can be separately arranged on different devices, and the type of the memory and the arrangement mode of the memory and the processor are not limited in the embodiments of the present application.
[0055] In a possible implementation, the at least one memory is located outside the server.
[0056] In another possible implementation, the at least one memory is located inside the server.
[0057] In yet another possible implementation, part of the at least one memory is located in the server and another part of the at least one memory is located outside the server.
[0058] In the present application, the processor and the memory can also be integrated into one device, i.e., the processor and the memory can also be integrated together.
[0059] In a sixth aspect, an embodiment of the present application provides a computer readable storage medium, wherein a computer program is stored in the computer readable storage medium, and when the instructions are executed on at least one processor, the method described in the foregoing first aspect or any possible implementation of the first aspect or the foregoing second aspect or any possible implementation of the second aspect is implemented.
[0060] In a seventh aspect, the present application provides a computer program product, wherein the computer program product comprises a computer program, and when the program is executed on at least one processor, the method described in the foregoing first aspect or any possible implementation of the first aspect or the foregoing second aspect or any possible implementation of the second aspect is implemented.
[0061] Optionally, the computer program product can be a software installation package, and when the foregoing method needs to be used, the computer program product can be downloaded and executed on a computing device.
[0062] The technical solutions provided by the third to seventh aspects of the present application have the beneficial effects of the technical solutions of the first and second aspects, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0063] The present application will be further described in detail below with reference to the accompanying drawings and preferred embodiments, but those skilled in the art will appreciate that the drawings are only drawn for the purpose of explaining the preferred embodiments and therefore should not be regarded as limiting the scope of the present application. In addition, unless specifically indicated, the drawings only schematically represent the composition or structure of the described objects and can contain exaggerated displays, and the drawings are not necessarily drawn to scale.
[0064] Figure 1 The framework schematic diagram of the face base library construction and updating method based on a mobile terminal is provided by the embodiments of the present application;
[0065] Figure 2 The flow schematic diagram of the face base library construction and updating method based on a mobile terminal is provided by the embodiments of the present application;
[0066] Figure 3 Another flow schematic diagram of the face base library construction and updating method based on a mobile terminal is provided by the embodiments of the present application;
[0067] Figure 4 The device schematic diagram of the mobile terminal-based face database building and updating method provided by the embodiment of the present application is shown in the figure.
[0068] Figure 5 The face space training model provided by the embodiment of the present application is shown in the figure.
[0069] Figure 6 The terminal schematic diagram of the mobile terminal-based face database building and updating provided by the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0070] The present application will be described in detail below with reference to the accompanying drawings and embodiments. Figures 1 to 6 The present application will be described in detail below with reference to the accompanying drawings and embodiments.
[0071] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0072] Please refer to Figure 1 shown in the figure, Figure 1 The architecture schematic diagram of the mobile terminal-based face database building and updating method provided by the embodiment of the present application is shown in the figure.
[0073] Specifically, when arranging the face recognition database, the user can directly shoot dynamic video, reference photos and collation photos through the mobile terminal 101, and upload them to the processing unit 103 through the application installed in the mobile terminal 101. The mobile terminal 101 can be a smart phone, a tablet computer, a computer or other mobile intelligent devices with video input function and camera function. The processing unit 103 extracts the three-dimensional geometric features of the face from the dynamic video, generates the corresponding three-dimensional feature face according to the three-dimensional geometric features and the reference photos, and stores the three-dimensional feature face in the face database 104, thereby completing the building work of the face database. At the same time, the user can complete the construction of the three-dimensional feature face by uploading only the dynamic video and the reference photos. However, in order to ensure that the three-dimensional feature face in the face database 104 can be updated in real time over time, on the one hand, the user can upload the reference photos through the mobile terminal 101, and on the other hand, the recognition device 102 arranged at the entrance of the community, campus and office park can obtain the face photo as the reference photo according to the successful verification of the user once.
[0074] Please refer to Figure 2 shown in the figure, Figure 2 The flowchart of the mobile terminal-based face database building and updating method provided by the embodiment of the present application is shown in the figure.
[0075] Specifically, in the operation process, first, the dynamic video and the face photo of the user are uploaded through the mobile terminal 101 in step S210. The three-dimensional geometric features of the face are extracted from the dynamic video according to step S220, and the corresponding three-dimensional feature face is generated according to the three-dimensional geometric features and the reference photo, and the three-dimensional feature face is stored in the face database. Thus, the operation of inputting the face into the face database is completed. In fact, the concept of building the face database provided by the present application is to establish a three-dimensional model of the face by the action of the dynamic video and the low-dimensional reference photo, thereby improving the similarity and accuracy of the face uploaded to the face database through the mobile terminal 101. However, in order to further improve the quality and real-time performance of the three-dimensional feature face in the face database, and to cope with the recognition misjudgment of users in different situations and different ages, the three-dimensional feature face is simulated according to the light intensity, angle, shielding state, skin color and age information of the correction photo according to step S230, and the similarity between the correction photo and the three-dimensional feature face after simulation is compared. In fact, when maintaining and updating the three-dimensional feature face, considering that the crowd has multiple possible lines when passing through the identification device 102, the light intensity is used to constrain the change of light caused by weather and time, the angle is used to constrain the angle between the face and the device when the user passes through the identification device 102, the shielding state is used to constrain the state of the user wearing a mask, wearing a hat, wearing glasses and the like, the skin color is used to constrain the skin color change caused by the user's makeup and light change, and the age is used to constrain the time when the three-dimensional feature face of the face database 104 is generated, so that the similarity comparison between the correction photo and the three-dimensional feature face after simulation has consistency, thereby improving the accuracy of the similarity comparison. On the basis of the above, whether the similarity between the correction photo and the three-dimensional feature face after simulation is less than a preset first similarity threshold value is judged through step S240 (the definition of each preset value has been indicated in the description, and specific embodiments will not be described again, and can be referred to the foregoing). If the similarity between the two is less than the preset first similarity threshold value, a new three-dimensional feature face is generated according to the dynamic video and the correction photo through step S241, and the three-dimensional feature face stored in the face database is replaced. Through this step, the real-time maintenance of the three-dimensional feature face in the face database 104 can be completed, and the user experience can be improved. It is worth noting that the correction photo is the photo uploaded by the user himself or the photo verified by face recognition. If the similarity between the two is greater than or equal to the preset first similarity threshold value, whether the similarity between the correction photo and the three-dimensional feature face after simulation is less than a preset second similarity threshold value is judged through step S250. If the similarity between the two is less than the preset second similarity threshold value, a new three-dimensional feature face is generated according to the reference photo, the correction photo and the dynamic video through step S251, and the three-dimensional feature face stored in the face database 104 is replaced. Through this judgment process, the three-dimensional feature face stored in the face database 104 can be gradually adjusted under the condition that the user's face changes little, so that it can adapt to the user group with face shape changes such as children and the elderly.If the similarity is greater than or equal to a preset second similarity threshold, it indicates that the face shape of the face has little change, and there is no need to update. Then, the original three-dimensional feature face is returned through step S260.
[0076] In fact, Figure 2 The flowchart of the provided method can complete real-time construction and updating of the three-dimensional feature face of the face database 104, improve the similarity and accuracy of the face uploaded to the face database 104 via the mobile terminal 101 and the actual face, thereby improving the accuracy of face recognition, and coping with changes of the face in different situations, such as dark light, age growth, skin color change, wearing a mask, wearing glasses, and the like.
[0077] In yet another embodiment, the recognition device 102 can obtain the three-dimensional feature face in the simulation state through the action of the processing unit 103 in the recognition process, and compare and analyze the three-dimensional feature face after simulation with the face to be identified and verified, thereby improving the recognition flexibility and recognition accuracy of the recognition device 102.
[0078] Specifically, the recognition device 102 further includes a preset third similarity threshold in the recognition process, which is used to determine whether the face and the three-dimensional coordinate face after simulation are similar to an acceptable minimum threshold, and the preset third similarity threshold is less than the first similarity threshold.
[0079] In yet another embodiment, when the number of times that the similarity between the proof photo and the three-dimensional feature face after simulation is less than the preset first similarity threshold exceeds the preset acceptable threshold, the user is prompted to update the dynamic video and the reference photo, and the three-dimensional feature face is regenerated.
[0080] In yet another embodiment, when comparing the similarity between the proof photo and the three-dimensional feature face after simulation, the two-dimensional reference geometric features of the face in the proof photo are extracted; the three-dimensional feature face after simulation is processed in low dimension and normalized, and the two-dimensional reference geometric features of the processed three-dimensional feature face are extracted; for the eight feature parts of the forehead, eye, ear, cheek, nose, mouth, lower jaw, and facial shadow, the first to eighth, a total of eight weight values are set; the similarity between the two-dimensional reference geometric features and the two-dimensional reference geometric features in the above eight feature parts is compared, and the similarity between the proof photo and the three-dimensional feature face after simulation is calculated according to the first to eighth, a total of eight weight values.
[0081] Specifically, limited by the recognition duration and shooting accuracy of the to-be-identified user in the identification of the access control, in order to adapt to the planar camera, in order to improve the consistency of the comparison, after the simulation of the three-dimensional feature face is completed, the three-dimensional feature face after the simulation is subjected to low-dimensional and normalization processing, so that the three-dimensional feature face is consistent with the format of the comparison photo, thereby improving the accuracy of the comparison. In addition, based on the advantages of the three-dimensional feature face, eight weight values corresponding to the first to eighth are set for the eight feature parts of the forehead, eye, ear, cheek, nose, mouth, lower jaw and facial shadow. In fact, in the identification process, different parts of the face have different influences on the identification result, so the weight formula is set as S1+S2+S3+S4+S5+S6+S7+S8=1, and the similarity D=d1*S1+d2*S2+d3*S3+d4*S4+d5*S5+d6*S6+d7*S7+d8*S8.
[0082] Please refer to Figure 3 , as shown in Figure 3 The embodiment of the application provides another flowchart of the face database construction and updating method based on a mobile terminal.
[0083] Specifically, in Figure 2 the flowchart, before step S230 is executed, it is judged whether the photo quality value is less than the preset quality threshold value through step S320. If the photo quality value is less than the preset quality threshold value, the comparison and updating operation of the three-dimensional feature face is terminated through step S321. In fact, light intensity, angle, shielding state, skin color and age can affect the feature extraction of the three-dimensional geometric features of the face. At the same time, when the light is insufficient, the face angle is deflected too large, the shielding is too much, the skin color is too dark or the age is increased, the face features will be blurred and unclear. If such a photo is used for three-dimensional feature face updating, the similarity between the three-dimensional feature face and the actual face will be reduced, thereby affecting the identification accuracy. Therefore, in order to improve the situation that the three-dimensional feature face is in high similarity with the face in real time, the comparison photo with high quality is needed to modify the three-dimensional feature face. In addition, age actually has the greatest influence on the face shape. In order to ensure the similarity of the three-dimensional feature face, age is taken as a parameter. With the increase of the age of the user, the age coefficient will become smaller and smaller, so that the user can be notified in time to update the dynamic video and the reference photo stored before through this step, and the three-dimensional feature face in the face database 104 is corrected when the device usage time is too long due to the change of the face of the user.
[0084] Meanwhile, the corresponding coefficients are calculated according to the light intensity, angle, occlusion state, skin color and age included in the comparison photo, and the photo quality value is obtained through the coefficients. In fact, a better photo quality is obtained when the light intensity is moderate, and the light intensity coefficient will become smaller when the light intensity is smaller or larger; similarly, the angle coefficient of the front photo is the largest, and the angle coefficient gradually decreases as the left or right angle gradually increases; the occlusion state coefficient of the non-occluded photo is the largest, and the occlusion state coefficient gradually decreases as the occlusion rate increases; the skin color coefficient is the largest when the skin color is moderate, and the skin color coefficient gradually decreases as the skin color becomes whiter or darker; the age coefficient of the dynamic video and the input age of the reference photo is the largest, and the age coefficient gradually decreases over time. The photo quality value is calculated through the setting of the above coefficients. For example, the reference photo is {light intensity is 3500Lux, face tilt angle is 0°, no occlusion, skin color No: 15, age 15 years old}, and the corresponding coefficients are {0.9, 1, 1, 0.95, 1}, then the initial photo quality value is Q 初 = 0.9*1*1*0.95*1 = 0.855, if the preset quality threshold Q 预 = 0.7, then the initial photo quality value Q 初 > Q 预 , the input of the reference photo is completed, and a three-dimensional feature face is generated; when the information of the comparison photo 1 is {light intensity is 3500Lux, face tilt angle is 35°, hair, skin color No: 20, age 15 years old}, the corresponding coefficients are {1, 0.9, 0.9, 0.9, 1}, then Q 校1 = 1*0.9*0.9*0.9*1 = 0.729, which is greater than the preset quality threshold Q 预 , the comparison photo can be used for simulation and subsequent comparison and updating operation of the three-dimensional feature face in the face database; when the information of the comparison photo 2 is {light intensity is 4000Lux, face tilt angle is 20°, no occlusion, skin color No: 20, age 22 years old}, the corresponding coefficients are {1, 0.95, 1, 0.9, 0.8}, then Q 校1 = 1*0.95*1*0.9*0.8 = 0.684, which is less than the preset quality threshold Q 预 , the comparison photo 2 cannot be used for simulation and subsequent comparison and updating operation of the three-dimensional feature face in the face database. Through the embodiment, the quality of the reference photo can be controlled, so as to ensure the quality of the three-dimensional feature face input into the face database. Meanwhile, the quality of the comparison photo can also be controlled, so as to ensure the updating quality of the three-dimensional feature face in the face database, and further improve the precision, recognition precision and comprehensiveness of the three-dimensional feature face.
[0085] In still another implementation method, the three-dimensional geometric features and the reference photo are input into a face space training model, and a corresponding three-dimensional feature face is obtained through the face space training model;
[0086] The face space training model is a model obtained by training and learning according to a plurality of three-dimensional geometric feature samples and corresponding reference photo samples and actual face modeling.
[0087] Specifically, the three-dimensional feature face is obtained through the face space training model, so as to ensure the similarity between the three-dimensional feature face and the actual face of the user.
[0088] Referring to Figure 4 as shown, Figure 4 The device schematic diagram of the face base library construction and updating method based on a mobile terminal is provided.
[0089] Specifically, the face base library construction and updating device 40 comprises a transceiving unit 401, a generating unit 402 and a processing unit 403.
[0090] The transceiving unit 401 is configured to acquire a dynamic video and a face photo uploaded through a mobile terminal, and input the generated three-dimensional feature face into a face base library.
[0091] The generating unit 402 is configured to generate a corresponding three-dimensional feature face according to the three-dimensional geometric feature and the reference photo and / or the proofreading photo.
[0092] The processing unit 403 is configured to compare the similarity between the proofreading photo and the three-dimensional feature face after simulation and emulation, and perform the following operations: when the similarity between the proofreading photo and the three-dimensional feature face after simulation and emulation is less than a preset first similarity threshold, a new three-dimensional feature face is generated according to the dynamic video and the proofreading photo, and the three-dimensional feature face stored in the face base library is replaced; when the similarity between the proofreading photo and the three-dimensional feature face after simulation and emulation is greater than or equal to the preset first similarity threshold and less than or equal to a preset second similarity threshold, a new three-dimensional feature face is generated according to the reference photo, the proofreading photo and the dynamic video, and the three-dimensional feature face stored in the face base library is replaced; and when the similarity between the proofreading photo is greater than the preset second similarity threshold, the original three-dimensional feature face is returned.
[0093] Specifically, the device realizes the method described in the foregoing first aspect or any possible implementation manner of the first aspect.
[0094] Referring to Figure 5 as shown, Figure 5 The face space training model is provided.
[0095] Specifically, the face space training model comprises a first acquiring unit 501, a first training unit 502 and a sending unit 503.
[0096] The first obtaining unit 501 is configured to obtain a dynamic video of a face, a corresponding face photo, and a corresponding three-dimensional face model, and a dynamic video sample, a corresponding face photo sample, and a corresponding three-dimensional face model;
[0097] The first training unit 502 is configured to train a learning method for obtaining a three-dimensional feature face according to the dynamic video of the face, the corresponding face photo, and the corresponding three-dimensional face model, and the dynamic video sample, the corresponding face photo sample, and the corresponding three-dimensional face model.
[0098] The sending unit 503 is configured to send the face space training model after training.
[0099] In fact, the training model is placed in a server, and is configured to send the face space training model to the processing unit 103.
[0100] In yet another embodiment, the processing unit 103 comprises a second obtaining unit and a second training unit. The second obtaining unit is configured to obtain a dynamic video of a face, a corresponding face photo, and a corresponding three-dimensional face model, and a dynamic video sample, a corresponding face photo sample, and a corresponding three-dimensional face model. The second training unit is configured to train a learning method for obtaining a three-dimensional feature face according to the dynamic video of the face, the corresponding face photo, and the corresponding three-dimensional face model, and the dynamic video sample, the corresponding face photo sample, and the corresponding three-dimensional face model. Thus, the processing unit 103 itself can train a learning method for obtaining a face space training model.
[0101] Please refer to Figure 6 as shown, Figure 6 The terminal schematic diagram for building and updating a face database based on a mobile terminal is provided by the embodiment of the present application.
[0102] Specifically, the terminal 60 comprises a communication interface 601, a processor 602, and a memory 603. The processor 601, the communication interface 602, and the memory 603 can be connected through a bus or other means, and the embodiment of the present application takes the connection through a bus as an example.
[0103] The processor 601 is a computing core and a control core of the terminal 60, which can analyze various instructions in the terminal 60 and various data of the terminal 60, for example, the processor 601 can be a central processing unit (CPU), which can transmit various interactive data between internal structures of the terminal 60, and the like. The communication interface 602 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, and the like), and can be used for receiving and transmitting data under the control of the processor 601; the communication interface 602 can also be used for transmitting and interacting internal signaling or instructions of the terminal 60. The memory 603 is a memory device in the terminal 60, which is used for storing programs and data. It can be understood that the memory 603 can include a built-in memory of the terminal 60, and of course can also include an expansion memory supported by the terminal 60. The memory 603 provides a storage space, which stores an operating system of the terminal 60, and also stores program codes or instructions required by the processor to perform corresponding operations, and optionally, the storage space can also store related data generated after the processor performs the corresponding operations.
[0104] In the embodiment of the application, the processor 601 runs executable program codes in the memory 603, and is used for performing the following operations:
[0105] The dynamic video and the face photo of the mobile terminal are uploaded, and the face photo at least includes a reference photo and a verification photo;
[0106] Three-dimensional geometric features of the face are extracted from the dynamic video, a corresponding three-dimensional feature face is generated according to the three-dimensional geometric features and the reference photo, and the three-dimensional feature face is stored in a face database;
[0107] The three-dimensional feature face is simulated according to light intensity, angle, shielding state, skin color and age information of the verification photo, and similarity between the verification photo and the three-dimensional feature face after simulation is compared;
[0108] When the similarity between the verification photo and the three-dimensional feature face after simulation is less than a preset first similarity threshold, a new three-dimensional feature face is generated according to the dynamic video and the verification photo, and the three-dimensional feature face stored in the face database is replaced;
[0109] When the similarity between the verification photo and the three-dimensional feature face after simulation is greater than or equal to a preset first similarity threshold and less than a preset second similarity threshold, a new three-dimensional feature face is generated according to the reference photo, the verification photo and the dynamic video, and the three-dimensional feature face stored in the face database is replaced; the second similarity threshold is greater than the first similarity threshold;
[0110] When the collation photo is greater than or equal to a preset second similarity threshold, the original three-dimensional feature face is returned.
[0111] It should be noted that the implementation of each operation can also correspond to the description of the method embodiments of the terminal side shown in Figure 2 and Figure 3 .
[0112] The embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program includes program instructions, the computer program when being executed by a processor makes the processor realize Figure 2 and Figure 3 The operation of the terminal in the embodiment, or realize Figure 2 and Figure 3 The operation of the server in the embodiment.
[0113] The embodiment of the present application further provides a computer program product, when the computer program product runs on the processor, realizes Figure 2 and Figure 3 The operation of the terminal in the embodiment, or realize Figure 2 and Figure 5 The operation of the server in the embodiment.
[0114] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by programs instructing relevant hardware, and the programs can be stored in a computer readable storage medium, and the programs can include the processes of the above-mentioned embodiments when being executed. The storage medium includes ROM, RAM, magnetic disc or optical disc and various program code storage media.
[0115] The present application is described in detail above, and the principle and implementation mode of the present application are described by using specific examples. The above embodiment is only used to help understand the present application and core idea. It should be pointed out that, for ordinary skilled in the art, without departing from the principle of the present application, the present application can be improved and modified in several ways, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for building and updating a face database based on a mobile terminal, characterized in that, The method includes: Upload dynamic videos and photos of faces via mobile devices, wherein the face photos include at least a reference photo and a verification photo; The three-dimensional geometric features of the face are extracted from the dynamic video, and a corresponding three-dimensional feature face is generated based on the three-dimensional geometric features and the reference photo. The three-dimensional feature face is then stored in the face database. The three-dimensional feature face is simulated based on the light intensity, angle, occlusion state, skin color and age information of the calibration photo, and the similarity between the calibration photo and the simulated three-dimensional feature face is compared. When the similarity between the proofread photo and the simulated 3D feature face is less than a preset first similarity threshold, a new 3D feature face is generated based on the dynamic video and the proofread photo, and the 3D feature face stored in the face database is replaced. When the similarity between the proofread photo and the simulated 3D feature face is greater than or equal to a preset first similarity threshold and less than a preset second similarity threshold, a new 3D feature face is generated based on the reference photo, proofread photo, and dynamic video, and the 3D feature face stored in the face database is replaced; the second similarity threshold is greater than the first similarity threshold. When the proofread photo is greater than or equal to a preset second similarity threshold, the original 3D feature face is returned; When comparing the similarity between the proofed photo and the simulated 3D eigenface, the process specifically includes: extracting the two-dimensional reference geometric features of the face in the proofed photo; performing low-dimensional and normalization processing on the simulated 3D eigenface, and extracting the two-dimensional reference geometric features of the processed 3D eigenface. For eight feature areas—forehead, eyes, ears, cheeks, nose, mouth, jaw, and facial shadows—corresponding weight values are assigned from first to eighth, for a total of eight values. By comparing the similarity between the two-dimensional reference geometric features and the two-dimensional baseline geometric features at the above eight feature locations, the similarity between the proofread photo and the simulated three-dimensional feature face is calculated according to the first to the eighth weight values, a total of eight. Before simulating the three-dimensional feature face based on the light intensity, angle, occlusion state, skin color, and age information of the calibration photo, the quality of the calibration photo is verified based on the light intensity, angle, occlusion state, skin color, and age information of the calibration photo, including: Based on the light intensity, angle, occlusion status, skin tone, and age of the proofread photo, set corresponding light intensity coefficients, angle coefficients, occlusion status coefficients, skin tone coefficients, and age coefficients, and calculate the photo quality value of the proofread photo using the above coefficients; When the image quality value is less than a preset quality threshold, the simulation of the three-dimensional feature face and subsequent comparison and update operations are terminated. When the image quality value is greater than or equal to a preset quality threshold, the simulation of the three-dimensional feature face and subsequent comparison and update operations are completed.
2. The method for building and updating a face database based on a mobile terminal as described in claim 1, characterized in that, When generating a new 3D eigenface based on the reference photo, the calibration photo, and the dynamic video, the specific steps include: Based on the image quality value, set the baseline weight and calibration weight of the baseline image and calibration image for the new 3D eigenface, and use the baseline weight and reference weight to calculate and obtain the new 3D eigenface.
3. The method for building and updating a face database based on a mobile terminal as described in claim 1, characterized in that, Based on the influence of the light intensity coefficient, angle coefficient, occlusion state coefficient, skin color coefficient, and age coefficient on the similarity between the proofread photo and the simulated 3D feature face, the eight weight values (first to eighth) are adjusted.
4. The method for building and updating a face database based on a mobile terminal as described in any one of claims 1-3, characterized in that, The proofread photos include at least photos uploaded by the user via mobile device or photos obtained by the user after successfully completing a facial recognition scan.
5. The method for building and updating a face database based on a mobile terminal as described in any one of claims 1-3, characterized in that, Generating a corresponding 3D feature face based on the aforementioned 3D geometric features and reference photograph specifically includes: The three-dimensional geometric features and the reference photo are input into the face space training model, and the corresponding three-dimensional feature face is obtained through the face space training model; The face space training model is a model obtained by training and learning based on multiple three-dimensional geometric feature samples, corresponding benchmark photo samples, and actual face modeling.
6. A device for building and updating a face database based on a mobile terminal, based on the face database building and updating method according to any one of claims 1-5, characterized in that, The device includes: The transceiver unit is used to acquire dynamic videos and facial photos uploaded by mobile devices, and to input the generated 3D feature face into the facial database. The generation unit generates corresponding 3D feature faces based on 3D geometric features and reference photos, and / or, by calibrating the photos; The processing unit is used to compare the similarity between the proofreading photo and the simulated 3D eigenface; simultaneously, it performs the following operations: when the similarity between the proofreading photo and the simulated 3D eigenface is less than a preset first similarity threshold, a new 3D eigenface is generated based on the dynamic video and the proofreading photo, and the 3D eigenface stored in the face database is replaced; when the similarity between the proofreading photo and the simulated 3D eigenface is greater than or equal to the preset first similarity threshold and less than or equal to a preset second similarity threshold, a new 3D eigenface is generated based on the reference photo, the proofreading photo, and the dynamic video, and the 3D eigenface stored in the face database is replaced; when the similarity between the proofreading photo and the simulated 3D eigenface is greater than the preset second similarity threshold, the original 3D eigenface is returned.
7. A terminal, characterized in that, The terminal includes at least one processor, a communication interface, and a memory. The communication interface is used to send and / or receive data, the memory is used to store computer programs, and the at least one processor is used to call at least one computer program stored in the memory to implement the method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a processor, implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Method for authenticating human-face identity based on three-dimensional model reconstruction
CN102254154B
3D face recognition method, terminal device, and readable storage medium
CN107622227A
Face recognition system and method for enhancing face recognition
CN110580435A
Face optimization recognition method and terminal
CN114241584A
Face recognition method and device and terminal equipment
CN114511914A