Identity verification method and device, computer device, and storage medium
Through the voice enhancement, illumination-invariant features and facial feature extraction of smart glasses, the accuracy problem of identity authentication using smart glasses in bank branch environments was solved, and stable recognition was achieved under noisy conditions and illumination changes.
Patent Information
- Application Number
- CN202411237426.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-09-04
AI Technical Summary
When using smart glasses for self-service cash withdrawals at bank branches, facial recognition accuracy is affected by wearing stability, background noise, and lighting conditions, resulting in inaccurate identity authentication.
The user's voice and face image are obtained through smart glasses, and voice enhancement processing, illumination-invariant feature extraction and facial feature extraction are performed to comprehensively generate identity authentication results.
It improves the robustness of voiceprint recognition in noisy environments, overcomes the impact of lighting changes, improves the accuracy and security of identity authentication, and reduces the risk of misidentification and deceptive attacks.
Smart Images

Figure CN119312308B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to an identity verification method and device, a computer device and a storage medium. BACKGROUND
[0002] When self-service withdrawing at a bank outlet, smart glasses have a significant advantage as an identity verification tool. Smart glasses can capture real-time facial images and voice through the built-in camera and microphone, improving the convenience and security of the withdrawal process. Compared with traditional cash machines, smart glasses provide a more intuitive interaction mode, and users can complete the operation without touching the screen.
[0003] However, in actual application, smart glasses also face many challenges. First, the wearing stability of smart glasses affects the accuracy of facial recognition. If the glasses slip or shift, it may lead to inaccurate capture of facial features, thereby affecting identity verification. The user's facial expression, angle and occlusion may also affect the accuracy of facial recognition. Second, the bank outlet environment is noisy, and background noise can seriously affect the performance of voice recognition. In addition, the lighting conditions at the bank outlet are complex and variable, which poses a challenge to the camera of smart glasses to capture clear facial images. All of these will affect the accuracy of identity verification based on smart glasses. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide an identity verification method, device, computer device and storage medium to solve the accuracy of identity verification based on smart glasses.
[0005] To solve the above technical problems, the embodiments of the present application provide an identity verification method, which adopts the technical scheme as follows:
[0006] Obtain the user's voice and face image of the user through the smart glasses;
[0007] Perform voice enhancement processing on the user voice to obtain an enhanced voice, and perform identity verification on the user according to the enhanced voice to obtain a first verification result;
[0008] Input the face image into a light-invariant feature extraction model to obtain a light-invariant feature of the face image, and perform identity verification on the user according to the light-invariant feature to obtain a second verification result;
[0009] Extract the facial features of the face image, and perform identity verification on the user according to the facial features to obtain a third verification result;
[0010] Generate the identity verification result of the user according to the first verification result, the second verification result and the third verification result.
[0011] To solve the above technical problems, the embodiment of the application further provides an identity verification device, which adopts the technical scheme as follows:
[0012] The user acquisition module is configured to acquire user speech and a face image of a user through smart glasses.
[0013] The speech verification module is configured to perform speech enhancement processing on the user speech to obtain enhanced speech, and perform identity verification on the user according to the enhanced speech to obtain a first verification result.
[0014] The illumination verification module is configured to input the face image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the face image, and perform identity verification on the user according to the illumination-invariant features to obtain a second verification result.
[0015] The face verification module is configured to extract face features of the face image, and perform identity verification on the user according to the face features to obtain a third verification result.
[0016] The result generation module is configured to generate an identity verification result of the user according to the first verification result, the second verification result and the third verification result.
[0017] To solve the above technical problems, the embodiment of the application further provides a computer device, which adopts the technical scheme as follows:
[0018] The user speech and the face image of the user are acquired through smart glasses.
[0019] The speech verification module is configured to perform speech enhancement processing on the user speech to obtain enhanced speech, and perform identity verification on the user according to the enhanced speech to obtain a first verification result.
[0020] The illumination verification module is configured to input the face image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the face image, and perform identity verification on the user according to the illumination-invariant features to obtain a second verification result.
[0021] The face verification module is configured to extract face features of the face image, and perform identity verification on the user according to the face features to obtain a third verification result.
[0022] The result generation module is configured to generate an identity verification result of the user according to the first verification result, the second verification result and the third verification result.
[0023] To solve the above technical problems, the embodiment of the application further provides a computer readable storage medium, which adopts the technical scheme as follows:
[0024] Acquire the user's voice and facial image through the smart glasses;
[0025] performing voice enhancement processing on the user's voice to obtain enhanced voice, and performing identity authentication on the user based on the enhanced voice to obtain a first authentication result;
[0026] Inputting the facial image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, and performing identity verification on the user based on the illumination-invariant features to obtain a second verification result;
[0027] extracting facial features from the face image, and performing identity verification on the user based on the facial features to obtain a third verification result;
[0028] An identity authentication result of the user is generated according to the first verification result, the second verification result and the third verification result.
[0029] Compared with the prior art, the embodiments of the present application have the following main beneficial effects: obtaining the user's voice and facial image through smart glasses; performing voice enhancement processing on the user's voice to obtain enhanced voice, thereby improving the robustness of voiceprint recognition, accurately identifying the user's identity even in a noisy environment, and generating a first verification result; inputting the facial image into the illumination-invariant feature extraction model to obtain the illumination-invariant feature of the facial image, which overcomes the influence of ambient lighting changes on facial recognition, ensures recognition stability under different lighting conditions, and ensures the accuracy of the second verification result; extracting facial features of the facial image, improving the accuracy of identity recognition under complex facial postures and expression changes, and obtaining a third verification result; combining the first verification result, the second verification result, and the third verification result to generate the user's final identity authentication result, reducing the risk of misidentification and deceptive attacks, and improving the accuracy of identity authentication based on smart glasses. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0032] Figure 2 A flowchart of an embodiment of an identity authentication method according to the present application;
[0033] Figure 3is a schematic structural diagram of an embodiment of an identity verification device according to the present application;
[0034] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0036] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0037] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0038] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables. In this application, terminal device 101 may also be smart glasses.
[0039] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0040] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012 or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0041] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0042] It should be noted that the identity authentication method provided in the embodiment of the present application is generally executed by a server, and accordingly, the identity authentication device is generally set in the server.
[0043] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0044] Continue to refer Figure 2 , shows a flow chart of an embodiment of an identity authentication method according to the present application. The identity authentication method comprises the following steps:
[0045] Step S201: Acquire the user's voice and face image through the smart glasses.
[0046] In this embodiment, the electronic device on which the identity verification method is executed (eg Figure 1 The server shown in the figure can communicate with the terminal device through a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, Wi-Fi connection, Bluetooth connection, Wi MAX connection, Zigbee connection, UWB (Ultra Wide Band) connection, and other wireless connection methods currently known or to be developed in the future.
[0047] Specifically, when a user performs operations such as withdrawing cash at a bank branch, user identity verification is required. The user can wear smart glasses, which capture the user's voice and face image.
[0048] The smart glasses can send the user voice and face image to a background server of the smart glasses, and the background server sends the user voice and face image to a server of the bank. Alternatively, the smart glasses send the user voice and face image to a terminal device of a bank outlet, and the terminal device sends the user voice and face image to the server of the bank. The smart glasses can also directly send the user voice and face image to the server of the bank.
[0049] In step S202, the user voice is subjected to speech enhancement processing to obtain an enhanced voice, and the user is authenticated based on the enhanced voice to obtain a first authentication result.
[0050] Specifically, the user voice is subjected to speech enhancement processing to eliminate background noise and other interference, ensuring the intelligibility and integrity of the voice. The enhanced voice can be used for voiceprint recognition, extracting voiceprint features from the voice and comparing them with a voiceprint feature library to obtain the first authentication result.
[0051] In step S203, the face image is input into a light-invariant feature extraction model to obtain light-invariant features of the face image, and the user is authenticated based on the light-invariant features to obtain a second authentication result.
[0052] Specifically, the face image is input into a light-invariant feature extraction model. The model adjusts the light influence in the face image based on the detected light conditions, thereby extracting light-invariant features in the face. The light-invariant features are not affected by changes in environmental lighting, allowing identity verification to remain stable and accurate under various lighting conditions.
[0053] The light-invariant features are matched with library-invariant features in a pre-established light-invariant feature library to calculate feature similarity. The identity authentication result of the light-invariant feature dimension is determined based on the feature similarity. If there is a feature similarity greater than a set threshold, the identity authentication is successful, otherwise the identity authentication fails. By comparing the light-invariant features, a second authentication result can be generated.
[0054] In step S204, facial features of the face image are extracted, and the user is authenticated based on the facial features to obtain a third authentication result.
[0055] Specifically, the facial features of the face image are extracted. This process may involve standardization of two-dimensional images, three-dimensional reconstruction, pose estimation and pose compensation, and other processing steps, ultimately obtaining the user's unique facial features. Based on the facial features, identity authentication in the third dimension is performed, and a third authentication result is generated. This step improves the accuracy of verification through multi-dimensional facial features such as three-dimensional geometric structure, texture and stereo features.
[0056] Step S205, generating the identity verification result of the user according to the first verification result, the second verification result and the third verification result.
[0057] Specifically, the first verification result, the second verification result and the third verification result are comprehensively analyzed to generate the final identity verification result. Generally, if the three verification results are all passed, the identity verification result of the user identity verification is generated.
[0058] This multi-modal and multi-stage identity verification method improves the accuracy and security of identity verification, preventing false positives or attacks that may occur in a single verification method.
[0059] In this embodiment, the user's voice and face image of the user are obtained through the smart glasses; the voice enhancement processing is performed on the user's voice to obtain the enhanced voice, which improves the robustness of voiceprint recognition, and can accurately identify the user's identity in a noisy environment and generate the first verification result; the face image is input into the illumination invariant feature extraction model to obtain the illumination invariant feature of the face image, which overcomes the influence of environmental illumination change on face recognition, ensures the recognition stability under different illumination conditions, and ensures the accuracy of the second verification result; the face features of the face image are extracted to improve the identity recognition accuracy under complex face poses and expression changes, and the third verification result is obtained; the first verification result, the second verification result and the third verification result are comprehensively analyzed to generate the final identity verification result of the user, which reduces the risk of false recognition and fraudulent attacks, and improves the accuracy of identity verification based on smart glasses.
[0060] Further, the above step S202 can include: preprocessing the user's voice to obtain a first voice; extracting frequency domain features of the first voice through a convolutional neural network; performing time domain modeling on the frequency domain features through a gated recurrent unit to obtain an enhanced voice; extracting voiceprint features of the enhanced voice through a voiceprint recognition model; comparing the voiceprint features with each library voiceprint in a voiceprint feature library to verify the identity of the user and obtain the first verification result.
[0061] Specifically, the user's voice is first preprocessed, including removing noise, normalizing volume and filtering unnecessary frequency bands to ensure the quality and consistency of the input voice. The first voice is obtained after preprocessing.
[0062] The first voice is input into a convolutional neural network (CNN) to extract frequency domain features of the voice signal. CNN captures the potential and representative frequency patterns in the voice signal through layer-by-layer convolution and pooling operations. These frequency domain features are very important for voice enhancement and voiceprint recognition. They effectively compress the frequency information of the voice and retain the parts that are recognizable for user identity.
[0063] Frequency-domain features are fed into a Gated Recurrent Unit (GRU) network for time-domain modeling. A GRU is a recurrent neural network (RNN) capable of processing sequential data and capturing temporal dependencies in speech signals. Through time-domain modeling, the temporal dynamics of the speech signal are combined with frequency-domain features to generate enhanced speech. This enhanced speech is optimized at the frequency level, resulting in greater stability and accuracy in terms of temporal continuity.
[0064] The enhanced speech is input into the voiceprint recognition model to extract its voiceprint features. Voiceprint features are unique attributes of the user's voice, including timbre, pronunciation habits, vocal tract shape, etc. These features are highly discriminatory in biometric recognition.
[0065] The voiceprint feature is compared with each stored voiceprint in the voiceprint feature library, for example, by calculating the similarity score between the voiceprint feature and each stored voiceprint using cosine similarity or Euclidean distance. If the similarity score exceeds a set threshold, verification succeeds, confirming that the user's voice is from a registered user. If the similarity score is below the threshold, verification fails. After the comparison, a first verification result is generated, indicating the result of the voice identity verification.
[0066] In this embodiment, preprocessing eliminates environmental noise and inconsistencies in voice collection, ensuring the high quality of the input voice; a convolutional neural network is used to extract the frequency domain features of the voice, capturing the most distinctive frequency information in the voice; time domain modeling is performed through a gated recurrent unit to process the temporal dynamic characteristics of the voice, enhance the continuity and authenticity of the voice, and improve the effect of voice enhancement; by extracting voiceprint features and comparing them with the various stored voiceprints in the voiceprint feature library, the user's identity can be quickly and accurately identified, maintaining efficient and accurate recognition capabilities in noisy or complex voice environments.
[0067] Furthermore, the above-mentioned step of inputting the facial image into the illumination-invariant feature extraction model to obtain the illumination-invariant features of the facial image may include: performing illumination detection on the facial image to obtain the illumination conditions of the facial image; inputting the facial image into the illumination-invariant feature extraction model to perform feature transformation on the facial image according to the illumination conditions through the illumination-invariant feature extraction model to obtain a first facial image; and performing feature extraction on the first facial image through the illumination-invariant feature extraction model to obtain illumination-invariant features.
[0068] Specifically, facial images may be affected by lighting conditions, which may manifest as uneven lighting or shadow areas. Image processing technology is used to perform lighting detection on facial images to detect their lighting conditions, for example, by calculating the brightness histogram of the image, detecting the light intensity and the position of the light source, etc. In one embodiment, the lighting conditions may include uniform lighting (the light distribution is relatively uniform and will not produce obvious shadows or bright spots on the face), strong light (direct illumination by the light source may cause facial highlight areas such as the forehead and the tip of the nose to be overexposed), shadows (improper light source position, resulting in obvious shadow areas, such as under the eye sockets) and backlight (the light source is located behind the shooting angle, resulting in insufficient lighting of facial features).
[0069] Different lighting conditions will cause the visual features of facial images to change, so different methods are needed to adjust these changes to ensure the stability and consistency of facial features.
[0070] The facial image is input into the illumination-invariant feature extraction model. The model transforms the image features by adjusting the image's brightness, contrast, and color balance. Depending on the lighting conditions, the model adopts different transformation strategies. For example, under uniform lighting conditions, additional illumination compensation is generally not required, and standard feature extraction algorithms (such as principal component analysis) are directly used to extract facial features. Under strong lighting conditions, image enhancement techniques (such as histogram equalization and gamma correction) are used to adjust overexposed areas and reduce the impact of highlights to balance the brightness distribution. Under shadowed conditions, shadow compensation techniques (such as local contrast enhancement and shadow compensation filtering) are applied to reduce the impact of shadow areas on facial features. Under backlit conditions, backlight compensation techniques (such as backlight correction and image compensation) are used to increase the brightness and visibility of the facial area, adjusting the brightness of the facial area in the image to a reasonable range for feature extraction. Through these feature transformations, a first facial image with uniform illumination is generated, making facial features more prominent and stable.
[0071] The illumination-invariant feature extraction model is used to further extract features from the first face image to obtain illumination-invariant features. The information reflected by the illumination-invariant features can remain consistent under various lighting conditions and can be used for identity authentication to ensure accurate user identification under different lighting conditions.
[0072] In this embodiment, illumination detection is performed on the facial image to identify the influence of the current lighting environment, providing a basis for subsequent image processing; the image is input into the illumination-invariant feature extraction model, and appropriate image transformation techniques are applied according to the illumination conditions to eliminate interference caused by illumination changes, thereby ensuring that the extracted facial features remain highly stable in various subsequent illumination environments; feature extraction is performed on the first facial image to obtain illumination-invariant features, which are consistent under different illumination conditions, can improve the accuracy of identity verification, and reduce recognition errors caused by illumination changes.
[0073] Furthermore, the above-mentioned step of extracting features from the first facial image through the illumination-invariant feature extraction model to obtain illumination-invariant features may include: extracting facial key points and geometric features of the first facial image through the illumination-invariant feature extraction model; converting the facial key points and geometric features into feature vectors to obtain illumination-invariant features.
[0074] Specifically, the illumination-invariant feature extraction model is used to extract facial key points from the first face image, including prominent facial feature points such as the corners of the eyes, the tip of the nose, and the corners of the mouth. The positions of these key points are relatively fixed and will not shift significantly due to changes in illumination. They can remain consistent under different lighting conditions.
[0075] Based on the facial key points, we further extract geometric features, such as the spatial relationship and distance between facial key points, such as the distance between the eyes and the distance from the nose tip to the corner of the mouth. These geometric features are also not easily affected by changes in lighting and can stably reflect the geometric structure of the face.
[0076] The extracted facial key points and geometric features are converted into feature vectors, which are used as illumination-invariant features for identity authentication, ensuring that users can still be accurately identified under different lighting conditions.
[0077] In this embodiment, facial key points in the first face image are extracted through an illumination-invariant feature extraction model, and then facial geometric features are extracted. Facial key points and facial geometric features have strong anti-interference capabilities against illumination changes; they are converted into feature vectors and used as illumination-invariant features for identity authentication. Under different lighting conditions, users can be effectively identified, reducing recognition errors caused by illumination changes and improving the robustness and reliability of identity authentication.
[0078] Furthermore, the above-mentioned step of extracting facial features from a face image may include: normalizing the face image to obtain a two-dimensional face image; performing three-dimensional reconstruction on the two-dimensional face image to obtain a three-dimensional model of the face; calculating the three-dimensional model based on a posture estimation algorithm to obtain posture parameters of the face; performing posture compensation on the two-dimensional face image according to the posture parameters to obtain a two-dimensional corrected image, and performing posture compensation on the three-dimensional model according to the posture parameters to obtain a three-dimensional corrected model; and performing feature extraction on the two-dimensional corrected image and the three-dimensional corrected model to obtain facial features.
[0079] Specifically, the facial image undergoes standardization processing, including denoising, contrast adjustment, and resizing, to produce a two-dimensional facial image. This 2D facial image is then reconstructed into three dimensions to generate a three-dimensional model that reflects the user's facial geometry. This 3D model preserves the depth information of the face and provides additional geometric details, facilitating a more accurate analysis of the user's facial features.
[0080] Apply a pose estimation algorithm (such as a rotation matrix or deep learning model) to the 3D model to calculate the face's pose parameters, including pitch, yaw, and roll. These pose parameters reflect the orientation and angle of the user's head when the image was captured.
[0081] Based on these pose parameters, pose compensation is performed on the original 2D facial image to produce a 2D corrected image. Similarly, pose compensation is performed on the 3D model to produce a 3D corrected model. Pose compensation corrects facial images from different angles to similar angles to maintain consistency during feature extraction.
[0082] For two-dimensional facial images, rotation, scaling, perspective correction and other transformations are performed on the two-dimensional facial images according to the posture parameters (pitch angle, yaw angle, roll angle), so that the face in the image is "straightened" to a standard perspective close to the front, reducing the distortion of facial features caused by the shooting angle, and facilitating subsequent comparison with facial images under the standard frontal perspective.
[0083] For a 3D model, the posture of the model is adjusted in 3D space to align it with the standard frontal view. In fact, the 3D model is rotated so that the frontal orientation of the model is consistent with the view.
[0084] After posture compensation is completed, feature extraction is performed on the two-dimensional corrected image and the three-dimensional corrected model to obtain facial features, which can effectively reflect the uniqueness of the user's face.
[0085] In this embodiment, the standardization process ensures the consistency of the input images, allowing subsequent processing under uniform standards and reducing recognition errors caused by differences in input image quality. Three-dimensional reconstruction provides depth information and geometric details for face recognition, enabling a more comprehensive analysis of users' facial features. Pose estimation and compensation address the impact of user head pose changes on recognition accuracy. By compensating for the pose, two-dimensional face images and three-dimensional models are corrected to a standard pose, allowing consistent feature extraction results from images taken at different angles. Feature extraction from the corrected images and models improves the accuracy of feature extraction.
[0086] Further, the step of extracting facial features from the two-dimensional corrected image and the three-dimensional corrected model can include: extracting texture features and depth features from the two-dimensional corrected image; extracting three-dimensional geometric features and stereo features from the three-dimensional corrected model; mapping the texture features to the three-dimensional corrected model to obtain a composite model; extracting composite geometric features and composite texture features of the composite model; and determining the texture features, the depth features, the three-dimensional geometric features, the stereo features, the composite geometric features, and the composite texture features as the facial features.
[0087] Specifically, texture features and depth features are extracted from the two-dimensional corrected image. Facial texture features are extracted from the compensated two-dimensional image using a convolutional neural network (CNN). These texture features include skin texture, local feature points (such as the shape and details of the eyes, nose, and mouth), etc. Depth features are higher-level feature vectors extracted by a deep learning model, which remain stable despite changes in facial pose and expression.
[0088] Three-dimensional geometric features and stereo features are extracted from the three-dimensional corrected model. Three-dimensional geometric features provide a detailed description of the facial structure, including curvature, edges, contours, and other geometric information. Stereo features include the relative positions and relationships of various parts of the face in three-dimensional space. These features reflect the inherent morphology of the user's face and are not affected by changes in expression or pose.
[0089] To better combine two-dimensional and three-dimensional information, texture features extracted from the two-dimensional corrected image are mapped onto the three-dimensional model to form a composite model. The composite model effectively fuses the surface details of the two-dimensional image with the spatial structure of the three-dimensional model, providing a more three-dimensional and realistic facial description.
[0090] Composite geometric features and composite texture features are extracted from the composite model. Composite geometric features combine the morphology of the three-dimensional corrected model with the mapped texture features. Composite texture features are enhanced surface detail information based on the three-dimensional structure. Composite features are more expressive than two-dimensional or three-dimensional features.
[0091] All the features extracted above, including the texture feature, the depth feature, the three-dimensional geometry feature, the stereo feature, the composite geometry feature and the composite texture feature, are determined as the facial features.
[0092] In this embodiment, the texture feature and the depth feature are extracted from the two-dimensional modified image, the surface details and the depth information are captured; the three-dimensional geometry feature and the stereo feature are extracted from the three-dimensional modified model, the understanding of the facial structure is enhanced, especially the spatial relationship and the morphology of each part of the face; the texture feature of the two-dimensional modified image is mapped onto the three-dimensional modified model to form a composite model, the advantages of two-dimensional and three-dimensional are fused, so that the composite geometry feature and the composite texture feature have stronger description ability and distinguishing ability; the comprehensive application of these features constructs an accurate and stable facial feature vector, greatly improves the reliability and robustness of the identity verification.
[0093] It should be emphasized that, in order to further ensure the privacy and security of the user voice and the face image, the user voice and the face image can also be stored in a node of a blockchain.
[0094] The blockchain referred to in the present application is a new application mode of distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm and other computer technologies. The blockchain is essentially a decentralized database, which is a series of data blocks associated using cryptographic methods, each data block contains the information of a batch of network transactions, used to verify the validity (anti-fake) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer and an application service layer, etc.
[0095] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0096] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0097] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0098] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0099] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of an identity verification device, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0100] like Figure 3 As shown, the identity verification device 300 of this embodiment includes: a user acquisition module 301, a voice verification module 302, a lighting verification module 303, a face verification module 304 and a result generation module 305, wherein:
[0101] The user acquisition module 301 is used to acquire the user's voice and face image through the smart glasses.
[0102] The voice verification module 302 is configured to perform voice enhancement processing on the user's voice to obtain enhanced voice, and perform identity verification on the user based on the enhanced voice to obtain a first verification result.
[0103] The illumination verification module 303 is used to input the face image into the illumination invariant feature extraction model to obtain the illumination invariant features of the face image, and perform identity verification on the user based on the illumination invariant features to obtain a second verification result.
[0104] The face verification module 304 is used to extract facial features of the face image and authenticate the user based on the facial features to obtain a third verification result.
[0105] The result generation module 305 is configured to generate an identity authentication result of the user according to the first verification result, the second verification result and the third verification result.
[0106] In this embodiment, the user's voice and facial image are obtained through smart glasses; the user's voice is enhanced to obtain enhanced voice, which improves the robustness of voiceprint recognition, accurately identifies the user's identity even in a noisy environment, and generates a first verification result; the facial image is input into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, which overcome the influence of ambient lighting changes on facial recognition, ensure recognition stability under different lighting conditions, and ensure the accuracy of the second verification result; the facial features of the facial image are extracted, which improves the accuracy of identity recognition under complex facial postures and expression changes, and obtains a third verification result; the first verification result, the second verification result, and the third verification result are combined to generate the user's final identity authentication result, which reduces the risk of misidentification and deceptive attacks and improves the accuracy of identity authentication based on smart glasses.
[0107] In some optional implementations of this embodiment, the voice verification module 302 may include:
[0108] The preprocessing submodule is used to preprocess the user's voice to obtain the first voice.
[0109] The frequency domain extraction submodule is used to extract the frequency domain features of the first speech through a convolutional neural network.
[0110] The time domain modeling submodule is used to perform time domain modeling on the frequency domain features through a gated recurrent unit to obtain enhanced speech.
[0111] The voiceprint extraction submodule is used to extract the voiceprint features of the enhanced speech through the voiceprint recognition model.
[0112] The voiceprint comparison submodule is used to compare the voiceprint feature with each stored voiceprint in the voiceprint feature library to authenticate the user and obtain a first verification result.
[0113] In some optional implementations of this embodiment, the lighting verification module 303 may include:
[0114] The illumination detection submodule is used to perform illumination detection on the face image to obtain the illumination conditions of the face image.
[0115] The image transformation submodule is configured to input the face image into the illumination-invariant feature extraction model, and perform feature transformation on the face image according to the illumination condition by using the illumination-invariant feature extraction model, to obtain a first face image.
[0116] The illumination extraction submodule is configured to perform feature extraction on the first face image by using the illumination-invariant feature extraction model, to obtain illumination-invariant features.
[0117] In some optional implementations of the embodiment, the illumination extraction submodule can include:
[0118] The extraction unit is configured to extract the face key points and the geometric features of the first face image by using the illumination-invariant feature extraction model.
[0119] The conversion unit is configured to convert the face key points and the geometric features into a feature vector, to obtain the illumination-invariant features.
[0120] In some optional implementations of the embodiment, the face verification module 304 can include:
[0121] The standardization processing submodule is configured to perform standardization processing on the face image, to obtain a two-dimensional face image.
[0122] The three-dimensional reconstruction submodule is configured to perform three-dimensional reconstruction on the two-dimensional face image, to obtain a three-dimensional model of the face.
[0123] The pose calculation submodule is configured to calculate the three-dimensional model based on a pose estimation algorithm, to obtain pose parameters of the face.
[0124] The pose compensation submodule is configured to perform pose compensation on the two-dimensional face image according to the pose parameters, to obtain a two-dimensional corrected image, and perform pose compensation on the three-dimensional model according to the pose parameters, to obtain a three-dimensional corrected model.
[0125] The feature extraction submodule is configured to perform feature extraction on the two-dimensional corrected image and the three-dimensional corrected model, to obtain face features.
[0126] In some optional implementations of the embodiment, the feature extraction submodule can include:
[0127] The two-dimensional extraction unit is configured to extract texture features and depth features from the two-dimensional corrected image.
[0128] The three-dimensional extraction unit is configured to extract three-dimensional geometric features and stereo features from the three-dimensional corrected model.
[0129] The texture mapping unit is configured to map the texture features to the three-dimensional corrected model, to obtain a composite model.
[0130] The composite extraction unit is configured to extract composite geometric features and composite texture features of the composite model.
[0131] determine the texture feature, the depth feature, the three-dimensional geometry feature, the stereo feature, the composite geometry feature and the composite texture feature as the face feature.
[0132] To solve the above technical problems, the embodiment of the present application further provides a computer device. For details, please refer to Figure 4 Figure 4 The basic structure block diagram of the computer device of the embodiment is shown in the figure.
[0133] The computer device 4 comprises a memory 41, a processor 42 and a network interface 43 which are connected to each other through a system bus. It should be noted that only the computer device 4 with the memory 41, the processor 42 and the network interface 43 is shown in the figure, but it should be understood that all the shown components are not required to be implemented, and more or less components can be alternatively implemented. Among them, the computer device herein is a device which can automatically perform numerical calculation and / or information processing according to pre-set or stored instructions, and the hardware thereof includes but is not limited to microprocessor, application specific integrated circuit (ASIC), field-programmable gate array (FPGA), digital signal processor (DSP), embedded device, etc.
[0134] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote controller, a touchpad or a sound control device.
[0135] The memory 41 includes at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or a memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both an internal storage unit and an external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store an operating system and various application software installed on the computer device 4, such as computer readable instructions of the identity verification method, etc. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.
[0136] The processor 42 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer readable instructions or process data stored in the memory 41, such as computer readable instructions of the identity verification method.
[0137] The network interface 43 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0138] In this embodiment, the user's voice and facial image are obtained through smart glasses; the user's voice is enhanced to obtain enhanced voice, which improves the robustness of voiceprint recognition, accurately identifies the user's identity even in a noisy environment, and generates a first verification result; the facial image is input into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, which overcome the influence of ambient lighting changes on facial recognition, ensure recognition stability under different lighting conditions, and ensure the accuracy of the second verification result; the facial features of the facial image are extracted, which improves the accuracy of identity recognition under complex facial postures and expression changes, and obtains a third verification result; the first verification result, the second verification result, and the third verification result are combined to generate the user's final identity authentication result, which reduces the risk of misidentification and deceptive attacks and improves the accuracy of identity authentication based on smart glasses.
[0139] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned identity authentication method.
[0140] In this embodiment, the user's voice and facial image are obtained through smart glasses; the user's voice is enhanced to obtain enhanced voice, which improves the robustness of voiceprint recognition, accurately identifies the user's identity even in a noisy environment, and generates a first verification result; the facial image is input into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, which overcome the influence of ambient lighting changes on facial recognition, ensure recognition stability under different lighting conditions, and ensure the accuracy of the second verification result; the facial features of the facial image are extracted, which improves the accuracy of identity recognition under complex facial postures and expression changes, and obtains a third verification result; the first verification result, the second verification result, and the third verification result are combined to generate the user's final identity authentication result, which reduces the risk of misidentification and deceptive attacks and improves the accuracy of identity authentication based on smart glasses.
[0141] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0142] Obviously, the above-described embodiments are only some embodiments of the present application, but not all the embodiments. The preferred embodiments of the present application are shown in the drawings, but do not limit the patent scope of the present application. The present application can be implemented in many different forms, and contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the patent protection scope of the present application.
Claims
1. An identity authentication method, characterized in that: The steps include: Acquire the user's voice and facial image through the smart glasses; performing voice enhancement processing on the user's voice to obtain enhanced voice, and performing identity authentication on the user based on the enhanced voice to obtain a first authentication result; Inputting the facial image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, and performing identity verification on the user based on the illumination-invariant features to obtain a second verification result; extracting facial features from the face image, and performing identity verification on the user based on the facial features to obtain a third verification result; Generating an identity authentication result of the user according to the first verification result, the second verification result, and the third verification result; The step of extracting facial features of the face image comprises: performing normalization processing on the facial image to obtain a two-dimensional facial image; Performing three-dimensional reconstruction on the two-dimensional face image to obtain a three-dimensional face model; Calculating the three-dimensional model based on a posture estimation algorithm to obtain posture parameters of the face; Performing posture compensation on the two-dimensional face image according to the posture parameters to obtain a two-dimensional corrected image, and performing posture compensation on the three-dimensional model according to the posture parameters to obtain a three-dimensional corrected model; Feature extraction is performed on the two-dimensional corrected image and the three-dimensional corrected model to obtain facial features.
2. The identity authentication method according to claim 1, wherein: The step of performing voice enhancement processing on the user voice to obtain enhanced voice, and performing identity authentication on the user based on the enhanced voice to obtain a first verification result includes: Preprocessing the user voice to obtain a first voice; Extracting frequency domain features of the first speech through a convolutional neural network; Performing time-domain modeling on the frequency-domain features through a gated recurrent unit to obtain enhanced speech; extracting voiceprint features of the enhanced speech through a voiceprint recognition model; The voiceprint feature is compared with each stored voiceprint in a voiceprint feature library to authenticate the user and obtain a first verification result.
3. The identity authentication method according to claim 1, wherein: The step of inputting the facial image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image comprises: Performing illumination detection on the facial image to obtain illumination conditions of the facial image; Inputting the facial image into an illumination-invariant feature extraction model, so as to perform feature transformation on the facial image according to the illumination conditions by the illumination-invariant feature extraction model to obtain a first facial image; Feature extraction is performed on the first facial image using the illumination-invariant feature extraction model to obtain illumination-invariant features.
4. The identity authentication method according to claim 3, wherein: The step of extracting features from the first face image using the illumination-invariant feature extraction model to obtain illumination-invariant features comprises: Extracting facial key points and geometric features of the first face image using the illumination invariant feature extraction model; The facial key points and the geometric features are converted into feature vectors to obtain illumination invariant features.
5. The identity authentication method according to claim 1, wherein: The step of extracting features from the two-dimensional corrected image and the three-dimensional corrected model to obtain facial features includes: Extracting texture features and depth features from the two-dimensional corrected image; extracting three-dimensional geometric features and stereoscopic features from the three-dimensional corrected model; Mapping the texture features to the three-dimensional corrected model to obtain a composite model; extracting composite geometric features and composite texture features of the composite model; The texture feature, the depth feature, the three-dimensional geometric feature, the stereo feature, the composite geometric feature and the composite texture feature are determined as facial features.
6. An identity verification device, characterized in that: include: A user acquisition module, used to acquire the user's voice and face image through the smart glasses; a voice verification module, configured to perform voice enhancement processing on the user's voice to obtain enhanced voice, and authenticate the user based on the enhanced voice to obtain a first verification result; an illumination verification module, configured to input the facial image into an illumination-invariant feature extraction model to obtain illumination-invariant features of the facial image, and authenticate the user based on the illumination-invariant features to obtain a second verification result; A facial verification module, configured to extract facial features from the face image and authenticate the user based on the facial features to obtain a third verification result; A result generating module, configured to generate an identity authentication result of the user according to the first verification result, the second verification result and the third verification result; The facial verification module includes: A standardization processing submodule, configured to perform standardization processing on the facial image to obtain a two-dimensional facial image; A three-dimensional reconstruction submodule, configured to perform three-dimensional reconstruction on the two-dimensional face image to obtain a three-dimensional model of the face; A posture calculation submodule, configured to calculate the three-dimensional model based on a posture estimation algorithm to obtain posture parameters of the face; a posture compensation submodule, configured to perform posture compensation on the two-dimensional face image according to the posture parameters to obtain a two-dimensional corrected image, and perform posture compensation on the three-dimensional model according to the posture parameters to obtain a three-dimensional corrected model; The feature extraction submodule is used to extract features from the two-dimensional corrected image and the three-dimensional corrected model to obtain facial features.
7. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the identity authentication method according to any one of claims 1 to 5 when executing the computer-readable instructions.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the identity authentication method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Identity authentication method and device
CN103475490A