Face recognition method, device and equipment and readable storage medium
By displaying specified content on the screen and allowing users to express themselves verbally, and by matching local facial regions with audio content, the problem of liveness detection that ordinary camera terminals cannot solve is solved, thus improving the accuracy and security of facial recognition.
Patent Information
- Application Number
- CN202110369362.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-06
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-04-06
AI Technical Summary
In existing technologies, ordinary camera terminals cannot perform liveness detection, resulting in low accuracy and poor security of facial recognition results, and they cannot prevent photo or video attacks.
By displaying specified content at random locations on the screen, users can express themselves verbally, and the system can verify their liveness by matching facial features with the audio content.
It improves the accuracy and security of facial recognition, ensures that participants are live users, and prevents image or video attacks.
Smart Images

Figure CN115171175B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face recognition, and in particular to a face recognition method, apparatus, device, and readable storage medium. Background Technology
[0002] Facial recognition is a highly efficient method of identity verification. Taking a user's handheld device as an example, the device's camera typically captures the user's facial image and matches it against facial images in a facial image database to complete the verification process. In some cases, to prevent users from completing facial recognition verification using static photos, liveness detection is also required during the facial recognition process.
[0003] In related technologies, when performing liveness detection, facial images are usually captured using cameras with depth information, such as infrared cameras and depth cameras. By attaching depth information while capturing facial images, it is possible to effectively intercept behaviors such as photo attacks and video re-enactments.
[0004] However, when implementing facial recognition in the above way, the terminal needs to be equipped with specific camera hardware. In scenarios where many terminals are equipped with ordinary cameras, liveness detection cannot be achieved, and attacks that attempt to identify the face through photos, videos, etc., cannot be prevented during the facial recognition process. As a result, the accuracy of the facial recognition results is low and the security is poor. Summary of the Invention
[0005] This application provides a face recognition method, apparatus, device, and readable storage medium, which can improve the accuracy and security of face recognition results. The technical solution is as follows:
[0006] On the one hand, a face recognition method is provided, the method comprising:
[0007] In response to the start of the face recognition process, a face recognition video stream and audio content are acquired, wherein the face recognition video stream and the audio content are content acquired based on specified display content, and the display position of the specified display content is determined from at least two candidate display positions;
[0008] A partial facial region is extracted from the facial recognition video stream, wherein the partial facial region is the area corresponding to the facial features that are expressed when the specified display content is spoken.
[0009] The facial region is matched with the specified display content to obtain a first matching result;
[0010] The audio content is matched with the specified display content to obtain a second matching result;
[0011] The face recognition result is determined based on the first matching result and the second matching result.
[0012] On the other hand, a face recognition method is provided, the method comprising:
[0013] Display a face recognition interface, which includes a captured face image;
[0014] The specified display content is displayed in the face recognition interface, and the display position of the specified display content is determined from at least two candidate display positions;
[0015] Voice prompts are displayed on the face recognition interface, and these voice prompts are used to instruct the user to express the specified content in voice.
[0016] The face recognition result is displayed on the face recognition interface based on the captured face image and the voice prompt information.
[0017] On the other hand, a facial recognition device is provided, the device comprising:
[0018] The acquisition module is used to acquire a face recognition video stream and audio content in response to the start of the face recognition process, wherein the face recognition video stream and the audio content are content acquired based on specified display content, and the display position of the specified display content is determined from at least two candidate display positions;
[0019] The acquisition module is further configured to extract a partial facial region from the facial recognition video stream, wherein the partial facial region is the region corresponding to the facial features that are expressed when the specified display content is spoken.
[0020] The matching module is used to match the local area of the face with the specified display content to obtain a first matching result;
[0021] The matching module is further configured to match the audio content with the specified display content to obtain a second matching result;
[0022] The determination module is used to determine the face recognition result based on the first matching result and the second matching result.
[0023] On the other hand, a facial recognition device is provided, the device comprising:
[0024] The display module is used to display the face recognition interface, which includes a captured face image.
[0025] The display module is also used to display specified display content in the face recognition interface, wherein the display position of the specified display content is determined from at least two candidate display positions;
[0026] The display module is also used to display voice prompt information in the face recognition interface, the voice prompt information being used to instruct the corresponding voice expression for the specified display content;
[0027] The display module is also used to display the face recognition result on the face recognition interface based on the captured face image and the voice prompt information.
[0028] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the face recognition method as described in any of the embodiments of this application above.
[0029] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the face recognition method as described in any of the embodiments of this application above.
[0030] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the face recognition methods described in the above embodiments.
[0031] The beneficial effects of the technical solutions provided in this application include at least the following:
[0032] By matching a local area of the face with the specified displayed content, changes in the face during speech are confirmed. By matching the audio content with the specified displayed content, the user's spoken content is confirmed. This allows for liveness detection at both the level of facial changes and the level of spoken content, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video. This improves the accuracy of face recognition and enhances the security of the protection function provided by face recognition. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 This is a schematic diagram of the overall process of face recognition provided in an exemplary embodiment of this application;
[0035] Figure 2 This is a schematic diagram illustrating the implementation environment of a face recognition method provided in an exemplary embodiment of this application;
[0036] Figure 3 This is a flowchart of a face recognition method provided in an exemplary embodiment of this application;
[0037] Figure 4 This is a flowchart of a face recognition method provided in another exemplary embodiment of this application;
[0038] Figure 5 Based on Figure 4 The illustrated embodiment provides a schematic diagram of the image sequence acquisition process during face recognition;
[0039] Figure 6 This is a schematic diagram of the overall process of face recognition provided in an exemplary embodiment of this application;
[0040] Figure 7 This is a flowchart of a face recognition method provided in an exemplary embodiment of this application;
[0041] Figure 8 Based on Figure 7 A schematic diagram of a face recognition interface provided in the illustrated embodiment;
[0042] Figure 9 This is a structural block diagram of a face recognition device provided in an exemplary embodiment of this application;
[0043] Figure 10 This is a structural block diagram of a face recognition device provided in another exemplary embodiment of this application;
[0044] Figure 11 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0046] First, a brief introduction to the terms used in the embodiments of this application:
[0047] Artificial intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0048] Computer vision (CV) technology refers to machine vision that uses cameras and computers to identify, track, and measure targets, replacing the human eye. It further processes images to create images more suitable for human observation or for transmission to instruments. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, and video content / behavior recognition. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0049] Facial recognition refers to the function of identifying the identity information of a face in a facial image. Optionally, during the facial recognition process, features are extracted from the facial region to be identified, and the extracted features are compared with features in a preset facial feature database to determine the identity information of the face in that region. Typically, facial recognition is applied in scenarios such as terminal unlocking, attendance tracking, and resource payment. For illustrative purposes, this embodiment uses a resource payment scenario as an example. During resource payment, the user scans their facial image using a payment device, matching the image with facial images in a preset facial database to complete the payment. To avoid the security risks associated with users using photos or videos for resource payment, facial liveness detection is also performed during facial recognition to ensure that the user is indeed collecting the facial image, and not a photo or video.
[0050] In related technologies, liveness detection typically involves one of the following methods: 1. Issuing commands such as blinking or tilting the head to the user and detecting whether the user performs the specified action; 2. Collecting depth information through an infrared camera or depth camera to avoid attacks using planar content such as photos or videos. However, in the first method, the liveness detection process can be easily bypassed by pre-recording video of the action; the second method requires the terminal to be equipped with specific camera hardware, which is not feasible in scenarios where many terminals are equipped with ordinary cameras.
[0051] In this embodiment of the application, specified content is displayed at a random position on the display screen, and the user expresses the specified content by voice. Liveness detection is performed based on facial expression features and voice expression content. Liveness detection is determined to be successful only when the user expresses the specified content correctly by voice and the facial features accurately correspond to the display position and content of the specified content.
[0052] To illustrate, let's take the example of displaying a single-digit number. Random single-digit numbers are displayed sequentially at random positions on the screen. The user expresses the information by voice based on the displayed numbers. During liveness detection, it is necessary to determine whether the user is a live person for facial recognition based on the accuracy of the user's voice expression, the matching of the user's lip movements with the displayed numbers, and the matching of the user's gaze with the position of the displayed numbers.
[0053] Figure 1 This is a schematic diagram of the overall process of face recognition provided in an exemplary embodiment of this application, as shown below. Figure 1 As shown, the terminal interface 100 displays the numbers 3 (top left), 6 (bottom right), 8 (bottom left), and 1 (top right) sequentially. The user verbally expresses the displayed numbers according to the terminal interface, that is, says 3, 6, 8, and 1 in sequence. Simultaneously, facial recognition video stream 110 and audio stream 120 are captured. During facial recognition, the part in the audio stream where the "3" is said is confirmed, and the corresponding video image frame when the "3" is said is located in the facial recognition video stream 110. The mouth area and eye area are cropped from the video image frame. The correlation between the mouth shape and the number "3" is determined, and the correlation between the gaze direction and the top left corner is determined. The same detection is performed for "6", "8", and "1", ultimately obtaining the liveness detection result.
[0054] The face recognition method provided in this application can be implemented by a terminal, by a server, or by a combination of both.
[0055] When the face recognition method is implemented by the terminal, the terminal randomly selects specified display content from the content library and displays it at a random location, and collects face recognition video stream and audio stream. Based on the display method of the specified display content, it matches it with the face recognition video stream and audio stream to obtain the liveness detection result, and matches the face image in the face recognition video stream with the preset image library to obtain the face recognition result.
[0056] In this embodiment, the method of implementing face recognition is illustrated by the cooperation of a terminal and a server. Figure 2 This is a schematic diagram illustrating the implementation environment of a face recognition method provided in an exemplary embodiment of this application, such as... Figure 2As shown, the implementation environment includes a terminal 210 and a server 220, wherein the terminal 210 and the server 220 are connected through a communication network 230.
[0057] Terminal 210 has an application with authentication functionality installed. When a user uses this application and authentication is required, terminal 210 sends an authentication request to server 220. Server 220 then returns a display scheme for the specified content to terminal 210. Terminal 210 displays the specified content according to the display scheme on the face recognition interface, captures a face recognition video stream via camera, and captures an audio stream via microphone.
[0058] Terminal 210 sends the collected face recognition video stream and audio stream to server 220 via communication network 230. Server 220 matches the received audio stream and face recognition video stream with the display scheme of the specified display content sent to terminal 210 to confirm the liveness detection result. In some embodiments, server 220 first determines the face recognition result and determines the liveness detection result when the face recognition result is correct, or server 220 obtains the face recognition result based on the liveness detection result.
[0059] After receiving the face recognition result and confirming that the liveness detection was successful, the server 220 sends the face recognition result back to the terminal 210, which then performs the next operation based on the face recognition result.
[0060] It is worth noting that the aforementioned communication network 230 can be implemented as a wired network or a wireless network, and the communication network 230 can be implemented as any one of a local area network, a metropolitan area network, or a wide area network. This application embodiment does not limit this.
[0061] The aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0062] The terminal can be a smartphone, tablet, resource payment device, attendance device, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these.
[0063] Based on the above description of terms and implementation environment, the face recognition method provided in this application embodiment will be described, and the method will be applied to, for example... Figure 2 The following explanation uses the server shown as an example. Figure 3 As shown, the method includes:
[0064] Step 301: In response to the start of the face recognition process, acquire the face recognition video stream and audio content.
[0065] The facial recognition video stream and audio content are collected based on specified display content. The display position of the specified display content is determined from at least two candidate display positions.
[0066] In some embodiments, when a terminal requires facial recognition, it sends a facial recognition request to the server, thereby initiating the facial recognition process. When the server determines that the facial recognition process has begun, it sends a display scheme specifying the content to be displayed to the terminal. This display scheme includes content data for the specified content and the specified display position of the content on the terminal's screen. The terminal then displays the specified content at the designated position on the screen based on the display scheme.
[0067] In some embodiments, the specified display content is content randomly determined by the server from a preset content library; the display position is a position randomly determined by the server from at least two pre-defined candidate display positions. For example, the preset content library includes ten numbers from 0 to 9. The server randomly determines four numbers from the preset content library, which may or may not be repeated. The server randomly determines the display positions for the four numbers and sends the four numbers and their display positions to the terminal as the specified display content to be displayed in sequence.
[0068] The terminal activates its camera to capture a video stream for facial recognition and activates its microphone to capture audio content. Simultaneously, the terminal displays the specified content designated by the server on its facial recognition interface. In some embodiments, the terminal's facial recognition interface also displays voice prompts instructing the user to express the specified content via voice. That is, when the terminal displays the specified content, the user expresses the content displayed on the facial recognition interface via voice, thereby the camera captures a facial video stream of the voice expression process, and the microphone captures the audio content of the voice expression process.
[0069] It is worth noting that the above process is illustrated using the server acquiring the facial recognition video stream and audio content as an example. In some embodiments, when the facial recognition method is implemented by the terminal, the process includes the following: When the user selects a function requiring facial recognition verification in the application, the facial recognition function is triggered, and the terminal determines that the facial recognition process has begun. First, the terminal randomly determines the specified display content from a preset content library and determines the display position of the specified display content from at least two preset candidate display positions, thereby displaying the specified display content at the display position on the interface. In some embodiments, the terminal's facial recognition interface also displays voice prompts, which are used to instruct the user to express the specified display content via voice. That is, when the terminal displays the specified display content, the user expresses the specified display content displayed on the facial recognition interface via voice. Simultaneously, the terminal captures the facial video stream of the voice expression process through a camera and the audio content of the voice expression process through a microphone.
[0070] It is worth noting that the above-mentioned preset content library uses the ten digits from 0 to 9 as an example. In some embodiments, the preset content library may also include text content, static image content, dynamic image content, etc., and this application embodiment does not limit this. For illustration, the preset content library includes animal pictures, and when displaying animal pictures, the user is instructed to verbally describe the animals in the pictures.
[0071] The aforementioned at least two display positions include any of the following: 1. Specifying at least two display positions on the display screen, such as specifying the top left corner, bottom left corner, top right corner, and bottom right corner as candidate display positions, thereby clearly distinguishing the user's line of sight through these four positions; 2. Dividing the display screen into a grid, with n grids serving as at least two candidate display positions on the display screen, where n is a positive integer. This application does not limit the method for determining the candidate display positions.
[0072] Step 302: Extract a local area of the face from the face recognition video stream.
[0073] A facial region is the area corresponding to the facial features that appear when a specified content is expressed verbally.
[0074] In some embodiments, when a user expresses the specified display content verbally, the mouth needs to make verbal expressions, so there is mouth movement; in addition, since the display position of the specified display content on the display screen is determined from at least two candidate display positions, that is, there are differences in display position between different specified display contents, there is also eye movement.
[0075] In some embodiments, the local area of the face includes the mouth area of the face. That is, the mouth area of the face is extracted from the face recognition video stream for lip-reading recognition to obtain the lip reading result, which is the vocal content obtained from the mouth area recognition.
[0076] In other embodiments, the face local area includes the face eye area, that is, the face eye area is extracted from the face recognition video stream for gaze recognition to obtain gaze recognition results, wherein the gaze recognition results indicate the gaze direction of the eye area.
[0077] In some embodiments, before extracting the facial region, the face recognition video is first segmented, and video segments during the user's voice expression are extracted, thereby extracting the facial region based on these video segments. Illustratively, audio features of the audio content are extracted, and the time period of the voice expression of the specified displayed content is determined based on these audio features, thereby locating the image sequence corresponding to the time period from the face recognition video stream, i.e., the aforementioned video segment. The facial region is then extracted from the image frames in the image sequence. Specifically, facial region extraction is performed sequentially on all image frames in the image sequence; or, facial region extraction is performed on a specific image frame in the image sequence.
[0078] Step 303: Match the local area of the face with the specified display content to obtain the first matching result.
[0079] For different situations involving local facial regions, the first matching result includes at least one of the following:
[0080] First, when the local area of a face includes the mouth area, the mouth area is the area obtained by cropping the mouth that expresses the specified display content, and the first matching result includes the mouth matching result.
[0081] Specifically, lip-reading is performed on the mouth area of a face to obtain lip-reading results, which represent the identified vocalizations. These lip-reading results are then matched with specified displayed content to obtain lip-matching results, which represent the correlation between the lip-reading results and the specified displayed content. The lip-reading results can be presented in either pinyin or actual content format. After obtaining the lip-reading results, the similarity between the lip-reading results and the specified displayed content is determined.
[0082] For example, if the specified content is the number "3", and lip reading is performed on the mouth area of the face, and the lip reading result is "san", then the lip reading result matches the specified content, and the similarity between the lip reading result and the specified content is 95%.
[0083] Second, when the local area of the face includes the eye area, the eye area is the area obtained by cropping the eye of the specified display content, and the first matching result includes the eye matching result.
[0084] Specifically, gaze recognition is performed on the eye region of the face to obtain a gaze recognition result, which indicates the direction of the recognized eye. This gaze recognition result is then matched with the display position of specified content to obtain an eye matching result, which indicates the correlation between the gaze recognition result and the specified content. The gaze recognition result can be expressed as a direction or as the position of the gaze point on the terminal display screen. After obtaining the gaze recognition result, the similarity between the gaze recognition result and the display position of the specified content is determined.
[0085] Indicatively, the content "3" is displayed at the upper left (5, 5) position on the screen. The gaze recognition of the face's eye area is performed, and the gaze recognition result is the landing point position (7, 6) on the screen. Thus, the distance between the two coordinates is determined to be 5 based on the coordinates (5, 5) and (8, 9). Based on the total length of the screen's diagonal of 50, the similarity between the two coordinates is 90%.
[0086] Step 304: Match the audio content with the specified display content to obtain the second matching result.
[0087] In some embodiments, features are extracted from the audio content to obtain audio features, and then speech recognition is performed based on the audio features.
[0088] Optionally, speech recognition of the audio content can be performed using a pre-trained neural network model. That is, audio features or audio content are input into the neural network model, and after the model performs speech recognition, a speech recognition result is output. The speech recognition result is then matched with specified display content to obtain a second matching result.
[0089] The neural network model can be obtained through supervised training or unsupervised training. Taking supervised training as an example, sample audio labeled with reference results is input into the neural network model, and the output is the prediction result. The model parameters in the neural network model are adjusted based on the difference between the prediction result and the reference result.
[0090] In some embodiments, the second matching result represents the similarity between the speech recognition result and the specified displayed content. For example, if the specified displayed content is "3" and the speech recognition result is "3", then the similarity between the speech recognition result and the specified displayed content is 100%.
[0091] Step 305: Determine the face recognition result based on the first matching result and the second matching result.
[0092] In some embodiments, the first matching result and the second matching result are weighted and summed to obtain a confidence score. Based on the relationship between the confidence score and a preset threshold, the liveness detection result and the face recognition result are determined. Specifically, the liveness detection result is first determined to be a live face, thereby determining the face recognition result; or, the face recognition result is first determined to match a face in the face recognition database, thereby determining the liveness detection result.
[0093] In summary, the face recognition method provided in this application confirms changes in the facial features during speech by matching local facial regions with specified display content, and confirms the user's vocal content by matching audio content with specified display content. This enables liveness detection at both the facial change and vocal content levels, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video, thus improving the accuracy of face recognition and enhancing the security of the protection function provided by face recognition.
[0094] In some embodiments, the facial region is obtained by cropping a specified display time period of the content, or both the cropping of the facial region and the recognition of the audio content are based on audio features. Figure 4 This is a flowchart of a face recognition method provided in another exemplary embodiment of this application. The method is illustrated using an example of its application in a server. Figure 4 As shown, the method includes:
[0095] Step 401: In response to the start of the face recognition process, acquire the face recognition video stream and audio content.
[0096] The facial recognition video stream and audio content are collected based on specified display content. The display position of the specified display content is determined from at least two candidate display positions.
[0097] Optionally, the display position of the specified content is randomly obtained from at least two candidate display positions, and the specified display content is randomly determined from a preset content library.
[0098] In some embodiments, when a terminal requires facial recognition, it sends a facial recognition request to the server, thereby initiating the facial recognition process. When the server determines that the facial recognition process has begun, it sends a display scheme specifying the content to be displayed to the terminal. This display scheme includes content data for the specified content and the specified display position of the content on the terminal's screen. The terminal then displays the specified content at the designated position on the screen based on the display scheme.
[0099] Step 402: Locate the image sequence corresponding to the display of the specified content from the face recognition video stream.
[0100] In some embodiments, audio features of the audio content are extracted, and the time period for the speech expression of the specified display content is determined based on the audio features, thereby locating the image sequence corresponding to the time period from the face recognition video stream.
[0101] In this process, a pre-trained feature extraction model is used to extract features from the audio content, thereby obtaining audio features. Based on these audio features, the corresponding time period in the audio content during which the user makes a speech expression is predicted.
[0102] To illustrate, after extracting the audio features corresponding to the audio content, the time period corresponding to the user's voice expression is predicted to be 00:05 to 00:08, thereby locating the image sequence corresponding to the time period of 00:05 to 00:08 from the face recognition video stream.
[0103] In some embodiments, the face recognition process includes at least two sequentially displayed specified content items. Therefore, when locating the image sequence in the face recognition video stream, it is necessary to locate the image sequences corresponding to at least two specified display items. That is, based on audio features, the i-th time period for the speech expression of the i-th specified display item is determined, where i is a positive integer, thereby locating the i-th group of image sequences corresponding to the i-th time period from the face recognition video stream.
[0104] This is illustrative; please refer to it. Figure 5During the face recognition process, the face recognition interface 500 displays four random numbers 3, 6, 8, and 1 in sequence. After extracting the audio features corresponding to the audio content 510, it is predicted that the user's verbal expression of the number "3" will be between 00:05 and 00:08. Therefore, the image sequence 521 corresponding to "3" is located from the face recognition video stream during the time period of 00:05 to 00:08. Similarly, it is predicted that the user's verbal expression of the number "6" will be between 00:08 and 00:10. Therefore, the image sequence 521 corresponding to "3" is located from the face recognition video stream during the time period of 00:05 to 00:08. Image sequence 522, located from the time interval 00:08 to 00:10, corresponds to "6". The predicted time interval for the user's verbal expression of the number "8" is 00:10 to 00:13, so image sequence 523, located from the face recognition video stream, corresponds to "8"; the predicted time interval for the user's verbal expression of the number "1" is 00:13 to 00:15, so image sequence 524, located from the face recognition video stream, corresponds to "1".
[0105] Step 403: Extract regions from the image frames in the image sequence to obtain local facial regions.
[0106] In some embodiments, when the face recognition process includes at least two sequentially displayed specified content, region cropping is performed on each image sequence to obtain the local face region corresponding to each image sequence.
[0107] To illustrate, when the local facial region includes the mouth region and the eye region, taking three specified display contents as an example, for display content A, the mouth region 1 and the eye region 1 of the face are extracted from the image sequence a corresponding to display content A; for display content B, the mouth region 2 and the eye region 2 of the face are extracted from the image sequence b corresponding to display content B; and for display content C, the mouth region 3 and the eye region 3 of the face are extracted from the image sequence c corresponding to display content C.
[0108] Step 404: Match the local area of the face with the specified display content to obtain the first matching result.
[0109] In some embodiments, when the face recognition process includes at least two sequentially displayed specified content items, the facial local region extracted from the i-th image sequence is matched with the i-th specified display content to obtain the i-th matching sub-result. The matching sub-results corresponding to at least two specified display content items are combined to obtain the first matching result. For example, taking n specified display content items, the weighted average of the n matching sub-results is taken to obtain the first matching result.
[0110] Alternatively, the content recognition results of n specified display contents corresponding to local facial regions are obtained, and a recognition result sequence is formed by connecting the n content recognition results. This recognition result sequence is then compared with a reference sequence formed by connecting the n specified display contents to obtain the first matching result, where n is a positive integer. For example, taking the mouth region of a face, lip-syncing is performed to obtain the vocal content. The recognition result sequence formed by connecting the n vocal contents is matched with a reference sequence formed by connecting the n specified display contents, and the similarity between the recognition result sequence and the reference sequence is calculated as the first matching result. Similarly, taking the eye region of a face, gaze recognition is performed to obtain the gaze direction. The recognition result sequence formed by connecting the n gaze directions is matched with a reference sequence formed by connecting the display positions of the n specified display contents, and the similarity between the recognition result sequence and the reference sequence is calculated as the first matching result.
[0111] Step 405: Match the audio content with the specified display content to obtain the second matching result.
[0112] In some embodiments, after extracting the audio features of the audio content in step 402 above, speech recognition is performed on the audio content based on the audio features to obtain a speech recognition result; the speech recognition result is matched with the specified display content to obtain a second matching result, wherein the second matching result is used to represent the degree of correlation between the recognition result and the specified display content.
[0113] In some embodiments, the face recognition process includes at least two sequentially displayed specified content items, and the speech recognition result includes a speech recognition sequence, which includes at least two sequentially arranged recognition sub-results. The recognition sub-results in the speech recognition sequence are then matched with the at least two specified display content items to obtain a content matching result. The second matching result described above is obtained based on the content recognition result.
[0114] In some embodiments, the recognition sub-results in the speech recognition sequence can be sequentially matched with at least two specified display contents to obtain a sequential matching result, and a second matching result can be obtained based on the content matching result and the sequential matching result.
[0115] In the speech recognition sequence, among at least two recognition sub-results, the m-th recognition sub-result corresponds to the m-th specified display content, where m is a positive integer.
[0116] As an illustration, the speech recognition sequence includes the sequentially arranged recognition sub-results "3, 6, 8, 7", while the sequence specifying the displayed content is "3, 6, 8, 1". The second matching result is: similarity 75%.
[0117] It is worth noting that, in the above embodiments, the sequential display of specified content is used as an example for explanation. In some embodiments, multiple different specified content can be displayed simultaneously in different display positions in the face recognition interface. The order and content of the speech recognition sequence, gaze change sequence, lip shape change sequence and the display position of the specified content are matched according to the user's speech recognition sequence, gaze change sequence and lip shape change sequence.
[0118] Step 406: Determine the face recognition result based on the first matching result and the second matching result.
[0119] In some embodiments, the first matching result and the second matching result are weighted and summed to obtain the liveness detection probability, which represents the probability that the face recognition process is completed by a live person; the face recognition result is determined based on the liveness detection probability.
[0120] In some embodiments, when the first matching result includes a face mouth matching result identified by the face mouth region and a face eye matching result identified by the face eye region, the face mouth matching result, the face eye matching result and the second matching result are weighted and summed, and the liveness detection probability is determined based on the weighted summation result.
[0121] To illustrate, the matching result for the mouth is 0.7, the matching result for the eyes is 0.75, and the second matching result is 1. The first weight for the mouth matching result is 0.3, the second weight for the eyes matching result is 0.4, and the third weight for the second matching result is 0.3. The final liveness detection probability is 0.81.
[0122] In some embodiments, the liveness detection probability is compared with a probability threshold to determine whether the current face recognition process is performed by a live person.
[0123] In some embodiments, when the liveness detection probability reaches a probability threshold, the liveness detection in the face recognition process is determined to be successful, and the face recognition result is obtained based on the successful liveness detection. The face recognition is performed by a face recognition model. The face recognition process and the liveness detection process can be performed by two parts of a single model, or by two independent models; this embodiment does not limit this.
[0124] In summary, the face recognition method provided in this application confirms changes in the facial features during speech by matching local facial regions with specified display content, and confirms the user's vocal content by matching audio content with specified display content. This enables liveness detection at both the facial change and vocal content levels, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video, thus improving the accuracy of face recognition and enhancing the security of the protection function provided by face recognition.
[0125] The method provided in this embodiment identifies the vocal content expressed by the user's lip movements through the mouth area of the face, and determines the liveness detection result based on the matching relationship between the vocal content and the specified display content, thus avoiding the problem that the lip movements do not correspond to the vocal content expressed in the actual speech.
[0126] The method provided in this embodiment identifies the user's gaze direction by using the eye region of the face, and matches the user's gaze direction with the actual display position of the specified display content. This avoids the attack problem where the user's gaze is not directed at the specified display content, but the actual speech content expressed is consistent with the specified display content.
[0127] Indicative, Figure 6 This is a schematic diagram of the overall process of face recognition provided in an exemplary embodiment of this application, using the display of random numbers on the interface as an example for illustration. Figure 6 As shown, the process includes the following steps:
[0128] Step 601: Enter the liveness detection interface.
[0129] In other words, when a user triggers the facial recognition function on the terminal, they enter the facial recognition interface, where a liveness detection is also performed. The liveness detection interface displays random numbers sequentially, and the display position of these random numbers is randomly determined.
[0130] Step 602 prompts the user to read the string of numbers aloud.
[0131] In some embodiments, voice prompts are displayed on the terminal interface to prompt the user to express the specified content displayed on the interface via voice.
[0132] Step 603: The microphone acquires the audio stream.
[0133] Audio streams are captured using the terminal microphone, and the user's voice expression in response to the specified displayed content is recorded.
[0134] Step 604: The camera acquires the video stream.
[0135] By capturing video streams through the terminal's camera, the facial features of the user are recorded when the user expresses their voice in response to specified displayed content.
[0136] Step 605: Extract voiceprint features.
[0137] Voiceprint features are extracted from the audio stream. Specifically, a pre-trained neural network model is used to extract voiceprint features from the audio stream.
[0138] Step 606: Load the neural network and obtain the speech recognition sequence.
[0139] Voiceprint features are loaded into a pre-trained neural network model to obtain a speech recognition sequence. The speech recognition sequence includes the speech content in the recognized speech expression, such as recognizing the numbers expressed by the user in sequence through speech.
[0140] Step 607: Obtain the first distance between the speech recognition result and the true value.
[0141] The truth value refers to the specified content actually displayed on the interface. That is, the first similarity between the speech recognition sequence and the content sequence in the specified display content is obtained by matching the speech recognition sequence with the content sequence in the specified display content.
[0142] Step 608: Locate the image sequence of a single digit reading interval.
[0143] Based on the aforementioned voiceprint features, the time period of a single digit reading is located, and the corresponding image sequence is located from the video stream.
[0144] Step 609: Obtain the mouth sub-image sequence.
[0145] The mouth area is a dynamically changing region during speech expression; therefore, mouth sub-image sequences are obtained to determine mouth shape.
[0146] Step 610: Use a neural network to obtain a lip-reading sequence.
[0147] The mouth sub-image sequence is input into a pre-trained neural network model to obtain a lip-reading sequence, which is to identify the lip-reading content corresponding to the user's mouth shape based on the mouth sub-image sequence.
[0148] Step 611: Obtain the second distance between the lip reading recognition result and the ground truth value.
[0149] The truth value refers to the specified content actually displayed on the interface. That is, the second similarity between the lip reading sequence and the content sequence in the specified display content is obtained by matching the lip reading sequence with the content sequence in the specified display content.
[0150] Step 612: Obtain the eye sub-image sequence.
[0151] Since the random numbers are displayed at random positions on the interface, the eye area is a dynamically changing region that observes the random numbers during the speech expression process. Therefore, the eye sub-image sequence is obtained to determine the eye gaze.
[0152] Step 613: Regress the line of sight orientation and obtain the line of sight change sequence.
[0153] The eye sub-image sequence is input into a pre-trained neural network model to obtain a gaze recognition sequence, which is to identify the orientation of the display screen corresponding to the user's gaze based on the eye sub-image sequence.
[0154] Step 614: Obtain the result of the line-of-sight change and the third distance of the true value.
[0155] The truth value refers to the actual display position of the specified content on the interface. In other words, the third similarity between the result of the change in gaze and the actual display position is obtained by matching the result of the change in gaze with the actual display position.
[0156] Step 615, Integration of the decision-making level.
[0157] The first similarity, second similarity and third similarity are weighted and fused through a preset decision layer to obtain the fusion result, which is the liveness detection probability, representing the probability that the current face recognition process is completed by a live person.
[0158] The weights used in the fusion process are obtained through model training or can be pre-set.
[0159] Step 616: Determine if it is greater than the confidence score.
[0160] The confidence score is preset. When the liveness detection probability is greater than the confidence score, the probability of the current face recognition being completed by a live person is relatively high. When the liveness detection probability does not reach the confidence score, the probability of the current face recognition being completed by a live person is relatively low.
[0161] Step 617: When the value is greater than the confidence score, the verification is successful.
[0162] Step 618: If the value is not greater than the confidence score, the verification fails.
[0163] In summary, the face recognition method provided in this application confirms changes in the facial features during speech by matching local facial regions with specified display content, and confirms the user's vocal content by matching audio content with specified display content. This enables liveness detection at both the facial change and vocal content levels, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video, thus improving the accuracy of face recognition and enhancing the security of the protection function provided by face recognition.
[0164] In some embodiments, the terminal side has a corresponding interface during the face recognition process. Figure 7 This is a flowchart of a face recognition method provided in an exemplary embodiment of this application. The method is illustrated using an example of its application in a terminal. Figure 7 As shown, the method includes:
[0165] Step 701: Display the face recognition interface.
[0166] The face recognition interface includes a face capture image, which refers to an image captured in real time by the terminal's camera. In some embodiments, the face recognition interface also includes a face reference frame, used to instruct the user to capture a face image within the range of the face reference frame.
[0167] Step 702: Display the specified content in the face recognition interface.
[0168] The specified display position for the content is determined from at least two candidate display positions.
[0169] As an illustration, candidate positions include the top left, bottom left, top right, and bottom right corners of the face recognition interface. When displaying specified content, the content will be randomly selected from these four corners for display.
[0170] In some embodiments, at least two specified display contents are sequentially switched in the face recognition interface based on preset switching conditions. The preset switching conditions include either an interval switching condition or a voice recognition switching condition. The interval switching condition identifies the display time interval between two adjacent specified display contents. The voice recognition switching condition indicates that when a voice expression for the k-th specified display content is recognized, the display switches to the (k+1)-th specified display content, where n is a positive integer.
[0171] Step 703: Display voice prompts on the face recognition interface.
[0172] Voice prompts are used to instruct users to provide corresponding voice responses to specified displayed content.
[0173] This is illustrative; please refer to it. Figure 8 It illustrates a schematic diagram of a face recognition interface provided in an exemplary embodiment of this application, such as... Figure 8 As shown, the face recognition interface 800 displays a specified display content 810, namely the number "3", and a voice prompt message 820, the content of which is "Please read out the number displayed on the screen".
[0174] Step 704: Display the face recognition result on the face recognition interface based on the captured face image and voice prompt information.
[0175] In some embodiments, when the liveness detection passes, the face recognition result is determined based on face recognition matching; when the liveness detection fails, it is directly determined that the face recognition failed.
[0176] In summary, the face recognition method provided in this application confirms changes in the facial features during speech by matching local facial regions with specified display content, and confirms the user's vocal content by matching audio content with specified display content. This enables liveness detection at both the facial change and vocal content levels, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video, thus improving the accuracy of face recognition and enhancing the security of the protection function provided by face recognition.
[0177] Figure 9 This is a structural block diagram of a face recognition device provided in an exemplary embodiment of this application, such as... Figure 9 As shown, the device includes:
[0178] The acquisition module 910 is used to acquire a face recognition video stream and audio content in response to the start of the face recognition process, wherein the face recognition video stream and the audio content are content acquired based on specified display content, and the display position of the specified display content is determined from at least two candidate display positions;
[0179] The acquisition module 910 is further configured to extract a partial facial region from the face recognition video stream, wherein the partial facial region is the region corresponding to the facial features that are expressed when the specified display content is spoken.
[0180] The matching module 920 is used to match the partial area of the face with the specified display content to obtain a first matching result;
[0181] The matching module 920 is further configured to match the audio content with the specified display content to obtain a second matching result;
[0182] The determination module 930 is used to determine the face recognition result based on the first matching result and the second matching result.
[0183] In an optional embodiment, such as Figure 10 As shown, the acquisition module 910 includes:
[0184] The positioning unit 911 is used to locate the image sequence corresponding to the display of the specified display content from the face recognition video stream;
[0185] The cropping unit 912 is used to crop the image frames in the image sequence to obtain the local area of the face.
[0186] In an optional embodiment, the positioning unit 911 is further configured to extract audio features of the audio content; determine the time period for the speech expression of the specified display content based on the audio features; and locate the image sequence corresponding to the time period from the face recognition video stream.
[0187] In an optional embodiment, the face recognition process includes at least two sequentially displayed specified content items;
[0188] The positioning unit 911 is further configured to determine the i-th time period for the voice expression of the i-th specified display content based on the audio features, where i is a positive integer; and to locate the i-th group of image sequences corresponding to the i-th time period from the face recognition video stream.
[0189] In an optional embodiment, the face local area includes a face mouth area, which is a region obtained by cropping the mouth that expresses the specified display content, and the first matching result includes a mouth matching result;
[0190] The matching module 920 includes:
[0191] Recognition unit 921 is used to perform mouth shape recognition on the mouth area of the face to obtain lip reading recognition result, and the lip reading recognition result is used to represent the vocal content of the mouth that has been recognized.
[0192] The matching unit 922 is used to match the lip reading result with the specified display content to obtain the mouth matching result, which is used to represent the correlation between the lip reading result and the specified display content.
[0193] In an optional embodiment, the partial facial region includes a facial eye region, which is a region obtained by cropping the eyes that are observing the specified display content, and the first matching result includes an eye matching result;
[0194] The matching module 920 includes:
[0195] The recognition unit 921 is used to perform gaze recognition on the eye region of the face and obtain a gaze recognition result, which is used to indicate the gaze direction of the eye region.
[0196] The matching unit 922 is used to match the gaze recognition result with the display position of the specified display content to obtain the eye matching result, which is used to represent the correlation between the gaze recognition result and the specified display content.
[0197] In an optional embodiment, the matching module 920 includes:
[0198] Recognition unit 921 is used to perform speech recognition on the audio content and obtain speech recognition results;
[0199] The matching unit 922 is used to match the speech recognition result with the specified display content to obtain the second matching result, which is used to represent the degree of correlation between the speech recognition result and the specified display content.
[0200] In an optional embodiment, the face recognition process includes at least two sequentially displayed specified display contents, and the speech recognition result includes a speech recognition sequence, which includes at least two sequentially arranged recognition sub-results;
[0201] The matching unit 922 is further configured to perform content matching between the recognition sub-result in the speech recognition sequence and the at least two specified display contents to obtain a content matching result; perform sequential matching between the recognition sub-result in the speech recognition sequence and the at least two specified display contents to obtain a sequential matching result; and obtain a second matching result based on the content matching result and the sequential matching result.
[0202] In an optional embodiment, the determining module 930 is further configured to perform a weighted summation of the first matching result and the second matching result to obtain a liveness detection probability, wherein the liveness detection probability is used to represent the probability that the face recognition process is completed by a live person; and to determine the face recognition result based on the liveness detection probability.
[0203] In an optional embodiment, the display position of the specified display content is randomly determined from at least two candidate display positions;
[0204] The specified display content is randomly determined from a preset content library.
[0205] In an optional embodiment, this application also provides a face recognition device, the device comprising:
[0206] The display module is used to display the face recognition interface, which includes a captured face image.
[0207] The display module is also used to display specified display content in the face recognition interface, wherein the display position of the specified display content is determined from at least two candidate display positions;
[0208] The display module is also used to display voice prompt information in the face recognition interface, the voice prompt information being used to instruct the corresponding voice expression for the specified display content;
[0209] The display module is also used to display the face recognition result on the face recognition interface based on the captured face image and the voice prompt information.
[0210] In an optional embodiment, the display module is further configured to sequentially switch the display of at least two specified display contents in the face recognition interface based on preset switching conditions;
[0211] The preset switching conditions include interval switching conditions and voice recognition switching conditions;
[0212] The interval switching condition refers to the display time interval between two adjacent specified display contents; the voice recognition switching condition refers to switching to displaying the (k+1)th specified display content when a voice expression for the kth specified display content is recognized, where k is a positive integer.
[0213] In summary, the face recognition device provided in this application confirms changes in the face during speech by matching a local area of the face with designated display content, and confirms the user's voice content by matching the audio content with designated display content. This enables liveness detection at both the face change and voice content levels, ensuring that the user participating in the face recognition process is a live user, rather than a flat image or video, thus improving the accuracy of face recognition and enhancing the security of the protection function provided by face recognition.
[0214] It should be noted that the face recognition device provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the face recognition device provided in the above embodiments belongs to the same concept as the face recognition method embodiments, and its specific implementation process can be found in the method embodiments, which will not be repeated here.
[0215] Figure 11 A structural block diagram of an electronic device 1100 provided in an exemplary embodiment of this application is shown. The electronic device 1100 may be a portable mobile terminal, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The electronic device 1100 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0216] Typically, electronic device 1100 includes a processor 1101 and a memory 1102.
[0217] Processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1101 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1101 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1101 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1101 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0218] The memory 1102 may include one or more computer-readable storage media, which may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 are used to store at least one instruction, which is executed by the processor 1101 to implement the face recognition method provided in the method embodiments of this application.
[0219] In some embodiments, the electronic device 1100 may also optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, memory 1102, and peripheral device interface 1103 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1103 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1104, a display screen 1105, a camera assembly 1106, an audio circuit 1107, a positioning assembly 1108, and a power supply 1109.
[0220] Peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1101 and memory 1102. In some embodiments, processor 1101, memory 1102 and peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1101, memory 1102 and peripheral device interface 1103 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0221] The radio frequency (RF) circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1104 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1104 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1104 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1104 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0222] Display screen 1105 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1105 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1101 for processing. In this case, display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, disposed on the front panel of electronic device 1100; in other embodiments, there may be at least two display screens, disposed on different surfaces of electronic device 1100 or in a folded design; in still other embodiments, display screen 1105 may be a flexible display screen, disposed on a curved or folded surface of electronic device 1100. Furthermore, display screen 1105 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1105 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0223] The camera assembly 1106 is used to acquire images or videos. Optionally, the camera assembly 1106 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1106 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0224] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1101 for processing, or input to the radio frequency circuit 1104 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the electronic device 1100. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1107 may also include a headphone jack.
[0225] Positioning component 1108 is used to locate the current geographical location of electronic device 1100 in order to enable navigation or LBS (Location Based Service). Positioning component 1108 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, or Russia's Galileo system.
[0226] Power supply 1109 is used to supply power to various components in electronic device 1100. Power supply 1109 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When power supply 1109 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0227] In some embodiments, the electronic device 1100 further includes one or more sensors 1110. The one or more sensors 1110 include, but are not limited to: an accelerometer 1111, a gyroscope 1112, a pressure sensor 1113, a fingerprint sensor 1114, an optical sensor 1115, and a proximity sensor 1116.
[0228] Accelerometer 1111 can detect the magnitude of acceleration on the three coordinate axes of a coordinate system established by electronic device 1100. For example, accelerometer 1111 can be used to detect the components of gravitational acceleration on the three coordinate axes. Processor 1101 can control display screen 1105 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1111. Accelerometer 1111 can also be used for games or for acquiring user motion data.
[0229] The gyroscope sensor 1112 can detect the orientation and rotation angle of the electronic device 1100. The gyroscope sensor 1112 can work in conjunction with the accelerometer sensor 1111 to collect the user's 3D movements on the electronic device 1100. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0230] Pressure sensor 1113 can be disposed on the side bezel of electronic device 1100 and / or on the lower layer of display screen 1105. When pressure sensor 1113 is disposed on the side bezel of electronic device 1100, it can detect the user's grip signal on electronic device 1100, and processor 1101 can perform left / right hand recognition or quick operation based on the grip signal collected by pressure sensor 1113. When pressure sensor 1113 is disposed on the lower layer of display screen 1105, processor 1101 can control operable controls on UI interface based on the user's pressure operation on display screen 1105. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0231] The fingerprint sensor 1114 is used to collect a user's fingerprint. The processor 1101 identifies the user based on the fingerprint collected by the fingerprint sensor 1114, or vice versa. When the user's identity is identified as trusted, the processor 1101 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 1114 can be located on the front, back, or side of the electronic device 1100. When the electronic device 1100 has a physical button or manufacturer logo, the fingerprint sensor 1114 can be integrated with the physical button or manufacturer logo.
[0232] An optical sensor 1115 is used to collect ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 based on the ambient light intensity collected by the optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 based on the ambient light intensity collected by the optical sensor 1115.
[0233] The proximity sensor 1116, also known as a distance sensor, is typically located on the front panel of the electronic device 1100. The proximity sensor 1116 is used to detect the distance between the user and the front of the electronic device 1100. In one embodiment, when the proximity sensor 1116 detects that the distance between the user and the front of the electronic device 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from a screen-on state to a screen-off state; when the proximity sensor 1116 detects that the distance between the user and the front of the electronic device 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from a screen-off state to a screen-on state.
[0234] Those skilled in the art will understand that Figure 11 The structure shown does not constitute a limitation on the electronic device 1100, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0235] Embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the face recognition method provided in the above-described method embodiments.
[0236] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the face recognition method provided in the above-described method embodiments.
[0237] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the face recognition methods described in the above embodiments.
[0238] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0239] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0240] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A face recognition method, characterized in that, The method includes: In response to the start of the face recognition process, a face recognition video stream and audio content are acquired, wherein the face recognition video stream and the audio content are content acquired based on specified display content, and the display position of the specified display content is determined from at least two candidate display positions; A partial facial region is extracted from the facial recognition video stream, wherein the partial facial region is the area corresponding to the facial features that are expressed when the specified display content is spoken. The facial region is matched with the specified display content to obtain a first matching result; The audio content is matched with the specified display content to obtain a second matching result; The face recognition result is determined based on the first matching result and the second matching result; The partial facial region includes the facial eye region, which is the region obtained by cropping the eyes that are observing the specified display content. The first matching result includes the eye matching result. The step of matching the partial facial region with the specified display content to obtain a first matching result includes: The eye region of the face is subjected to gaze recognition to obtain a gaze recognition result, which is used to indicate the gaze direction of the eye region. The eye-tracking recognition result is matched with the display position of the specified display content to obtain the eye matching result, which is used to represent the correlation between the eye-tracking recognition result and the specified display content.
2. The method according to claim 1, characterized in that, Extracting a partial facial region from the facial recognition video stream includes: Locate the image sequence corresponding to the display of the specified content from the face recognition video stream; The image frames in the image sequence are cropped to obtain the local area of the face.
3. The method according to claim 2, characterized in that, The step of locating the image sequence corresponding to the display of the specified content from the face recognition video stream includes: Extract the audio features of the audio content; The time period for the speech expression of the specified display content is determined based on the audio features; Locate the image sequence corresponding to the time period from the face recognition video stream.
4. The method according to claim 3, characterized in that, The face recognition process includes at least two designated display contents displayed sequentially; The step of determining the time period for voice expression of the specified displayed content based on the audio features includes: Based on the audio features, the i-th time period for the speech expression of the i-th specified display content is determined, where i is a positive integer; The step of locating the image sequence corresponding to the time period from the face recognition video stream includes: Locate the i-th image sequence corresponding to the i-th time period from the face recognition video stream.
5. The method according to any one of claims 1 to 4, characterized in that, The step of matching the audio content with the specified display content to obtain a second matching result includes: The audio content is subjected to speech recognition to obtain the speech recognition result; The speech recognition result is matched with the specified display content to obtain the second matching result, which is used to represent the degree of correlation between the speech recognition result and the specified display content.
6. The method according to claim 5, characterized in that, The face recognition process includes at least two specified display contents displayed sequentially, and the speech recognition result includes a speech recognition sequence, which includes at least two sequentially arranged recognition sub-results. The step of matching the speech recognition result with the specified display content to obtain the second matching result includes: The recognition sub-result in the speech recognition sequence is matched with the at least two specified display contents to obtain a content matching result; The recognition sub-results in the speech recognition sequence are sequentially matched with the at least two specified display contents to obtain a sequential matching result; The second matching result is obtained based on the content matching result and the order matching result.
7. The method according to any one of claims 1 to 4, characterized in that, The step of determining the face recognition result based on the first matching result and the second matching result includes: The first matching result and the second matching result are weighted and summed to obtain the liveness detection probability, which is used to represent the probability that the face recognition process is completed by a live person; The face recognition result is determined based on the liveness detection probability.
8. The method according to any one of claims 1 to 4, characterized in that, The display position of the specified display content is randomly determined from at least two candidate display positions; The specified display content is randomly determined from a preset content library.
9. A face recognition method, characterized in that, The method includes: Display a face recognition interface, which includes a captured face image; The specified display content is displayed in the face recognition interface, and the display position of the specified display content is determined from at least two candidate display positions; Voice prompts are displayed on the face recognition interface, and these voice prompts are used to instruct the user to express the specified content in voice. The face recognition result is displayed on the face recognition interface based on the captured face image and the voice prompt information, and the face recognition result is determined according to the face recognition method according to claim 1.
10. The method according to claim 9, characterized in that, The step of displaying specified content on the face recognition interface includes: In the face recognition interface, at least two of the specified display contents are sequentially switched and displayed based on preset switching conditions; The preset switching conditions include interval switching conditions and voice recognition switching conditions; The interval switching condition refers to the display time interval between two adjacent specified display contents; the voice recognition switching condition refers to switching to displaying the (k+1)th specified display content when a voice expression for the kth specified display content is recognized, where k is a positive integer.
11. A face recognition device, characterized in that, The device includes: The acquisition module is used to acquire a face recognition video stream and audio content in response to the start of the face recognition process, wherein the face recognition video stream and the audio content are content acquired based on specified display content, and the display position of the specified display content is determined from at least two candidate display positions; The acquisition module is further configured to extract a partial facial region from the facial recognition video stream, wherein the partial facial region is the region corresponding to the facial features that are expressed when the specified display content is spoken. The matching module is used to match the local area of the face with the specified display content to obtain a first matching result; The matching module is further configured to match the audio content with the specified display content to obtain a second matching result; The determining module is used to determine the face recognition result based on the first matching result and the second matching result. The partial facial region includes the facial eye region, which is the region obtained by cropping the eyes that are observing the specified display content. The first matching result includes the eye matching result. The step of matching the partial facial region with the specified display content to obtain a first matching result includes: The eye region of the face is subjected to gaze recognition to obtain a gaze recognition result, which is used to indicate the gaze direction of the eye region. The eye-tracking recognition result is matched with the display position of the specified display content to obtain the eye matching result, which is used to represent the correlation between the eye-tracking recognition result and the specified display content.
12. A face recognition device, characterized in that, The device includes: The display module is used to display the face recognition interface, which includes a captured face image. The display module is also used to display specified display content in the face recognition interface, wherein the display position of the specified display content is determined from at least two candidate display positions; The display module is also used to display voice prompt information in the face recognition interface, the voice prompt information being used to instruct the corresponding voice expression for the specified display content; The display module is further configured to display the face recognition result on the face recognition interface based on the face acquisition image and the voice prompt information, wherein the face recognition result is determined according to the face recognition method according to claim 1.
13. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the face recognition method as described in any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or instruction set is loaded and executed by a processor to implement the face recognition method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Face recognition method and recognition system
CN104966053A
Video processing method, device and system, terminal equipment and storage medium
CN110688911A