Face image processing method, device, apparatus, storage medium and program product
By capturing and magnifying facial images in a facial recognition system, and guiding users to increase their distance to capture background information, the system identifies information from a second terminal device, thus solving the security problem of high-definition video bypassing facial recognition and improving recognition accuracy and system security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ANT KUAI TECHNOLOGY CO LTD
- Filing Date
- 2026-02-27
- Publication Date
- 2026-07-07
AI Technical Summary
In existing technologies, facial recognition systems have limited ability to defend against risky behaviors during high-definition video playback, making them vulnerable to attacks from pre-recorded videos and leading to information security issues.
The system captures facial images using the camera of the first terminal device, magnifies them, and displays a partial facial image in the preview area. This guides the user to increase the distance between their face and the camera, creating conditions for the camera to capture background information around the face. During this transition, the system identifies information from the second terminal device to determine if a fake face is present.
It improves the accuracy and security of facial recognition, effectively defends against risky behaviors during high-definition video playback, and enhances the system's resistance to attacks.
Smart Images

Figure CN122347820A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of information processing technology, and in particular to a method, device, apparatus, storage medium, and program product for processing facial images. Background Technology
[0002] With the widespread adoption of facial recognition technology on mobile devices, its security faces increasingly serious challenges. For example, during facial recognition on one terminal device, a user can bypass the facial recognition system by using the screen of another terminal device to display a high-definition image or video of the user to be authenticated.
[0003] Currently, requiring users to perform certain facial recognition actions, such as blinking, opening their mouth, and shaking their head, can identify some injection risks in static images. However, its defense capabilities against risks in high-definition video playback are limited, and it is easily compromised by pre-recorded videos, leading to information security issues. Summary of the Invention
[0004] This specification provides a face image processing method, device, apparatus, storage medium, and program product to improve the accuracy of identifying risky behaviors and enhance the security of face recognition.
[0005] This specification provides a face image processing method applied to a first terminal device, comprising: responding to a face recognition trigger operation and displaying a face acquisition page, the face acquisition page including a preview area; acquiring a first face image using a camera on the first terminal device, the first face image including a target face; enlarging the first face image to display a partial face image in the preview area, and outputting guidance information to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image; during the transition, if the face image displayed in the preview area meets the set first face condition, acquiring a second face image using the camera, the second face image including the target face and background information around the target face; if the background information in the second face image includes information of the second terminal device, determining that the target face is a fake face.
[0006] This specification also provides a face image processing device, including: a display module, a capture module, a magnification module, an output module, and a determination module; the display module is used to respond to a face recognition trigger operation and display a face capture page, which includes a preview area; the capture module is used to capture a first face image using a camera on a first terminal device, the first face image including a target face; the magnification module is used to magnify the first face image to display a partial face image in the preview area; the output module is used to output guidance information to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image; the capture module is used to, during the transition process, if the face image displayed in the preview area meets the set first face condition, to capture a second face image using the camera, the second face image including the target face and background information around the target face; the determination module is used to determine that the target face is a fake face if the background information in the second face image includes information from the second terminal device.
[0007] This specification also provides an electronic device, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute one or more computer instructions to perform the steps in the method provided in this specification.
[0008] This specification also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps of the method provided in this specification.
[0009] This specification also provides a computer program product, including: a computer program / instructions, which, when executed by a processor, can implement the steps of the method provided in this specification.
[0010] In this embodiment, a face image is captured by the camera of a first terminal, and the face image is magnified. The magnified partial face image is displayed in the preview area, and the user is guided to increase the distance between the face and the camera. This causes the preview area to transition from displaying a partial face image to displaying a complete face image, actively creating conditions for the camera to capture background information around the face. During the transition, if the face image displayed in the preview area meets the set face conditions, the camera captures an image including background information around the face. If the background information includes information from the second terminal, it indicates that the currently captured face image may come from a high-definition image or video played by the second terminal. In this case, the face in the face image is determined to be a fake face, indicating a risky behavior in the face recognition process, thus improving the accuracy and security of face recognition. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a schematic flowchart of a face image processing method provided as an exemplary embodiment of this specification.
[0012] Figure 2 This is a schematic diagram of the interactive process of a face image processing system provided in one embodiment of this specification.
[0013] Figure 3 This is a schematic diagram of the structure of a face image processing apparatus provided for an exemplary embodiment of this specification.
[0014] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this specification. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0016] It should be noted that, in the cases involving user information in this application's embodiments, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. The various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0017] To address the aforementioned technical issues, in some embodiments of this specification, a facial image is captured by the camera of a first terminal, magnified, and displayed in a preview area as a partial image of the face. The user is guided to increase the distance between their face and the camera, causing the preview area to transition from displaying a partial image to displaying a complete image. This proactively creates conditions for the camera to capture background information surrounding the face. During this transition, if the facial image displayed in the preview area meets the set facial criteria, the camera captures an image including background information surrounding the face. If the background information includes information from a second terminal, it indicates that the currently captured facial image may be from a high-definition image or video played by the second terminal. In this case, the face in the image is determined to be a fake face, indicating a risky behavior in the facial recognition process, thus improving the accuracy and security of facial recognition.
[0018] In addition, by separating image preview from image processing, the first face image after magnification is displayed in the preview area, guiding the user to increase the distance between the face and the camera. As the distance increases, a second face image is captured, and image processing is performed on the second face image to identify risky behaviors and achieve concealment of defensive intentions, which greatly improves the security and resistance to attacks of the solution.
[0019] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0020] Figure 1 This is a schematic flowchart of a face image processing method provided in an exemplary embodiment of this specification, such as... Figure 1 The method shown includes: 101. Respond to the face recognition trigger operation and display the face capture page, which includes a preview area.
[0021] 102. Use the camera on the first terminal device to capture a first face image, the first face image including the target face.
[0022] 103. Enlarge the first face image to display a partial face image in the preview area, and output guiding information to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image.
[0023] 104. During the transition, if the face image displayed in the preview area meets the set first face condition, the camera is used to capture the second face image, which includes the target face and the background information around the target face.
[0024] 105. If the background information in the second face image includes information about the second terminal device, the target face is determined to be a fake face.
[0025] In the embodiments of this specification, the implementation method of the first terminal device is not limited. For example, the first terminal device may include, but is not limited to, terminal devices with facial recognition requirements such as smartphones, tablets, facial recognition access control devices, elevator facial recognition devices, and facial payment devices.
[0026] In the embodiments of this specification, the implementation method of the first terminal device responding to a face recognition trigger operation and displaying a face collection page is not limited. For example, the first terminal device may have a target application installed, which could be a financial payment application, a government service application, or a digital identity verification application, etc. The application page of the target application displays controls such as face recognition login, face recognition authentication, or face recognition door opening. When these controls are triggered, a face recognition trigger operation is initiated, and the first terminal device responds to the face recognition trigger operation, displaying the face collection page in the target application. Alternatively, the face recognition trigger operation can be initiated via Near Field Communication (NFC), QR code, or card swiping, and the first terminal device responds to the face recognition trigger operation, displaying the face collection page in the target application.
[0027] In the embodiments of this specification, the face acquisition page includes a preview area. The preview area is used to display the acquired face image during the face image acquisition process, such as a magnified partial face image. The shape of the preview area is not limited. For example, the shape of the preview area may include, but is not limited to, circles, squares, rectangles, and ellipses.
[0028] In the embodiments of this specification, a first face image is captured using a camera on a first terminal device. The first face image includes a target face. This target face may be obtained by the camera capturing a real human face, or it may be a face displayed in a high-definition video played by a second terminal device. The target face in the first face image can be a complete face or a partial face; there is no limitation on this. Preferably, the target face in the first face image is a complete face. The camera used to capture the first face image is not limited. For example, the camera capturing the first face image can be a front-facing camera on the first terminal device. The front-facing camera can be a fixed-focus wide-angle lens with a fixed physical focal length (or optical focal length). Alternatively, the camera capturing the first face image can also be a rear-facing camera on the first terminal device. The rear-facing camera can include, but is not limited to, a main camera, an ultra-wide-angle lens, and a telephoto lens. For example, the main camera has an optical zoom of 1x, the ultra-wide-angle lens has an optical zoom of 0.5x, and the telephoto lens has optical zooms of 2x, 3x, 5x, and 10x, with each telephoto lens corresponding to a fixed optical zoom.
[0029] In the embodiments of this specification, the first face image is magnified to display a partial face image in the preview area. The partial face image includes a portion of the target face. For example, a partial face indicates that the complete face is not presented; a complete face includes eyes, nose, mouth, and a complete outline. For example, a partial face may include, but is not limited to: "the first face image includes the user's nose and mouth," "the first face image includes half of the user's face," and "the first face image includes the top of the head or chin," etc. The facial features displayed in the preview area are fewer than those in the first face image. For example, the first face image includes a complete face, but the preview area displays a portion of the face. Another example is that the first face image includes a partial face, such as eyes, mouth, nose, and chin, but does not include the forehead, but the preview area displays the mouth and nose.
[0030] In the case of a single first face image, the preview area can display a partial face image. If there are multiple consecutive first face images, the preview area can continuously display multiple partial face images to form a video stream.
[0031] In the embodiments of this specification, guidance information is output. This guidance information is used to instruct the user to increase the distance between the target face and the camera, causing the preview area to transition from displaying a partial view of the face to displaying the full face. This guidance information can be output through one or more of the following methods: voice, text, image, light, and touch.
[0032] In the embodiments described in this specification, the user, guided by the instructions, increases the distance between their face and the camera. In non-risk scenarios, the user moves their phone away from their face. In risky scenarios, the second terminal device playing high-definition video moves away from the camera of the first terminal device. During this process of increasing the distance between the face and the camera, the camera continuously captures and magnifies the first face image. As the distance between the face and the camera increases, the proportion of the face in the preview area gradually decreases, the camera's visible area gradually increases, and the preview area gradually transitions from displaying a "partial face image" to displaying a "complete face image."
[0033] In the embodiments of this specification, as the user increases the distance between their face and the camera, it is determined whether the face image displayed in the preview area meets the set first face condition. If the face image displayed in the preview area does not meet the set first face condition, guidance information continues to be output until the face image in the preview area meets the set first face condition. If the face image displayed in the preview area meets the set first face condition, a second face image is captured using the camera. The first face condition may include, but is not limited to: the proportion of the face in the preview area meets a set proportion requirement, the face size meets a set size requirement, the clarity exceeds a set clarity threshold, the face is not obscured, and a complete face is detected.
[0034] In this case, when the user increases the distance between their face and the camera, the camera's visible area increases. Therefore, the second face image captured by the camera includes not only the face but also the background information around the face.
[0035] The definition of the area surrounding the face is not limited. For example, it can be an area centered on the face and expanded outwards by a specified multiple, which can include, but is not limited to, 1.5x, 2.0x, or 2.5x. The greater the distance between the user's face and the camera, the more background information is available around the face. This background information can include, but is not limited to, walls, access control panels, sky, trees, other people, moving objects, light sources, reflections, shadows, indoor furniture, billboards, and text signs.
[0036] In the embodiments of this specification, during the face recognition process performed by the first terminal device for the user to be verified, the real face recognition for identity verification can be as follows: the user to be verified faces the first terminal device, and the camera of the first terminal device captures the face image of the user to be verified, and identity verification is performed based on the face image. In the identity verification process involving risky behavior, a second user can play a high-definition video of the user to be verified through a second terminal device, causing the camera of the first terminal device to capture the face image in the high-definition video, thereby bypassing identity verification. In this case, the second terminal device is considered a risky device.
[0037] Therefore, if the background information in the second face image includes information from the second terminal device, and the target face in either the first or second face image is determined to be a fake face, or if the first or second face image is a forged face, then the face recognition process carries a risk. For example, the first and second face images might be forged images obtained by playing a high-definition video containing the target face on the screen of the second terminal device; in other words, the entire face recognition process is susceptible to presentational attacks. The second terminal device might be a device that plays a high-definition image or video of the user to be verified during the authentication process. The implementation of the second terminal device can include, but is not limited to, television screens, smart screens, smartphones, tablets, laptops, and desktop computers. The information of the second terminal device can include, but is not limited to, border information (e.g., rectangular lines or straight line edge information), screen-specific reflections, moiré patterns, etc.
[0038] In the embodiments of this specification, a face image is captured by the camera of the first terminal, the face image is magnified, and a magnified partial face image is displayed in the preview area. The user is guided to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image, actively creating shooting conditions for the camera to capture background information around the face. During the transition, if the face image displayed in the preview area meets the set face conditions, the camera captures an image including background information around the face. If the background information includes information from the second terminal, it indicates that the currently captured face image comes from a high-definition image or high-definition video played by the second terminal. In this case, the face in the face image is determined to be a fake face, and the face recognition process involves risky behavior, thereby improving the accuracy and security of face recognition.
[0039] In one optional embodiment, the method of magnifying the first face image to display a partial face image in the preview area is not limited. An exemplary description follows.
[0040] In one example, the first face image is magnified while the size of the preview area remains unchanged. Therefore, the magnified partial face image is displayed in the preview area.
[0041] In another example, a pre-set first digital zoom ratio is obtained, and the first face image is magnified based on the first digital zoom ratio to obtain a partial face image. Based on this, an implementation method for magnifying a first face image to display a partial face image in a preview area includes: obtaining a first digital zoom ratio, where the first digital zoom ratio represents a cropping size; the cropping size represents the pixel size of the cropped area, which corresponds to the cropped partial face area; cropping the first face image based on the cropping size represented by the first digital zoom ratio to obtain a partial face area; interpolating the partial face area to obtain a partial face image with the same size as the first face image, and displaying the partial face image in the preview area.
[0042] The first digital zoom ratio can be preset based on empirical values. Any digital zoom ratio that magnifies the first face image and allows a portion of the face to be displayed in the preview area after magnification is applicable. For example, the first digital zoom ratio can include, but is not limited to, 1.5x or 2.0x. A first digital zoom ratio greater than 1.0x results in the face being cropped and magnified in the preview area, meaning the field of view is reduced.
[0043] The calculation method for the cropping size represented by the first digital zoom ratio is not limited. For example, the pixel size of the first face image is W×H, where W represents the width and H represents the height. The first digital zoom ratio is represented by Z, and the width and height of the cropping size are represented as W / Z and H / Z, respectively. The cropping method can include, but is not limited to, center cropping, center-left cropping, center-right cropping, center-top cropping, and center-bottom cropping, etc., and is not limited thereto.
[0044] Interpolation refers to using image scaling algorithms (such as bilinear interpolation or bicubic interpolation) to enlarge a cropped local face region (with a small pixel size) to the same pixel size as the first face image, thereby generating a local face image.
[0045] For example, if the resolution of the local face region is 5×5 and the pixel size of the first face image is 10×10, a blank image with a pixel size of 10×10 is created. Based on the pixel size of the local face region and the pixel size of the first face image, the magnification ratio is determined. For any pixel in the blank image, this pixel is mapped to the local face region according to the magnification ratio to obtain the target pixel corresponding to that pixel. For example, if the Nearest Neighbor Interpolation algorithm is used for interpolation in the local face region, the nearest neighbor pixel of the target pixel is found in the local face region, and the pixel value of this nearest neighbor pixel is assigned to that pixel in the blank image to form a local face image. As another example, if bilinear interpolation is used for interpolation in the local face region, four candidate pixels surrounding the target pixel are found in the local face region, and the weighted average of the pixel values of these four candidate pixels is calculated as the pixel value of that pixel in the blank image, thus obtaining the local face image. For example, if bicubic interpolation is used to interpolate in a local face region, 16 candidate pixels around the target pixel are found in the local face region, and the weighted average of the pixel values of the 16 candidate pixels is calculated as the pixel value of any pixel in the blank image, thereby obtaining a local face image.
[0046] Optionally, before interpolating the local face region, the method further includes: identifying whether the local face region meets the set second face condition; if the local face region does not meet the set second face condition, increasing the first digital zoom ratio and re-cropping the first face image according to the cropping size represented by the increased first digital zoom ratio, until the cropped local face region meets the set second face condition.
[0047] The second face condition may include, but is not limited to: the proportion of a face part (such as a nose or eyes) in the second sub-region is greater than a set proportion threshold, no face outline edge is detected in the second sub-region and the second sub-region includes a face part, and the second sub-region includes a part of a face, etc.
[0048] In one alternative embodiment, with the optical focal length of the camera remaining unchanged, the camera can use the same digital zoom ratio to capture the first face image and the second face image, or it can use different digital zoom ratios.
[0049] For example, in one embodiment, the first face image and the second face image can use the same second digital zoom ratio. Based on this, a method for capturing the first face image using a camera on a first terminal device includes: capturing the first face image based on the second digital zoom ratio corresponding to the camera; an implementation method for capturing the second face image using a camera includes: capturing the second face image based on the second digital zoom ratio corresponding to the camera. Wherein, the second digital zoom ratio is less than the first digital zoom ratio. For example, if the first digital zoom ratio is 1.5x or 2.0x, the third digital zoom ratio is 0.9x, 1.0x, or 1.1x, etc.
[0050] In another example, the first and second face images can use different digital zoom ratios. For instance, the first face image is captured using a second digital zoom ratio, and the second face image is captured using a third digital zoom ratio. The third digital zoom ratio is less than the first digital zoom ratio.
[0051] In one optional embodiment, the implementation method for determining whether the face image displayed in the preview area meets the set first face condition is not limited, and an exemplary description is given below.
[0052] For example, if a complete face is detected in the preview area, the face image displayed in the preview area is determined to meet the set first face condition. As another example, for the detected face features displayed in the preview area, these features may include, but are not limited to, the distance between the eyes, eyes, nose, or mouth; if the face feature decreases from a first size to a second size, and the second size is less than a set size threshold, then the face image displayed in the preview area is determined to meet the set first face condition, where the first size is greater than the size threshold. For example, the size of the face feature can be a physical size or a pixel size, without limitation. The size threshold can be 1cm or 2cm, or it can be 5px or 8px, etc. As yet another example, for the detected face features displayed in the preview area, if the face feature decreases from a first size to a second size, and the difference between the second size and the first size is greater than a set difference threshold, then the face image displayed in the preview area is determined to meet the set first face condition.
[0053] In one optional embodiment, the method for determining whether the target face is a fake face is not limited if the background information of the second terminal device appears in the background information of the second face image. For example, by calling a face image processing model to extract the background information around the face from the second face image, if the background information includes the information of the second terminal device, then the target face is determined to be a fake face, and the face recognition process involves risky behavior. If the background information does not include the information of the second terminal device, then the target face is determined to be a real face, and the face recognition process does not involve risky behavior at this time, and subsequent face recognition processes can continue to be executed, for example, to confirm whether the face in the second face image is the face of the user to be verified.
[0054] Optionally, the information of the second terminal device includes any one of border information, environmental information, and texture information. For example, border information may include, but is not limited to, rectangular borders, straight lines, and right angles; environmental information may include, but is not limited to, screen reflections; and texture information may include, but is not limited to, moiré patterns. Regardless of the screen resolution of the second terminal device, the physical border of the second terminal device always exists. Therefore, the method provided in the embodiments of this application has strong robustness against risky behaviors when playing on high-definition and ultra-high-definition screens.
[0055] The deployment location of the facial image processing model is not limited. For example, the model can be deployed on a first terminal device. The first terminal device can directly call the model to extract background information around the face from the second facial image. If the background information includes information from the second terminal device, a risky behavior is identified in the facial recognition process. Alternatively, the model can be deployed on a server-side device. The first terminal device sends the second facial image to the server-side device. The server-side device calls the model to extract background information around the face. If the background information includes information from the second terminal device, a first notification indicating the target face is a fake face is generated and sent to the first terminal device. Upon receiving the first notification from the server-side device, the first terminal device is redirected from the facial capture page to a new page displaying an authentication failure message. If the background information does not include information from the second terminal device, a second notification indicating the target face is a real face is generated and sent to the first terminal device. Upon receiving the second notification from the server-side device, the first terminal device continues with the subsequent facial recognition process.
[0056] In the embodiments of this specification, the number of model parameters supported by the face image processing model is not limited, with the goal of meeting application requirements. If there are relatively more model parameters, the face image processing model will be relatively larger in scale and have relatively better performance. Of course, it will consume more time and resources during inference and training. If there are relatively fewer model parameters, the face image processing model will be relatively smaller in scale. Under the condition that the performance requirements are met, the model is more lightweight and consumes relatively less time and resources during inference and training.
[0057] The face image processing model can be implemented as a classification model. The input to the model is a second face image, and the output can be the probability of the presence of risky behavior and the probability of the absence of risky behavior. Sample data and its annotation results can be pre-constructed. The sample data includes positive and negative samples; positive samples are labeled as indicating the absence of risky behavior, and negative samples are labeled as indicating the presence of risky behavior. The positive and negative samples, along with their respective annotation results, are used as training samples to train the face image processing model.
[0058] The model architecture for the face image processing model is not limited. For example, a face image processing model may include: a feature extraction layer, an attention layer, a multi-scale feature fusion layer, and a classification layer. Edge maps, spectrograms, and spectrophotomasks are generated based on the second face image. The edge map is a binary image extracted from the second face image using an edge detection algorithm. White pixels represent detected significant edges, and black pixels represent non-edge regions. The edge map guides the model to focus on structural anomalies; for example, real face edges are soft and predominantly curved, while structural anomalies refer to rectangles, straight lines, or right angles around the face in the image. The spectrogram is an amplitude spectrum obtained by converting the second face image from the spatial domain to the frequency domain, reflecting the energy distribution of different frequency components in the second face image. Moiré patterns appear as regular bright spots or highly directional energy stripes in the spectrogram; therefore, the spectrogram guides the model to detect moiré patterns in the background information. For example, the spectral energy of a real face is concentrated in the center (low frequency), while the spectrogram of a screen playback image shows high-frequency peaks in a specific direction; therefore, it can identify whether a screen exists in the background information. A specular mask is a binary image that marks areas of specular reflection in the original image, typically caused by screen reflections, glass overlays, or strong light sources. Specular masks can be used as attention guides or additional input channels to improve a model's sensitivity to reflections.
[0059] The second face image, edge map, spectrogram, and spectrophotometer are input into the feature extraction layer for feature extraction and concatenation to obtain the first feature vector. This first feature vector is then input into the attention layer for attention calculation to obtain the second feature vector, which guides the model to focus on the face edges and background regions. The third feature vector is input into the multi-scale feature fusion layer for feature fusion to obtain the fourth feature vector. During training, the multi-scale feature fusion layer learns from a large amount of data which background feature combinations predict risky behavior to obtain the fused features, i.e., the fourth feature vector. This fourth feature vector is then input into the classification layer for category prediction to determine whether the second face image contains risky behavior.
[0060] In an optional embodiment, when the first terminal device includes a depth sensor, such as a time-of-flight (ToF) sensor and a structured light sensor, the depth sensor can also be used to emit infrared light and measure the time or shape of the returned infrared light to construct a depth map of the scene, which corresponds to the second face image. A real face is three-dimensional, while the screen of the second terminal device is two-dimensional. Therefore, by analyzing the flatness of the depth map and whether the background information includes information from the second terminal device, it is possible to effectively distinguish whether the currently acquired face image originates from a real person or is being played in high definition on the screen of the second terminal device. For example, if the flatness of the face region in the depth map is lower than a set first flatness threshold (e.g., 5% or 10%), or if the background information includes information from the second terminal device, then the target face is determined to be a fake face. If the flatness of the face region in the depth map is higher than a set second flatness threshold (e.g., 50%, 80%), and the background information includes information from the second terminal device, then the target face is determined to be a real face. The above method implements a combination of software and hardware.
[0061] The following provides a facial image processing system, which includes a first terminal device and a cloud server. The first terminal device has a target application installed, which includes a facial recognition module for executing... Figure 1 The method is illustrated. For example, the face recognition module is implemented as a Software Development Kit (SDK) corresponding to the target application. The user performs face recognition using a terminal device. The cloud server checks whether the background information of the second face image includes information from the second terminal device.
[0062] 1) Startup and initialization: The user triggers the facial recognition function in the target application (App), such as... Figure 2As shown. The target application calls the face recognition module, and the face recognition module completes its initialization.
[0063] 2) Camera and Dual-Stream Setup: The face recognition module requests and obtains camera access, then creates and configures two video streams: a preview stream and a capture stream, such as... Figure 2 As shown.
[0064] The preview stream describes the video stream from which the user previews a face image. The face recognition module calls the camera's Application Programming Interface (API). It sets a first digital zoom level (e.g., 1.5x or 2.0x), crops and enlarges the first face image captured by the camera to obtain a partial face image, and displays this partial face image in the preview area of the face capture page, so that the user sees only a portion of the face in the preview area.
[0065] The capture stream describes a second face image captured by the camera, provided that the face displayed in the preview area meets the set first face criteria. The second face image is not displayed in the preview area. For example, based on a digital zoom ratio of 1.0x, i.e., without any digital zoom, the second face image is captured while maintaining the camera's native wide-angle field of view.
[0066] 3) Guiding user behavior When a partial face is displayed in the preview area, the face recognition module calls the user interface (UI) module to display guiding information on the face capture page, such as... Figure 2 As shown. For example, prompts such as "Please place your face completely within the preview frame" or "Please move your phone further away." Guided by both visual (partial face) and informational prompts, users actively increase the distance between themselves and their primary device until their face is fully displayed in the preview area.
[0067] 4) Distance judgment and timing of shooting The face recognition module and the face pose and distance monitoring module continuously analyze the preview stream. This module determines whether the faces displayed in the preview area meet the set first face criteria, such as... Figure 2 As shown. For example, by calculating the pixel area occupied by the face in the image or the distance between the eyes, the physical distance between the face and the camera is assessed in real time. When the physical distance exceeds a set threshold, the face image displayed in the preview area is considered to meet the set first face condition. Another example is when the face pose and distance monitoring module detects that the face size decreases from a first size to a second size, and the first size is greater than a size threshold while the second size is less than a set size threshold, it determines that the user has increased the distance between their face and the camera as guided.
[0068] If the face displayed in the preview area meets the set first face condition, the face recognition module triggers the single-frame wide-angle image acquisition module to acquire the second face image.
[0069] 5) Covert wide-angle acquisition: The acquisition module obtains the image data of the current frame, i.e., the second face image, from the background acquisition stream, such as... Figure 2 As shown.
[0070] Because the capture stream is always kept at a wide angle (e.g., 1.0x zoom), and the physical distance between the first terminal device and the face is relatively far, the second face image not only contains a clear face but also has a high probability of including a wider background environment around the face. If other users are using the second terminal device to play a high-definition video of the user to be verified, then the screen bezel, notch, earpiece, and even the fingers of the other user holding the second terminal device may also be captured and included in the second face image.
[0071] 6) End-to-end cloud joint analysis: The captured second facial image, after being encrypted, is uploaded to a cloud server, such as... Figure 2 As shown.
[0072] The cloud server calls the facial image processing model to extract background information around the face from the second facial image. If the background information includes information from the second terminal device, it determines that there is a risky behavior in the facial recognition process, such as... Figure 2 As shown. The face image processing model is mainly used to analyze whether there are rectangular or straight edges, bangs, earpieces, screen-specific reflections, moiré patterns, and user fingers in the background information.
[0073] 7) Result Return and Processing The cloud server returns either a "first notification message indicating the presence of risky behavior" or a "second notification message indicating the absence of risky behavior" to the first terminal device, such as... Figure 2 As shown, the face recognition module on the first terminal device executes subsequent service logic based on the received notification information. For example, if the face recognition module receives the first notification information, it outputs an authentication failure message; if it receives the second notification information, it outputs an authentication success message.
[0074] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 104 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
[0075] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0076] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving data sent by B, or it can be understood as A indirectly receiving data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending data directly to A, or it can be understood as B indirectly sending data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.
[0077] Figure 3 A schematic diagram of a face image processing apparatus provided as an exemplary embodiment of this specification is shown below. Figure 3 As shown, the device includes: a display module 31, a data acquisition module 32, a magnification module 33, an output module 34, and a determination module 35.
[0078] The display module is used to respond to face recognition trigger operations and display the face capture page, which includes a preview area; The acquisition module is used to acquire a first face image using the camera on the first terminal device, wherein the first face image includes the target face; The zoom module is used to zoom in on the first face image to display a partial face image in the preview area; the output module is used to output guidance information to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image. The acquisition module is used to acquire a second face image using a camera if the face image displayed in the preview area meets the set first face condition during the transition process. The second face image includes the target face and the background information around the target face. The determination module is used to determine that the target face is a fake face if the background information in the second face image includes information of the second terminal device.
[0079] In an optional embodiment, the magnification module is specifically used to: obtain a first digital zoom ratio, the first digital zoom ratio representing the cropping size; crop the first face image based on the cropping size represented by the first digital zoom ratio to obtain a local face region; interpolate the local face region to obtain a local face image with the same size as the first face image, and display the local face image in the preview area.
[0080] Optionally, before interpolating the local face region, the magnification module is further configured to: if the local face region is found not to meet the set second face condition, increase the first digital zoom ratio, and re-crop the first face image according to the cropping size represented by the increased first digital zoom ratio, until the cropped local face region meets the set second face condition.
[0081] In one optional embodiment, the acquisition module is specifically used to acquire a first face image based on a second digital zoom ratio corresponding to the camera; the acquisition module is specifically used to: acquire a second face image based on a second digital zoom ratio corresponding to the camera while keeping the optical focal length of the camera unchanged; wherein, the second digital zoom ratio is less than the first digital zoom ratio.
[0082] Optionally, the determining module is also used for: If a complete human face is detected in the preview area, it is determined that the human face image displayed in the preview area meets the set first human face condition. If the face feature displayed in the detected preview area decreases from the first size to the second size, and the second size is less than the set size threshold, then it is determined that the face image displayed in the preview area meets the set first face condition, and the first size is greater than the size threshold. If the face feature displayed in the detected preview area decreases from the first size to the second size, and the difference between the second size and the first size is greater than a set difference threshold, then the face image displayed in the preview area is determined to meet the set first face condition.
[0083] In an optional embodiment, the determining module is specifically used for: sending the second face image to the server device; wherein the server device calls a face image processing model to extract background information around the face from the second face image; if the background information includes information of the second terminal device, then generates notification information indicating that the target face is a fake face and sends it to the first terminal device; and receiving the notification information sent by the server device.
[0084] Optionally, the information of the second terminal device may include any one of the following: border information, environment information, and texture information.
[0085] For detailed descriptions of the implementation methods and effects of the above-mentioned device, please refer to the foregoing embodiments, which will not be repeated here.
[0086] Figure 4 This specification illustrates a schematic diagram of an electronic device provided in an exemplary embodiment, which is applicable to the face image processing method provided in the foregoing embodiments. Figure 4 As shown, the electronic device 700 mainly consists of a communication interface 702, a user interface 704, a processor 706, and a memory 708. These components are interconnected and communicate with each other through a system bus, network, or other connection mechanism 410. The communication interface 702 enables the device 700 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 702 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 702 can also be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi (Wireless Fidelity), Bluetooth, Global Positioning System (GPS), or wide-area wireless interface such as WiMAX (Wireless Maximum) or LTE (Long Term Evolution). Of course, the communication interface 702 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 702 may also include multiple physical communication interfaces, such as a Wi-Fi interface, a Bluetooth interface, and a wide-area wireless interface.
[0087] User interface 704 includes receiving user input and providing output to the user. Therefore, user interface 704 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), display using DLP (Digital Light Processing) technology, printer, and other known or future similar devices. User interface 704 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other known or future similar devices. In some embodiments, user interface 704 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, electronic device 700 may support remote access from other devices via communication interface 702 or another physical interface (not shown). User interface 704 can be configured to receive user input, the position and movement of which can be indicated by an indicator or cursor described herein. User interface 704 can also be configured as a display device for rendering or displaying text fragments.
[0088] Processor 706 may include one or more general-purpose processors and / or special-purpose processors. Memory 708 may include one or more volatile and / or non-volatile memory components and may be integrated wholly or partially with processor 706. Memory 708 may include removable and non-removable components.
[0089] The processor 706 is capable of executing program instructions 718 (e.g., compiled or uncompiled program logic and / or machine code) stored in memory 708 to perform the various functions described herein.
[0090] Memory 708 may contain non-transitory computer-readable media, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Memory 708 stores program instructions that, when executed by device 700, enable device 700 to perform any of the methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Processor 706 executing program instructions 718 may cause processor 706 to use data 712.
[0091] For example, program instructions 718 may include an operating system 722 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 700 and one or more applications 720 (e.g., a browser, social application, or game application). Similarly, data 712 may include operating system data 716 and application data 714. Operating system data 716 is primarily accessible to the operating system 722, while application data 714 is primarily accessible to one or more applications 720. Application data 714 may reside in a file system visible or hidden from the user of device 700.
[0092] Application 720 can communicate with operating system 722 through one or more application programming interfaces (APIs). These APIs help application 720 read and / or write application data 714, transmit or receive information via communication interface 702, receive or display information on user interface 704, etc.
[0093] In some terminology, application 720 may be simply referred to as "app". Furthermore, application 720 can be downloaded to device 700 through one or more online app stores or app markets. However, applications can also be installed on device 700 in other ways, such as through a web browser or a physical interface on electronic device 700 (e.g., a USB port).
[0094] Accordingly, embodiments of this specification also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, embodiments of this specification also provide a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above-described method embodiments. It should be understood that each step or combination of steps in the above-described method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.
[0095] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes that element.
[0096] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.
[0097] The terminology used in the embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one. “A plurality” generally includes at least two, but does not exclude the inclusion of at least one.
[0098] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0099] The above are merely embodiments of this specification and are not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A face image processing method, characterized in that, Applied to a first terminal device, the method includes: In response to a face recognition trigger, a face capture page is displayed, which includes a preview area; A first face image is captured using the camera on the first terminal device, and the first face image includes the target face; The first face image is magnified to display a partial face image in the preview area, and guidance information is output to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image. During the transition, if the face image displayed in the preview area meets the set first face condition, the camera is used to capture a second face image, which includes the target face and the background information around the target face. If the background information in the second face image includes information about the second terminal device, the target face is determined to be a fake face.
2. The method according to claim 1, characterized in that, Enlarging the first face image to display a partial face image in the preview area includes: Obtain the first digital zoom ratio, where the first digital zoom ratio represents the cropping size; Based on the cropping size represented by the first digital zoom ratio, the first face image is cropped to obtain a local face region; Interpolate the local face region to obtain a local face image with the same size as the first face image, and display the local face image in the preview area.
3. The method according to claim 2, characterized in that, Before interpolating the local face region, the method further includes: If the local face region is found to not meet the set second face condition, the first digital zoom ratio is increased, and the first face image is re-cropped according to the cropping size represented by the increased first digital zoom ratio until the cropped local face region meets the set second face condition.
4. The method according to claim 2 or 3, characterized in that, Acquiring a first face image using the camera on the first terminal device includes: acquiring the first face image based on the second digital zoom level corresponding to the camera; Acquiring a second face image using the camera includes: acquiring a second face image based on a second digital zoom ratio corresponding to the camera, while keeping the optical focal length of the camera constant; Wherein, the second digital zoom ratio is less than the first digital zoom ratio.
5. The method according to claim 4, characterized in that, It also includes at least one of the following operations: If a complete human face is detected in the preview area, it is determined that the human face image displayed in the preview area meets the set first human face condition. For the detected facial features displayed in the preview area, if the facial features decrease from a first size to a second size, and the second size is less than a set size threshold, then it is determined that the facial image displayed in the preview area meets the set first facial condition, where the first size is greater than the size threshold. For the detected facial features displayed in the preview area, if the facial features decrease from a first size to a second size, and the difference between the second size and the first size is greater than a set difference threshold, then it is determined that the facial image displayed in the preview area meets the set first facial condition.
6. The method according to any one of claims 1-3 and 5, characterized in that, If information about a second terminal device appears in the background information of the second face image, determining that the target face is a fake face includes: The second face image is sent to the server device; wherein, the server device calls the face image processing model to extract the background information around the face from the second face image, and if the background information includes the information of the second terminal device, a notification message indicating that the target face is a fake face is generated and sent to the first terminal device; Receive notification information sent by the server device.
7. The method according to claim 6, characterized in that, The information of the second terminal device includes any one of the following: border information, environment information, and texture information.
8. A face image processing device, characterized in that, include: Display module, acquisition module, zoom-in module, output module, and confirmation module; The display module is used to respond to the face recognition trigger operation and display the face collection page, which includes a preview area; The acquisition module is used to acquire a first face image using a camera on a first terminal device, wherein the first face image includes a target face; The magnification module is used to magnify the first face image to display a partial face image in the preview area; the output module is used to output guidance information to guide the user to increase the distance between the face and the camera, so that the preview area transitions from displaying a partial face image to displaying a complete face image. The acquisition module is used to acquire a second face image using the camera if the face image displayed in the preview area meets the set first face condition during the transition process. The second face image includes the target face and the background information around the target face. The determining module is used to determine that the target face is a fake face if the background information in the second face image includes information of the second terminal device.
9. An electronic device, characterized in that, include: A memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions for: performing the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it is able to perform the steps of the method described in any one of claims 1-7.
11. A computer program product, characterized in that, include: A computer program / instruction that, when executed by a processor, enables the implementation of the steps in the method described in any one of claims 1-7.