Face recognition method, apparatus, device, medium and program product
By collecting multi-dimensional spatial data for spatial information preprocessing and risk detection in the face recognition system, and quantifying and assigning values to determine strategies, the problem of Hook technology injecting fictitious photos or videos is solved, thereby improving the security and authenticity of face recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2025-08-15
- Publication Date
- 2026-06-02
AI Technical Summary
Existing facial recognition systems cannot effectively identify attacks that inject fake photos or videos through hooking technology, resulting in insufficient security and authenticity.
Multidimensional spatial data, including image data and device spatial data, is collected by multidimensional sensing components. Spatial information preprocessing and risk detection are performed, and spatial scores are obtained by quantification. Based on the scores, facial recognition strategies are determined, and strategies such as multispectral verification, dynamic behavior analysis, and multi-factor composite authentication are used for facial recognition.
Effectively defends against attacks that inject fictitious data, improves the security and authenticity of facial recognition, and enhances the system's security protection capabilities.
Smart Images

Figure CN122135443A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of biometrics, artificial intelligence and information security, and more specifically to a face recognition method, device, equipment, medium and program product. Background Technology
[0002] In the fintech field, facial recognition serves as a core line of defense for identity verification. By accurately comparing biometric features, it ensures that the operation matches the account holder, curbing account theft at its source. In high-risk scenarios such as fund transfers, its verification function can intercept unauthorized operations, enhancing security. Furthermore, facial recognition aligns with regulatory requirements for identity authenticity, providing dual protection for the security and compliance of financial transactions.
[0003] Current facial recognition systems have security risks. Attackers can use hooking technology to inject fake photos or videos into cameras, bypassing existing security measures. Current technology cannot effectively identify such fake attacks and lacks sufficient defense against these advanced attacks, making it difficult to ensure the authenticity and security of facial recognition. Summary of the Invention
[0004] In view of the above problems, this application provides face recognition methods, devices, equipment, media and program products that improve security protection capabilities.
[0005] According to a first aspect of this application, a face recognition method is provided, comprising: in response to a user-authorized face recognition request, acquiring multidimensional spatial data through a multidimensional sensing component; wherein the multidimensional spatial data includes image data and device spatial data; performing spatial information preprocessing on the image data to obtain background environment information; performing spatial risk detection on the background environment information and the device spatial data, and quantifying and assigning a value to the spatial risk detection result to obtain a spatial score; and determining a face recognition strategy based on the spatial score, and performing face recognition on the user through the face recognition strategy.
[0006] According to an embodiment of this application, the image data includes a first image and a second image, and the device spatial data includes first distance information, second distance information, current terminal posture, first location information, and second location information; the acquisition of multi-dimensional spatial data through multi-dimensional sensing components includes: acquiring the first image based on a front-facing camera component; acquiring the second image based on a rear-facing camera component; performing optical dimension ranging through the front-facing camera component to obtain the first distance information; performing acoustic ranging through an acoustic-to-electrical conversion component to obtain the second distance information; acquiring the current terminal posture based on an inertial measurement sensor; acquiring the first location information based on a positioning component; and acquiring the second location information based on user identification card location information.
[0007] According to an embodiment of this application, the spatial risk detection of the background environment information and the device spatial data includes: optical environment detection based on the background environment information; face distance detection based on the first distance information and the second distance information; geographical location detection based on the first location information and the second location information; and posture detection based on the current terminal posture.
[0008] According to an embodiment of this application, the face distance detection based on the first distance information and the second distance information includes: determining a distance difference based on the first distance information and the second distance information; and determining that the face distance detection result is abnormal when the distance difference is greater than a preset distance threshold.
[0009] According to an embodiment of this application, the background environment information includes the first environment information and the second environment information; the optical environment detection based on the background environment information includes: determining environmental difference information based on the first environment information and the second environment information; and determining the optical environment detection result as an environmental anomaly when the environmental difference information meets the environmental anomaly conditions.
[0010] According to an embodiment of this application, determining a face recognition strategy based on the spatial score includes: determining the face recognition strategy as a static recognition strategy when the spatial score is greater than a preset score threshold; and determining the face recognition strategy as a high-intensity recognition strategy when the spatial score is less than or equal to the preset score threshold, wherein the high-intensity recognition strategy includes one or more of a multispectral verification strategy, a dynamic behavior analysis strategy, and a multi-factor composite authentication strategy.
[0011] According to an embodiment of this application, the step of preprocessing the image data with spatial information to obtain background environment information includes: processing the first image through a deep learning model to obtain the first environment information; wherein the deep learning model is trained based on a set of indoor and outdoor scene images; and processing the second image through the deep learning model to obtain the second environment information.
[0012] According to an embodiment of this application, the step of quantifying and assigning values to the spatial risk detection results to obtain a spatial score includes: determining abnormal risk situations based on the spatial risk detection results; and calculating the spatial score based on the abnormal risk situations and a preset initial spatial score.
[0013] A second aspect of this application provides a face recognition device, the device comprising: a data acquisition module, configured to acquire multidimensional spatial data via a multidimensional sensing component in response to a user-authorized face recognition request; wherein the multidimensional spatial data includes image data and device spatial data; a preprocessing module, configured to perform spatial information preprocessing on the image data to obtain background environment information; a risk detection module, configured to perform spatial risk detection on the background environment information and the device spatial data, and quantify and assign a spatial risk detection result to obtain a spatial score; and a face recognition module, configured to determine a face recognition strategy based on the spatial score, and perform face recognition on the user using the face recognition strategy.
[0014] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0015] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0016] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0017] In the embodiments of this application, based on the collected multidimensional spatial data, spatial scores are obtained through spatial information preprocessing, spatial risk detection, and quantification. This allows for the determination of the authenticity and credibility of the current user's facial image and spatial environment, detecting the degree of spatial risk, and identifying the risk of injecting fictitious data. Selecting a facial recognition strategy based on the spatial scores and utilizing this strategy for facial recognition can effectively defend against fictitious data injection attacks, enhance security capabilities, and improve the authenticity and security of facial recognition. Attached Figure Description
[0018] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1 The illustrations depict application scenarios of face recognition methods, apparatuses, devices, media, and program products according to embodiments of this application.
[0020] Figure 2 A flowchart illustrating a face recognition method according to an embodiment of this application is shown schematically.
[0021] Figure 3 This illustration schematically shows a data acquisition flowchart of a face recognition method according to an embodiment of the present application;
[0022] Figure 4 A schematic diagram illustrating the spatial risk detection flowchart of a face recognition method according to an embodiment of this application is shown.
[0023] Figure 5 A schematic diagram of a face recognition system according to an embodiment of this application is shown.
[0024] Figure 6 A schematic diagram illustrating the structure of a face recognition device according to an embodiment of this application is shown; and
[0025] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a face recognition method according to an embodiment of this application. Detailed Implementation
[0026] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0030] In recent years, facial recognition has become an important front line in the fight against illicit industries. Currently, the mainstream attack methods of illicit industries against facial recognition mainly involve two steps. The first step is that malicious attackers use technologies such as deepfake and dynamic animation to create fake videos, generating fake facial verification videos from static photos of victims obtained through illegal channels. The second step is to hijack the system's camera input source based on Hook technology (using tools such as virtual cameras) and inject the fake video stream into the hijacked camera channel.
[0031] Based on the attack method of injecting fictitious images / videos, the security of facial recognition is strengthened mainly by enhancing the detection of the device environment (identifying hook behavior) and strengthening the identification of fictitious videos from the backend algorithm level of facial recognition.
[0032] Regarding device environment detection, malicious attackers on client systems commonly use root access tools combined with kernel root schemes to effectively conceal their hooking behavior by leveraging high system privileges. Normal applications (APPs), however, are constrained by the operating system and run within its user privileges. This gives attackers a natural advantage in concealment and resistance to detection, making it difficult for current APP detection technologies to effectively identify the characteristics of such tools. Furthermore, after reverse engineering, these detection methods are easily bypassed by malicious attackers using higher privileges. At the facial recognition backend algorithm level, analysis is primarily conducted in the spatial and frequency domains, such as capturing subtle artifacts in local facial regions and identifying the frequency domain features of fictitious images. However, with the development of facial recognition fictitious technology, malicious attackers utilize deep fictitious techniques based on feature fusion to bypass relevant detections. This has led to a continuous increase in the strength of liveness detection based on action and lighting, significantly impacting the user experience.
[0033] Therefore, in the field of facial recognition, how to effectively resist attacks that inject fictitious photos or videos into the system's camera through hook technology to circumvent facial verification, thereby enhancing the security of facial recognition and strengthening the system's security protection capabilities, remains an urgent problem to be solved.
[0034] This application provides a face recognition method, comprising: responding to a user's authorized face recognition request, acquiring multi-dimensional spatial data through a multi-dimensional sensing component; wherein the multi-dimensional spatial data includes image data and device spatial data; performing spatial information preprocessing on the image data to obtain background environment information; performing spatial risk detection on the background environment information and device spatial data, and quantifying and assigning a spatial risk detection result to obtain a spatial score; and determining a face recognition strategy based on the spatial score, and performing face recognition on the user through the face recognition strategy. In this application's embodiments, based on the acquired multi-dimensional spatial data, by obtaining a spatial score through spatial information preprocessing, spatial risk detection, and quantification, the authenticity and credibility of the current user's face image and spatial environment can be determined, and the degree of spatial risk can be detected to identify the risk of injecting fictitious data. Selecting a face recognition strategy based on the spatial score and using the face recognition strategy for face recognition can effectively defend against attacks injecting fictitious data, improve security protection capabilities, and enhance the authenticity and security of face recognition.
[0035] Figure 1 The illustrations depict application scenarios of face recognition methods, apparatuses, devices, media, and program products according to embodiments of this application.
[0036] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0037] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0038] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0039] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0040] It should be noted that the face recognition method provided in this application embodiment can be executed by the first terminal device 101, the second terminal device 102, the third terminal device 103, or the server 105. Correspondingly, the face recognition device provided in this application embodiment can generally be located in the first terminal device 101, the second terminal device 102, the third terminal device 103, or the server 105. The face recognition method provided in this application embodiment can also be executed by a server or server cluster that is different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the face recognition device provided in this application embodiment can also be located in a server or server cluster that is different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0042] The following will be based on Figure 1 The described scene, through Figures 2-4 The face recognition method according to the embodiments of this application will be described in detail.
[0043] Figure 2 A flowchart illustrating a face recognition method according to an embodiment of this application is shown.
[0044] like Figure 2 As shown, the face recognition method of this embodiment includes operations S210 to S240. This face recognition method does not limit the specific executing entity. The executing entity can be any electronic device, such as a terminal device (e.g., a mobile terminal, desktop device, IoT terminal, and industrial control terminal) or a server device, etc. The executing entity can also be any software application or client. For ease of explanation, a terminal device will be used as an example below.
[0045] During operation S210, in response to a user-authorized face recognition request, multi-dimensional spatial data is collected through multi-dimensional sensing components; the multi-dimensional spatial data includes image data and device spatial data.
[0046] In the embodiments of this application, user consent or authorization is required before initiating the face recognition process. For example, a face recognition request can be sent to the user, and operation S210 is executed if the user authorizes or agrees that the terminal device can obtain the user's face and spatial information.
[0047] In facial recognition scenarios (such as account security and identity verification, transaction and payment security, remote business processing, etc.), a facial recognition request is initiated to the user. After the user authorizes the facial recognition request, the system responds to the user's authorization of the facial recognition request and performs the corresponding facial recognition operation based on the authorized facial recognition request.
[0048] A face recognition request is an interactive instruction initiated by a terminal device to a user, triggering the face recognition process. Its core purpose is to obtain user authorization and initiate processes such as face collection, spatial risk monitoring, and face recognition.
[0049] A facial recognition request should include at least the following: information about the initiator (such as an account management app, identity verification service, password retrieval service, etc.), a description of the scenario (such as confirming account ownership, confirming identity, etc.), permissions granted by the user (such as allowing access to the camera, allowing analysis of facial feature data, allowing temporary use of location information, etc.), restrictions on the use of data (such as specifying the specific scope of use of facial data), a data protection commitment (such as encrypted transmission of images, local storage without uploading, etc.), rejection information (such as cancellation, temporary non-verification, etc.), and a user agreement (such as the content, purpose, types of information collected, methods of information use, user rights, security guarantees, etc. of the facial recognition service).
[0050] For example, a facial recognition request is initiated to the user, and relevant content corresponding to the request is displayed. This provides the user with an entry point to choose whether to agree to or refuse facial recognition. That is, before facial recognition is performed on the user's information, the user can provide their consent or refusal instruction through the corresponding entry point. If the user agrees to facial recognition, the user's information enters the facial recognition process, and operation S210 is executed. It should be noted that, apart from the images captured by the camera, other information does not need to be uploaded to the server for facial recognition verification. Spatial risk detection can be completed locally on the terminal device, without involving data transmission and without infringing on user privacy. It can also be stated in the user agreement that image collection is necessary for the facial recognition service.
[0051] In the technical solution of this application, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0052] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this application all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0053] During facial recognition, multidimensional spatial data of the current user is collected through multidimensional sensor components integrated in the terminal device. These multidimensional sensor components include, but are not limited to, front-facing camera components (such as a front-facing camera), rear-facing camera components (such as a rear-facing camera), audio-to-electrical conversion components (such as an earpiece and microphone), inertial measurement sensors (such as a gyroscope), positioning components (such as a Global Positioning System (GPS)), a Subscriber Identity Module Card (SIM card), and an Ultra Wide Band (UWB) chip. The multidimensional spatial data includes image data and device spatial data. The image data includes a first image and a second image, and the device spatial data includes first distance information, second distance information, current terminal posture, first location information, and second location information.
[0054] For example, multidimensional sensors can be integrated into an app in the form of a software development kit (SDK). The SDK contains data acquisition logic, which includes calling various sensors in the operating system to collect multidimensional spatial data.
[0055] During operation S220, spatial information preprocessing is performed on the image data to obtain background environment information.
[0056] According to an embodiment of this application, in operation S220, the image data is preprocessed with spatial information to obtain background environment information, including: processing a first image through a deep learning model to obtain first environment information; wherein the deep learning model is trained based on a set of indoor and outdoor scene images; and processing a second image through a deep learning model to obtain second environment information.
[0057] Based on the first image captured by the front-facing camera and the second image captured by the rear-facing camera, a deep learning model can obtain background environment information labels corresponding to the current front and rear cameras, namely, first environment information and second environment information. Background environment information includes indoor, outdoor, black screen, and mirrored images, etc. The deep learning model can determine whether the environment of an image is indoor or outdoor. This deep learning model falls under the category of indoor and outdoor scene recognition in scene classification tasks. It can use a convolutional neural network model, employing a set of indoor and outdoor scene images (including indoor images, outdoor images, and black screen images, etc.) for feature extraction and modeling, and then train the convolutional neural network model to recognize the environmental information corresponding to the image.
[0058] In the embodiments of this application, a deep learning model is used to improve the accuracy of spatial environment recognition, which facilitates the effective identification of indoor attacks and improves the reliability of face recognition.
[0059] During operation of S230, spatial risk detection is performed on background environmental information and equipment space data, and the spatial risk detection results are quantified and assigned values to obtain spatial score values.
[0060] After performing spatial risk detection based on background environment information and device spatial data, the spatial risk detection results need to be further quantified and assigned values to obtain spatial scores. Spatial risk detection includes, but is not limited to, optical environment detection, face distance detection, geographic location detection, posture detection, and paired device location detection.
[0061] According to an embodiment of this application, in operation S230, quantifying and assigning values to the spatial risk detection results to obtain a spatial score, the steps include: determining abnormal risk conditions based on the spatial risk detection results; and calculating the spatial score based on the abnormal risk conditions and a preset initial spatial score.
[0062] First, initial spatial scores are preset for each type of spatial risk detection. For example, optical environment detection, face distance detection, geographic location detection, posture detection, and paired device location detection are assigned 5, 4, 3, 2, and 1 points respectively. Second, when any spatial risk detection category is detected as abnormal, the score for that category is automatically reset to zero. For instance, if the optical environment detection result is an environmental anomaly, the corresponding score of 5 is reset to zero. Finally, the remaining valid scores are accumulated to calculate the spatial score. For example, if the optical environment detection result is an environmental anomaly, the face distance detection result is a distance anomaly, and the geographic location detection, posture detection, and paired device location detection results are all normal, the spatial score is 6 points.
[0063] In the embodiments of this application, spatial risk is transformed into a quantifiable score by combining a preset initial spatial risk score with abnormal risk conditions, which accurately reflects the current spatial security status and facilitates the selection of facial recognition strategies.
[0064] In operation S240, a face recognition strategy is determined based on spatial score values, and the user's face is recognized using the face recognition strategy.
[0065] In the embodiments of this application, based on the collected multidimensional spatial data, spatial scores are obtained through spatial information preprocessing, spatial risk detection, and quantification. This allows for the determination of the authenticity and credibility of the current user's facial image and spatial environment, detecting the degree of spatial risk, and identifying the risk of injecting fictitious data. Selecting a facial recognition strategy based on the spatial scores and utilizing this strategy for facial recognition can effectively defend against fictitious data injection attacks, enhance security capabilities, and improve the authenticity and security of facial recognition.
[0066] According to an embodiment of this application, in operation S240, which determines a face recognition strategy based on spatial score values, the method includes: determining a static recognition strategy when the spatial score value is greater than a preset score threshold; and determining a high-intensity recognition strategy when the spatial score value is less than or equal to the preset score threshold. The high-intensity recognition strategy includes one or more of a multispectral verification strategy, a dynamic behavior analysis strategy, and a multi-factor composite authentication strategy.
[0067] The facial recognition strategy is determined by comparing the spatial score with a preset threshold. If the spatial score is less than or equal to the threshold (e.g., 10 points), a significant anomaly is considered to exist in the spatial data of the current terminal device. In this case, the facial recognition strategy needs to be strengthened, particularly the liveness detection, to reduce the risk of attackers impersonating users. Therefore, a high-intensity recognition strategy, such as multispectral verification, dynamic behavior analysis, and multi-factor authentication, is employed. If the spatial score is greater than the threshold (e.g., 10 points), the spatial data of the current terminal device is considered normal, and a static recognition strategy can be used for facial recognition.
[0068] Multispectral verification strategy: Add light-based liveness detection for facial recognition (e.g., by flashing different colored RGB lights on the screen, and judging whether the user is real based on the color and texture of the light reflected from the face), and strengthen the authentication strength of the backend algorithm. For faces with indistinct light emission features, the transaction will be directly rejected.
[0069] Dynamic behavior analysis strategy: Increase the number of live actions during the face recognition process, and upload the entire live action process video to the server for further risk identification, such as determining whether the video has been edited or has AI editing features.
[0070] Multi-factor authentication strategy: Add other traditional facial recognition authentication methods to the corresponding business scenarios, such as adding SMS verification codes and remote agent verification.
[0071] Static recognition strategy: Detect and locate faces, extract face image feature vectors, compare the similarity of the vectors with feature vectors in the database, determine whether authentication is successful based on the similarity threshold, and finally output the authentication conclusion.
[0072] In the embodiments of this application, the face recognition strategy is dynamically adjusted according to the current spatial risks, taking into account both efficiency and security, effectively defending against fictitious attacks, and improving the reliability and adaptability of face recognition.
[0073] Figure 3 A schematic flowchart illustrating the data acquisition process of a face recognition method according to an embodiment of this application is shown.
[0074] According to embodiments of this application, the image data includes a first image and a second image, and the device spatial data includes first distance information, second distance information, current terminal posture, first location information, and second location information. For example... Figure 3 As shown, the process of acquiring multidimensional spatial data through the multidimensional sensing component in operation S210 includes operations S310 to S370.
[0075] When operating the S310, the first image is acquired based on the front-facing camera component.
[0076] The current facial recognition photo, i.e., the first image, can be obtained through the front-facing camera.
[0077] When operating the S320, a second image is acquired based on the rear camera component.
[0078] The second image can be obtained by using the rear camera to capture the current camera's photo.
[0079] When operating the S330, optical distance measurement is performed through the front-facing camera to obtain the first distance information.
[0080] Facial feature points can be located using face detection algorithms. By combining the focal length of the front-facing camera, the size of the optical sensor, and the pixel displacement of the feature points, the distance between the current face and the front-facing camera can be calculated using the principle of triangulation, which is the first distance information.
[0081] When operating the S340, acoustic ranging is performed through the acoustic-to-electric conversion component to obtain the second distance information.
[0082] Ultrasonic signals can be emitted through the earpiece built into the terminal device, and the reflected ultrasonic waves can be received by the microphone. By calculating the time difference or phase difference between the emission and reception, the distance information between the obstacle in front of the terminal device and the terminal device can be obtained, i.e., the second distance information.
[0083] When operating the S350, the current terminal attitude is obtained based on the inertial measurement sensor.
[0084] The current terminal device attitude can be obtained through the gyroscope, i.e., the current terminal attitude.
[0085] When operating the S360, the first location information is obtained based on the positioning component.
[0086] The current geographical location information of the terminal device, i.e., the primary location information, can be obtained through GPS.
[0087] When operating S370, second location information is obtained based on the user identification card location information.
[0088] The approximate location information of the current device can be obtained through the location information service of the operator's SIM card, i.e., the secondary location information.
[0089] As an optional embodiment, the process of collecting multi-dimensional spatial data through multi-dimensional sensing components in operation S210 also includes obtaining location distance information based on the pairing status of wearable devices using ultra-wideband (UWB) technology. For example, it checks whether the terminal device has been paired with a smart wearable device (such as a smartwatch, bracelet, etc.) and obtains the current location distance between the user's wearable device and the mobile terminal device, i.e., location distance information, using UWB technology.
[0090] In the embodiments of this application, multi-dimensional spatial data is obtained, which can integrate information such as images, distance, terminal posture, and location to make up for the limitations of single data recognition, improve the accuracy and robustness of face recognition, and enhance the face recognition capability in complex scenarios.
[0091] Figure 4 A schematic flowchart illustrating the spatial risk detection process of a face recognition method according to an embodiment of this application is shown.
[0092] like Figure 4 As shown, the spatial risk detection of background environment information and equipment space data in operation S230 includes operations S410 to S440.
[0093] When operating the S410, optical environment detection is performed based on background environment information.
[0094] Based on background environment information, optical environment detection is used to check whether there are any obvious anomalies in the current environment of the terminal device.
[0095] According to an embodiment of this application, the background environment information includes first environment information and second environment information; in operation S410, optical environment detection is performed based on the background environment information, including: determining environmental difference information based on the first environment information and the second environment information; and determining the optical environment detection result as an environmental anomaly when the environmental difference information meets the environmental anomaly conditions.
[0096] If the environmental difference information meets the conditions for an abnormal environment, the optical environment detection result is determined to be normal. If the environmental difference information does not meet the conditions for an abnormal environment, the optical environment detection result is determined to be abnormal.
[0097] An abnormal environment condition is defined as the presence of any of the abnormal scenarios, in which case the optical environment detection result is considered to be an abnormal environment. Abnormal scenarios include: (1) the first environmental information is indoor or outdoor, while the second environmental information is a black screen, indicating that the front camera displays normal face recognition, while the rear camera displays a black screen; (2) the first environmental information is indoor or outdoor, while the second environmental information is an indoor mirror or an outdoor mirror, indicating that the front camera displays a normal face recognition environment, while the rear camera displays a mirror image of the front camera; (3) the first environmental information is outdoor, while the second environmental information is indoor, indicating that the front camera displays an outdoor environment, while the rear camera displays an indoor environment. Since attackers generally carry out attacks indoors, detecting abnormal scenarios in indoor and outdoor environments can effectively detect fictitious attacks.
[0098] In the embodiments of this application, by comparing the image environment information of the front and rear camera components, the environment in which the user's face is located is identified, and environmental anomalies are effectively detected by utilizing abnormal environmental conditions, thereby identifying fictitious attacks.
[0099] When operating S420, face distance detection is performed based on the first distance information and the second distance information.
[0100] Based on the first distance information and the second distance information, the face distance detection is used to check whether there are any obvious abnormalities in the current face distance of the terminal device.
[0101] According to an embodiment of this application, in the operation S420 of performing face distance detection based on first distance information and second distance information, the method includes: determining a distance difference based on the first distance information and second distance information; and determining that the face distance detection result is abnormal if the distance difference is greater than a preset distance threshold.
[0102] Attackers use a terminal device connected to other devices via a data cable to directly inject photos / videos into the camera stream of the terminal device to achieve facial recognition attacks. In this type of attack scenario, there are unlikely to be any obstacles in front of the terminal device. In this case, there will be a significant difference between the first and second distance information. If the difference between the first and second distance information is greater than a preset distance threshold (e.g., 20cm), a facial distance risk is identified, and the facial distance detection result is considered abnormal. Conversely, if the distance difference is less than or equal to the preset distance threshold, the facial distance detection result is considered normal.
[0103] In the embodiments of this application, based on attack scenarios such as replacing the system camera through hook technology or using customized root methods, the attacker does not need to perform face recognition actions by holding the terminal. This results in a large difference between the distance between the face and the front camera (first distance information) and the distance between the terminal and the obstacle in front (second distance information). By detecting the distance difference, abnormal scenarios can be effectively detected, and attacks such as hook technology injection and customized root can be effectively resisted.
[0104] In operation S430, geographic location detection is performed based on the first location information and the second location information.
[0105] The system can compare the first and second location information to determine if there are any abnormal locations. For example, if the GPS-obtained current device location information is compared with the operator's SIM card location information, and a discrepancy is found between the location information and the card location information, or if the terminal device is found to have no SIM card inserted when reading the card location information, then a location risk is considered present, and the location detection result is "location abnormal." If no location abnormality is found, the location detection result is "location normal."
[0106] When operating S440, attitude detection is performed based on the current terminal attitude.
[0107] The system can check the current terminal posture obtained through the gyroscope to determine if it conforms to common postures used in face recognition. For example, it can check for subtle positional movements expected when the device is held by a real person during face recognition. If the current terminal posture is a horizontal position (within 180 degrees ± 10 degrees) and there is no shaking during face recognition, then a risk is considered, and the posture detection result is "posture abnormal." If the current terminal posture conforms to common postures used in face recognition, the posture detection result is also "posture abnormal."
[0108] As an optional embodiment, the spatial risk detection in operation S230, which checks the background environment information and device spatial data, also includes pairing device location detection based on location distance information. For example, it checks whether the location distance between the current user's wearable device and the mobile terminal device is within the normal usage range. If the user has a paired wearable device that is in normal working condition, but the location of the wearable device is in a different location (e.g., different prefecture-level cities) from the current face recognition terminal device, then a risk is identified, and the pairing device location detection result is "pairing device abnormal." If the wearable device is within the normal usage range, the pairing device location detection result is "pairing device normal."
[0109] In the embodiments of this application, the security and accuracy of face recognition are improved by preventing non-liveness attacks such as photo and video injection through spatial risk detection, and ensuring that face recognition is based on real faces for authentication.
[0110] As one embodiment, based on the above-described face recognition method, this application also provides a face recognition system. The following will combine... Figure 5 The system is described in detail.
[0111] Figure 5 A diagram of a face recognition system according to an embodiment of this application is shown schematically.
[0112] like Figure 5 As shown, the face recognition system includes a device risk assessment system and a back-end risk control system. The device risk assessment system includes a multi-dimensional sensor module and a strategy execution module, while the back-end risk control system includes a device spatial risk information preprocessing module and a device spatial risk scoring module.
[0113] The multi-dimensional sensor module is integrated into the APP in the form of an SDK. The SDK calls multiple sensors in the terminal operating system to collect multi-dimensional spatial data.
[0114] The device space risk information preprocessing module preprocesses the image data sent by the multi-dimensional sensor module and obtains the background environment information labels of the current front camera and rear camera through deep learning algorithms, including indoor, outdoor, black screen, mirror, etc.
[0115] The device spatial risk scoring module compares the risk information transmitted from the front-end module and scores the results according to the following strategy: It assigns scores to five categories of spatial risk detection results: optical environment, face distance, geographical location, posture, and paired device location. The initial spatial scores are set to 5, 4, 3, 2, and 1 points respectively. When any spatial risk detection result is abnormal, the score for that category is reset to zero. The valid scores are then tallied to obtain the spatial score. If the spatial score is less than or equal to 10 points, the device's spatial information is considered to have a significant anomaly. After obtaining the spatial score, the module returns different instructions to the strategy execution module, such as successful risk authentication, failed risk authentication, or the need for additional enhanced authentication measures.
[0116] The strategy execution module determines the face recognition strategy based on the spatial score and executes the face recognition strategy. The strategy execution module can also perform front-end result processing based on instructions from the device spatial risk scoring module, such as displaying error messages and invoking enhanced authentication logic.
[0117] Based on the above-described face recognition method, this application also provides a face recognition device. The following will combine... Figure 6 The device is described in detail.
[0118] Figure 6 A schematic block diagram of a face recognition device according to an embodiment of this application is shown.
[0119] like Figure 6 As shown, the face recognition device 600 in this embodiment includes a data acquisition module 610, a preprocessing module 620, a risk detection module 630, and a face recognition module 640.
[0120] The data acquisition module 610 is used to collect multi-dimensional spatial data through multi-dimensional sensing components in response to a user-authorized face recognition request; wherein, the multi-dimensional spatial data includes image data and device spatial data. In one embodiment, the data acquisition module 610 can be used to perform the operation S210 described above, which will not be repeated here.
[0121] The preprocessing module 620 is used to perform spatial information preprocessing on the image data to obtain background environment information. In one embodiment, the preprocessing module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0122] The risk detection module 630 is used to perform spatial risk detection on background environment information and equipment spatial data, and to quantify and assign values to the spatial risk detection results to obtain spatial scores. In one embodiment, the risk detection module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0123] The face recognition module 640 is used to determine a face recognition strategy based on spatial score values, and to perform face recognition on the user using the face recognition strategy. In one embodiment, the face recognition module 640 can be used to perform the operation S240 described above, which will not be repeated here.
[0124] According to an embodiment of this application, the image data includes a first image and a second image, and the device spatial data includes first distance information, second distance information, current terminal posture, first location information, and second location information; the data acquisition module 610 includes: a first data acquisition unit for acquiring the first image based on a front-facing camera component; a second data acquisition unit for acquiring the second image based on a rear-facing camera component; a third data acquisition unit for obtaining the first distance information by performing optical dimension ranging through the front-facing camera component; a fourth data acquisition unit for obtaining the second distance information by performing acoustic ranging through an acoustic-to-electrical conversion component; a fifth data acquisition unit for acquiring the current terminal posture based on an inertial measurement sensor; a sixth data acquisition unit for acquiring the first location information based on a positioning component; and a seventh data acquisition unit for acquiring the second location information based on user identification card location information.
[0125] According to an embodiment of this application, the risk detection module 630 includes: a first risk detection unit for optical environment detection based on background environment information; a second risk detection unit for face distance detection based on first distance information and second distance information; a third risk detection unit for geographical location detection based on first location information and second location information; and a fourth risk detection unit for posture detection based on the current terminal posture.
[0126] According to an embodiment of this application, the second risk detection unit includes: a distance difference calculation subunit, used to determine a distance difference based on first distance information and second distance information; and a distance anomaly determination subunit, used to determine that the face distance detection result is a distance anomaly when the distance difference is greater than a preset distance threshold.
[0127] According to an embodiment of this application, the background environment information includes first environment information and second environment information; the first risk detection unit includes: an environmental difference determination subunit, used to determine environmental difference information based on the first environment information and the second environment information; and an environmental anomaly determination subunit, used to determine that the optical environment detection result is an environmental anomaly when the environmental difference information meets the environmental anomaly conditions.
[0128] According to an embodiment of this application, a face recognition module 640 includes: a first face recognition unit, configured to determine that the face recognition strategy is a static recognition strategy when the spatial score is greater than a preset score threshold; and a second face recognition unit, configured to determine that the face recognition strategy is a high-intensity recognition strategy when the spatial score is less than or equal to the preset score threshold, wherein the high-intensity recognition strategy includes one or more of a multispectral verification strategy, a dynamic behavior analysis strategy, and a multi-factor composite authentication strategy.
[0129] According to an embodiment of this application, the preprocessing module 620 includes: a first image processing unit, configured to process a first image using a deep learning model to obtain first environmental information; wherein the deep learning model is trained based on an indoor and outdoor scene image set; and a second image processing unit, configured to process a second image using a deep learning model to obtain second environmental information.
[0130] According to an embodiment of this application, the risk detection module 630 includes: an anomaly determination unit, used to determine risk anomalies based on spatial risk detection results; and a score calculation unit, used to calculate a spatial score based on the risk anomalies and a preset initial spatial score.
[0131] According to embodiments of this application, any multiple modules among the data acquisition module 610, preprocessing module 620, risk detection module 630, and face recognition module 640 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the data acquisition module 610, preprocessing module 620, risk detection module 630, and face recognition module 640 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the data acquisition module 610, preprocessing module 620, risk detection module 630, and face recognition module 640 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0132] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a face recognition method according to an embodiment of this application.
[0133] like Figure 7 As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0134] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 902 and / or RAM 903. It should be noted that programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.
[0135] According to embodiments of this application, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.
[0136] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0137] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.
[0138] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the face recognition method provided in the embodiments of this application.
[0139] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0140] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0141] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0142] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0144] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
Claims
1. A face recognition method, characterized in that, The method includes: In response to a user-authorized facial recognition request, multi-dimensional spatial data is collected through a multi-dimensional sensing component; wherein, the multi-dimensional spatial data includes image data and device spatial data; The image data is preprocessed with spatial information to obtain background environment information; Spatial risk detection is performed on the background environment information and the device spatial data, and the spatial risk detection results are quantified and assigned values to obtain spatial scores; and A face recognition strategy is determined based on the spatial score, and the user's face is recognized using the face recognition strategy.
2. The method according to claim 1, characterized in that, in, The image data includes a first image and a second image, and the device spatial data includes first distance information, second distance information, current terminal posture, first location information, and second location information; The acquisition of multidimensional spatial data through multidimensional sensing components includes: The first image is obtained based on the front-facing camera component; The second image is obtained based on the rear camera component; The first distance information is obtained by optical dimension ranging through the front-facing camera component; The second distance information is obtained by acoustic ranging using an acoustic-to-electrical conversion component; The current terminal attitude is obtained based on an inertial measurement sensor; The first location information is obtained based on the positioning component; and The second location information is obtained based on the user identification card location information.
3. The method according to claim 2, characterized in that, The spatial risk detection of the background environment information and the device spatial data includes: Optical environment detection is performed based on the aforementioned background environment information; Face distance detection is performed based on the first distance information and the second distance information; Geographic location detection is performed based on the first location information and the second location information; and Attitude detection is performed based on the current terminal attitude.
4. The method according to claim 3, characterized in that, The face distance detection based on the first distance information and the second distance information includes: Based on the first distance information and the second distance information, determine the distance difference; and If the distance difference is greater than a preset distance threshold, the face distance detection result is determined to be an anomaly.
5. The method according to claim 3, characterized in that, in, The background environment information includes the first environment information and the second environment information; the optical environment detection based on the background environment information includes: Based on the first environmental information and the second environmental information, environmental difference information is determined; If the environmental difference information meets the conditions for environmental anomaly, the optical environment detection result is determined to be an environmental anomaly.
6. The method according to claim 1, characterized in that, The method for determining the face recognition strategy based on the spatial score includes: If the spatial score is greater than a preset score threshold, the face recognition strategy is determined to be a static recognition strategy; and If the spatial score is less than or equal to the preset score threshold, the face recognition strategy is determined to be a high-intensity recognition strategy. The high-intensity recognition strategy includes one or more of the following: multispectral verification strategy, dynamic behavior analysis strategy, and multi-factor composite authentication strategy.
7. The method according to claim 5, characterized in that, The step of preprocessing the image data spatial information to obtain background environment information includes: The first image is processed by a deep learning model to obtain the first environmental information; wherein the deep learning model is trained based on a set of indoor and outdoor scene images; and The second image is processed by the deep learning model to obtain the second environmental information.
8. The method according to claim 1, characterized in that, The process of quantifying and assigning values to the spatial risk detection results to obtain spatial scores includes: Based on the space risk detection results, abnormal risk situations are identified; and The spatial score is calculated based on the aforementioned risk anomaly and the preset initial spatial score.
9. A face recognition device, characterized in that, The device includes: The data acquisition module is used to collect multi-dimensional spatial data through multi-dimensional sensing components in response to a user-authorized face recognition request; wherein, the multi-dimensional spatial data includes image data and device spatial data; The preprocessing module is used to perform spatial information preprocessing on the image data to obtain background environment information; The risk detection module is used to perform spatial risk detection on the background environment information and the equipment spatial data, and to quantify and assign values to the spatial risk detection results to obtain a spatial score; and The face recognition module is used to determine a face recognition strategy based on the spatial score value, and to perform face recognition on the user through the face recognition strategy.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 8.