Method and apparatus for intercepting face recognition attack, and device and storage medium
By acquiring transaction data and live-scanned facial images, and utilizing unsupervised clustering and pixel consistency comparison, combined with geographic attribute data, the system identifies and blocks attack images from risky users, solving the problem of low accuracy in facial recognition attack interception and achieving efficient attack identification and interception.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CHINA UNIONPAY
- Filing Date
- 2025-09-23
- Publication Date
- 2026-07-30
AI Technical Summary
In existing technologies, the accuracy of intercepting facial recognition attacks is relatively low, especially in scenarios with high timeliness requirements. It is difficult to effectively intercept potential or missed attacks, which affects the experience of normal users.
By acquiring transaction data and live-scanned facial images, image clusters are formed based on the transaction data. Unsupervised clustering algorithms and pixel consistency comparisons are used to identify and intercept attack images from risky users. Geographic attribute data is then combined to conduct risk assessment, thereby improving the accuracy of attack identification.
While ensuring the pass rate for normal users, it can quickly and accurately identify and block facial recognition attacks from risky users, improving the accuracy of blocking facial recognition attacks and avoiding potential or missed attacks.
Smart Images

Figure CN2025123173_30072026_PF_FP_ABST
Abstract
Description
Methods, devices, equipment, and storage media for intercepting facial recognition attacks
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese patent application CN202510114354.7, filed on January 23, 2025, entitled “Method, Apparatus, Device and Storage Medium for Intercepting Face Recognition Attacks”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of information security technology, and in particular relates to a method, apparatus, device and storage medium for intercepting facial recognition attacks. Background Technology
[0004] With the widespread application of identity verification across various fields, facial recognition has become a primary means of verification. During the collection of facial data, malicious actors may employ techniques such as synthetic attacks and injection attacks to compromise facial recognition results. For example, using artificial intelligence synthesis technology, a user's facial features can be fitted onto a realistic background image and animated to create fake biopsy videos, thus attacking the facial recognition process and disrupting the results.
[0005] In related technologies, for facial recognition scenarios with high timeliness requirements, such as order transaction scenarios and account login scenarios, in order not to affect the experience of normal users, a high pass rate is usually prioritized. However, this can easily reduce the accuracy of intercepting facial recognition attacks. Summary of the Invention
[0006] This application provides a method, apparatus, device, and storage medium for intercepting face recognition attacks, which can solve the problem of low accuracy in intercepting face recognition attacks in related technologies.
[0007] In a first aspect, embodiments of this application provide a method for intercepting facial recognition attacks, which may include:
[0008] Obtain transaction data and live face images of N users within the first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1.
[0009] Based on transaction data, the live face images of N users are divided to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background regions, and the background regions are non-face regions in the live face images.
[0010] Identify the interception data for risky users corresponding to biopsied face images in the image cluster. The interception data is used to indicate the interception of face recognition attacks related to risky users within the second time window.
[0011] Secondly, embodiments of this application provide an interception device for face recognition attacks, which may include:
[0012] The acquisition module is used to acquire transaction data and live face images of N users within the first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1.
[0013] The segmentation module is used to segment the live face images of N users based on transaction data to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background regions, and the background regions are non-face regions in the live face images.
[0014] The determination module is used to determine the interception data of risky users corresponding to the live face images in the image cluster. The interception data is used to indicate the interception of face recognition attacks related to risky users within the second time window.
[0015] Thirdly, embodiments of this application provide a computer device, which includes: a processor and a memory storing computer program instructions;
[0016] When the processor executes computer program instructions, it implements the method for intercepting face recognition attacks as described in the first aspect.
[0017] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the method for intercepting face recognition attacks as described in the first aspect.
[0018] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method for intercepting face recognition attacks as shown in the first aspect.
[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method for intercepting face recognition attacks as described in the first aspect.
[0020] The method, apparatus, device, and storage medium for intercepting facial recognition attacks according to embodiments of this application can acquire transaction data and live face images of N users within a first time window. The transaction data includes at least one of the following: transaction time and transaction location. Based on the transaction data, the live face images of the N users are divided to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background areas, where the background area is a non-face area in the live face image. Then, the interception data of the risky user corresponding to the live face image in the image cluster is determined. The interception data is used to indicate the interception of facial recognition attacks related to the risky user within a second time window. In this way, potential attack images can be quickly and accurately discovered and identified from a large amount of daily transaction data, targeting both time and / or space dimensions. From the perspective of post-transaction handling, this approach fully utilizes the homogeneous background characteristics of illegal facial recognition attacks, combined with the clustering of attack image features that naturally distinguish them from the discrete features of normal user live-check facial images. By clustering live-check facial images, the live-check facial images of high-risk users and their corresponding interception data can be identified. This achieves the interception of facial recognition attacks related to high-risk users, flexibly controls the risks of facial recognition services, and, while ensuring the success rate of identity verification for normal users, intercepts facial recognition attacks from high-risk users as much as possible after a transaction occurs, avoiding potential or missed facial recognition attacks and improving the accuracy of facial recognition attack interception. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 is a flowchart of a method for intercepting face recognition attacks provided in an embodiment of this application;
[0023] Figure 2 is a schematic diagram of the reference geographic attribute data in a method for intercepting face recognition attacks provided in an embodiment of this application;
[0024] Figure 3 is a flowchart illustrating a method for intercepting face recognition attacks provided in an embodiment of this application;
[0025] Figure 4 is a schematic diagram of the structure of an interception device for face recognition attacks provided in an embodiment of this application;
[0026] Figure 5 is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0027] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0028] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0029] The acquisition, storage, use, and processing of data (including but not limited to features and information mentioned in this document) in the technical solution of this application all comply with the relevant provisions of national laws and regulations.
[0030] Liveness detection is a method used in identity verification scenarios to determine the true physiological characteristics of an object. In facial recognition applications, liveness detection can verify whether a user is a real, living person by using technologies such as facial landmark localization and facial tracking through combined actions such as blinking, opening the mouth, shaking the head, and nodding. It can effectively resist common attack methods such as photos, videos, face swapping, masks, occlusion, and screen re-photographing, and determine the authenticity of the face, thereby helping users identify fraudulent behavior and ensuring user information security.
[0031] Facial recognition has been a primary means of identity verification for many years. However, with the advancement of artificial intelligence and the use of internet technology for cyberattacks, information theft, extortion, and fraud, attacks on facial recognition have evolved from low-tech methods like re-photographing and using masks to more complex and realistic methods such as synthesis and injection. During facial detection, malicious actors employ techniques like synthesis and injection attacks to compromise the recognition results. For example, in the liveness detection phase, attackers might fit a captured image including a face onto a realistic background image provided by the attacker, animate corresponding actions, and generate a fake liveness detection video to undermine the facial recognition process.
[0032] Because these background images share a common background characteristic—meaning multiple attacked user videos are fitted using the same background image—the fitted images exhibit consistent background features. This is particularly evident in pre-synthesized and real-time synthesized injection attacks. Therefore, to effectively identify attacks using background images with common characteristics, potential common background images can typically be detected during the transaction process. However, for time-sensitive facial recognition scenarios such as order transactions and account logins, a high pass rate is usually prioritized to avoid impacting the experience of legitimate users. This can easily reduce the accuracy of intercepting facial recognition attacks.
[0033] To address the problems in related technologies, this application proposes a method to improve the systematic capabilities of face recognition prevention and control. This method involves post-transaction analysis of face recognition attacks and interception of face recognition attacks from high-risk users, thereby avoiding potential or missed face recognition attacks and improving the accuracy of face recognition attack interception.
[0034] Based on this, the methods, apparatus, computer devices, and storage media of the embodiments of this application will be described in detail below with reference to Figures 1 to 5. It should be noted that these embodiments are not intended to limit the scope of disclosure of this application.
[0035] Figure 1 is a flowchart of a method for intercepting face recognition attacks provided in an embodiment of this application.
[0036] As shown in Figure 1, this method for intercepting face recognition attacks can be applied to computer devices, and the interception method can specifically include the following steps:
[0037] Step 110: Obtain transaction data and live face images of N users within the first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1. Step 120: Based on the transaction data, divide the live face images of the N users to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background areas, where the background area is a non-face area in the live face image. Step 130: Determine the interception data of the risky user corresponding to the live face image in the image cluster. The interception data is used to indicate the interception of face recognition attacks related to the risky user within the second time window.
[0038] In this way, potential attack images can be quickly and accurately discovered and identified from a large amount of daily transaction data, targeting both time and / or space dimensions. From the perspective of post-transaction handling, this approach fully utilizes the homogeneous background characteristics of illegal facial recognition attacks, combined with the clustering of attack image features that naturally distinguish them from the discrete features of normal user live-check facial images. By clustering live-check facial images, the live-check facial images of high-risk users and their corresponding interception data can be identified. This achieves the interception of facial recognition attacks related to high-risk users, flexibly controls the risks of facial recognition services, and, while ensuring the success rate of identity verification for normal users, intercepts facial recognition attacks from high-risk users as much as possible after a transaction occurs, avoiding potential or missed facial recognition attacks and improving the accuracy of facial recognition attack interception.
[0039] The steps described above are explained in detail below.
[0040] First, regarding step 110, in some embodiments of this application, during the actual service delivery process, the server can obtain service transaction data sent by the user. This service transaction data includes transaction data from N users and live-check facial images used for identity verification during the transaction. The transaction data includes at least one of the following: transaction time and transaction location. The service transaction data can be service transaction data from time-sensitive facial recognition scenarios such as money transfers, binding payment cards, account logins, and facial recognition payment orders.
[0041] Next, regarding step 120, in some embodiments of this application, step 120 may specifically include steps 1201 and 1202.
[0042] Step 1201: Based on the transaction data, extract features from the background regions of the live face images of N users to obtain the background feature data of each user's live face image.
[0043] For example, if the transaction data is the transaction location, features can be extracted from the background regions of the live face images within the same transaction location; if the transaction data is the transaction time, features can be extracted from the background regions of the live face images within the same transaction time interval; if the transaction data includes both the transaction location and the transaction time, features can be extracted from the background regions of the live face images within the same transaction time interval and the same transaction location. For example, if the transaction time interval is 5 minutes, the transaction location can be represented by latitude and longitude information, that is, the geographical location of the transaction location is a rectangular area with a side length of 100m. The live face images of users within the same time interval and transaction location are then clustered.
[0044] In some embodiments, feature extraction can be performed using a pre-trained convolutional network with a classic structure. In other embodiments, histogram of oriented gradients (HOG) features can be used instead to save computational costs.
[0045] It should be noted that, in addition to performing the background separation process described above after acquiring transaction data and live face images, the background separation process can also be introduced during the hack detection stage before performing specific facial recognition or identity verification services. The background image in the live face image can be separated and retained simultaneously with the hack detection process. In this way, the background feature data of the user's live face image can be obtained at the same time as acquiring transaction data and live face images, so as to improve clustering efficiency and thus improve the efficiency of intercepting facial recognition attacks.
[0046] Furthermore, in the facial recognition aspect of this application, "hack" refers to the act of bypassing the security verification of the facial recognition system through technical means to illegally access or impersonate others. This behavior may involve using photos, videos, or other forms of disguise to deceive the facial recognition system and obtain unauthorized access. The process of detecting such hacking behavior is called hack detection.
[0047] Step 1202: Based on the background feature data of each user's biopsy face image, cluster the biopsy face images of N users to obtain at least one image cluster.
[0048] For example, considering the homogeneous background of illegal face recognition attacks, and the clustering of attack image features that are naturally different from the discrete features of normal user biopsy face images, unsupervised clustering methods such as DBSCAN, OPTICS, and BIRCH can be used to cluster the background feature data.
[0049] Thus, the image clusters that can be significantly aggregated have the characteristic of smaller average centroid distance within the class, and the corresponding biopsy face images have a higher risk level. This can improve the recognition effect of biopsy face images as much as possible, avoid potential or missed face recognition attacks, and improve the accuracy of intercepting face recognition attacks.
[0050] Furthermore, in some embodiments, step 1202 described above may specifically include:
[0051] Using an unsupervised clustering algorithm, based on the similarity between the background feature data of at least two users' live face images, live face images of N users are clustered to obtain at least one initial image cluster;
[0052] At least one image cluster is determined from at least one initial image cluster; wherein the image cluster satisfies at least one of the following conditions: the number of biopsy face images in the image cluster is greater than or equal to the number of reference images, and the range of the average centroid distance interval of the image cluster belongs to the range of the reference average centroid distance interval.
[0053] Here, in the embodiments of this application, the similarity of the background regions of at least two frames of biopsy face images in the initial image cluster is greater than or equal to a preset similarity.
[0054] For example, BIRCH clustering is performed on all background feature data within the group, the number of reference images within the image cluster and the reference average centroid distance range are set, the initial image clusters falling within the range are labeled, and the background region of the biopsy face image within the image cluster is obtained.
[0055] Pixel consistency is compared in the background regions of the biopsied face images within the image cluster. The normalized sum of squared differences algorithm is used to evaluate the pixel consistency between at least two background regions. Specifically, it can be calculated using the following formula (1):
[0056] Where I(x, y) represents the x and y coordinates of pixel i in one background region, and I′(x, y) represents the x and y coordinates of pixel i in another background region.
[0057] In this way, the sum of squared differences between corresponding pixels in at least two background regions is calculated. For images using the same background template, the difference should be close to 0. A relevant threshold is set, and images below the threshold are considered to have pixel-level consistency. The matched background images are added to the attack background library, and the associated users are marked with risks. It is worth noting that such associations are often many-to-many clustered. Thus, error fluctuations can be further eliminated by the number of clusters.
[0058] Then, regarding step 130, the time in the second time window in this embodiment of the application is later than the time in the first time window.
[0059] In this way, potential attack images can be quickly and accurately discovered and identified from a large amount of daily transaction data, targeting the time and / or spatial dimensions. From the perspective of post-transaction handling, the homogeneous background of illegal face recognition attacks can be fully utilized. Combined with the clustering of attack image features that are naturally different from the discrete features of normal user live face images, live face images can be clustered to determine the live face images of users with high risk and their corresponding interception data.
[0060] In some embodiments of this application, the transaction data includes the transaction time. Based on this, before step 130, the method for intercepting face recognition attacks may also include steps 2101 and 2102.
[0061] Step 2101: Perform pixel consistency comparison on the background regions of the biopsied face images of at least two users in the image cluster to obtain pixel comparison data of at least two background regions.
[0062] Step 2102: If the data value of the pixel comparison data is greater than or equal to the first preset threshold, determine at least two users corresponding to the pixel comparison data as risk users.
[0063] For example, in order to further reduce errors, it is necessary to perform pixel consistency judgment on the background area of the live face images within these image clusters, and issue a unified risk warning to the associated users who exceed the consistency threshold, so as to make corresponding strategy adjustments to the live detection process of at least two users within the image cluster.
[0064] Furthermore, the pixel consistency judgment mentioned above can be determined by using a normalized difference square algorithm and a normalized correlation coefficient evaluation algorithm to determine the consistency between the background regions of at least two users' live face images. Based on this, in some embodiments, step 2101 may specifically include:
[0065] The comparison algorithm performs pixel consistency comparison on the background regions of at least two users' live face images in the image cluster to obtain pixel comparison data of at least two background regions; wherein the comparison algorithm includes at least one of the following: normalized difference square algorithm, normalized correlation coefficient evaluation algorithm.
[0066] In some embodiments of this application, the transaction data includes the transaction location. Based on this, before step 130, the method for intercepting face recognition attacks may also include steps 2103 and 2104.
[0067] Step 2103: Based on the reference geographic attribute data of the transaction location, perform correlation matching on the background areas of the live face images of at least two users in the image cluster to obtain content matching data.
[0068] It should be noted that the reference geographic attribute data in the embodiments of this application includes at least one of the following: weather condition attribute data, indoor and outdoor environment attribute data, background complexity attribute data, scene function attribute data, and administrative division attribute data.
[0069] Step 2104: If the data value of the content matching data is less than or equal to the second preset threshold, determine at least two users corresponding to the content matching data as risk users.
[0070] For example, a spatiotemporal domain, i.e., a geographic location element, can be introduced. This involves verifying the background regions of transaction data and live face images of N users within the same spatiotemporal range to clarify the risk scope. Since the attribute data reflected in the background regions of normal users within the same geographic location tends to be consistent, while the background regions of risky users using the same set of attack images often fail to adjust to changes in geographic location, leading to conflicts. Therefore, correlation matching can be performed on the background regions of live face images of at least two users in an image cluster using reference geographic attribute data, as shown in Figure 2. If the data value of the content matching data is less than or equal to a second preset threshold, then at least two users corresponding to the content matching data are identified as risky users.
[0071] In some embodiments of this application, the transaction data includes the transaction location and transaction time. Based on this, before step 130, the method for intercepting face recognition attacks may also include steps 2105 to 2107.
[0072] Step 2105: Perform pixel consistency comparison on the background regions of the live face images of at least two users in the image cluster to obtain pixel comparison data of at least two background regions; and, based on the reference geographic attribute data of the transaction location, perform correlation matching on the background regions of the live face images of at least two users in the image cluster to obtain content matching data.
[0073] Step 2106: Determine the risk score data for each of at least two users based on the pixel comparison data, the first risk confidence level corresponding to the pixel comparison data, the content matching data, and the second risk confidence level corresponding to the content matching data.
[0074] Step 2107: If the risk score data is greater than or equal to the third preset threshold, determine the user corresponding to the risk score data as a risk user.
[0075] For example, based on a fixed time zone, the element of geographical location is introduced to extract background features from business transaction images within the same time interval and spatial range. Data with conflicting features are analyzed to clarify the scope of risk. Since the background attribute characteristics of normal users at a fixed geographical location within the same time period tend to be consistent, while the backgrounds of risky users using the same attack template often cannot adjust to changes in geographical location, resulting in conflicts. For two sets of conflicting data, risk confirmation can be achieved through methods such as adjacent location reference, comparison of historical transaction data, and manual screening.
[0076] It should be noted that, in order to improve the accuracy of identifying risky users, the sharing users identified by any of the above methods can be re-examined through the following steps. For example, for two sets of data that conflict with each other, risk re-examination can be carried out based on the reference geographical attribute data of adjacent transaction locations, live facial images of adjacent transaction times, or manual screening. Based on this, before step 130, the method for intercepting facial recognition attacks may also include steps 2108 and 2109.
[0077] Step 2108: Based on the risk verification data, re-verify the background areas of the live face images of at least two users in the image cluster to obtain re-verification result data; wherein, the risk verification data includes at least one of the following: reference geographical attribute data of adjacent transaction locations of the transaction location, and live face images of adjacent transaction times of the transaction time.
[0078] Step 2109: If the data value of the re-verification result data is higher than or equal to the fourth preset threshold, determine the user corresponding to the content matching data as a risk user.
[0079] For example, to improve the accuracy of identifying high-risk users, reference geographic attribute data of adjacent transaction locations and live-shot facial images of adjacent transaction times can be used to further clarify the true attributes. Then, users with abnormal values within the problem group are marked as risky, and manual verification can be performed if costs are controllable. In this embodiment, for broad-ranging attribute features such as weather and time, adjacent locations are more suitable for judgment, while for narrow-ranging attribute features, similar times are more suitable.
[0080] In addition, after step 130, a step of using intercepted data is also included. In this embodiment, the intercepted data can be used to intercept at least one of the following face recognition attacks: face recognition attacks related to the risky user, and face recognition attacks related to the risky user's electronic device. Based on this, the method for intercepting face recognition attacks in this embodiment may further include:
[0081] Step 140: Based on the intercepted data, send an interception command to the detection terminal corresponding to the risky user;
[0082] The detection end includes at least one of the following: the electronic device of the risky user, and the face detection system used to examine the live-shot face image of the risky user;
[0083] The interception instruction includes at least one of the following: a first interception instruction and a second interception instruction; the first interception instruction is used to instruct the electronic device of the risky user to increase the difficulty level of liveness detection for the risky user; the second interception instruction is used to instruct the face detection system to increase the hack detection interception level of the liveness detection face image of the risky user.
[0084] For example, risk labeling requires labeling both users and devices. This two-way association between devices and users can uncover more potential or missed attack behaviors. The risk-labeled users or devices can be adjusted according to specific business scenarios, and these risk strategies can respectively affect the difficulty of liveness detection during the front-end data collection process and the response level of the back-end hack detection service.
[0085] Therefore, this application embodiment can utilize two detection methods—pixel-level consistency analysis and background attribute conflict—to quickly and accurately identify potential attack targets within a large volume of daily transaction data, taking into account the temporal and spatial characteristics of illegal attacks on facial recognition. From a post-incident perspective, it marks target risky users and devices. Unlike traditional methods, it fully leverages the attack characteristics of facial recognition scenarios, focuses on localized abnormal data, reduces attack screening costs, and, in conjunction with multi-stage risk control parameter distribution, flexibly controls the risks of facial recognition services, implementing interception as efficiently as possible while ensuring the pass rate for legitimate users.
[0086] It should be noted that, in this embodiment of the application, the live face image is obtained through a liveness detection process. Specifically, during the liveness detection stage, the electronic device prompts the user to adjust the position of the electronic device so that the face is moved within the coverage area of the face capture preview control displayed on the electronic device. The user holding the electronic device will cause the electronic device to undergo obvious deflection or acceleration changes, so that the electronic device can capture the change information of the electronic device. When the user adjusts the position of the electronic device so that it can capture the user's complete face image, at least one frame of the live face image in the face capture preview control can be captured.
[0087] It should be noted that the method for intercepting facial recognition attacks provided in this application embodiment can be applied to scenarios that require live facial image verification for identity authentication, such as financial payment, access control, account login authentication, and other scenarios that require verification of the user's real identity.
[0088] For ease of understanding, the following describes the process of intercepting face recognition attacks, including the method applied to computer devices in the embodiments of this application.
[0089] Figure 3 is a flowchart illustrating a method for intercepting face recognition attacks provided in an embodiment of this application.
[0090] As shown in Figure 3, flow images can be obtained from all business scenarios, and background separation can be performed on the flow images. The background area is divided according to the transaction time and transaction location, so as to evaluate the risk labeling of users and devices associated with the images, which can guide the configuration of flexible face risk strategies.
[0091] Based on this, the interception method may include steps f1 to f10.
[0092] In step f1, transaction data and live-detection facial images of N users are acquired. In actual production operations, transaction data and live-detection facial images submitted by the user terminals can be acquired. For details, please refer to the relevant description of step 110 in the above embodiments, which will not be repeated here.
[0093] In step f2, background feature data is extracted. Specifically, background features are extracted from the background regions of the biopsy face images of N users to obtain the background feature data of each user's biopsy face image.
[0094] Here, transaction data may include transaction time and transaction location. Therefore, it can be processed separately, namely, pixel consistency analysis is performed in steps f3 to f4, and multi-factor evaluation of reference geographic attribute data is performed in steps f5 to f7, as detailed below.
[0095] Step f3: Based on the background feature data of each user's live face image, cluster the live face images of N users according to the transaction time to obtain at least one image cluster.
[0096] Step f4: Perform pixel consistency comparison on the background regions of the biopsied face images of at least two users in the image cluster to obtain pixel comparison data of at least two background regions.
[0097] Step f5: Based on the background feature data of each user's live face image, cluster the live face images of N users according to the transaction location to obtain at least one image cluster.
[0098] Step f6: For the transaction location, based on the reference geographic attribute data of the transaction location, correlation matching can be performed on the background areas of the live face images of at least two users in the image cluster to obtain content matching data.
[0099] Step f7: Determine the first risk score data based on the pixel comparison data and the first risk confidence level corresponding to the pixel comparison data; and determine the second risk score data based on the content matching data and the second risk confidence level corresponding to the content matching data.
[0100] Step f8: Based on the first risk score data and the second risk score data, summarize the risk score data for each user.
[0101] In this way, the qualitative extremes before risk labeling of users can be verified by cross-evaluation of the two methods. Users who hit both questions have a higher risk confidence level, which can appropriately reduce the subsequent manual verification process or further increase the risk level rating, providing more reference information for risk decision-making.
[0102] Step f9: If the risk score data is greater than or equal to the third preset threshold, determine the user corresponding to the risk score data as a risk user.
[0103] Step f10: Based on the intercepted data, send an interception command to the detection terminal corresponding to the risky user. For details, please refer to the relevant description of step 130 in the above embodiment, which will not be repeated here.
[0104] Therefore, by targeting the time and / or spatial dimensions, potential attack images can be quickly and accurately discovered and identified from a large amount of daily transaction data. From the perspective of post-transaction handling, this method fully utilizes the homogeneous background characteristics of illegal facial recognition attacks, combined with the clustering of attack image features that naturally distinguish them from the discrete features of normal user live face images. By clustering live face images, the live face images of high-risk users and their corresponding interception data can be identified. This achieves the interception of facial recognition attacks related to high-risk users, flexibly controls the risks of facial recognition services, and, while ensuring the identity verification success rate of normal users, intercepts facial recognition attacks from high-risk users as much as possible after the transaction occurs, avoiding potential or missed facial recognition attacks and improving the accuracy of facial recognition attack interception.
[0105] This application also provides a device for intercepting facial recognition attacks, which will be described in detail with reference to Figure 4.
[0106] Figure 4 is a schematic diagram of the structure of an interception device for face recognition attacks provided in an embodiment of this application.
[0107] In some embodiments of this application, the face recognition attack interception device shown in FIG4 can be installed in the computer equipment provided in the embodiments of this application.
[0108] As shown in Figure 4, the interception device 40 for face recognition attacks may specifically include:
[0109] The acquisition module 401 is used to acquire the transaction data and live face images of N users within the first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1.
[0110] The segmentation module 402 is used to segment the live face images of N users based on transaction data to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background regions, and the background regions are non-face regions in the live face images.
[0111] The determination module 403 is used to determine the interception data of the risk user corresponding to the live face image in the image cluster. The interception data is used to indicate the interception of face recognition attacks related to the risk user within the second time window.
[0112] Thus, the face recognition attack interception device in this embodiment can acquire transaction data and live face images of N users within a first time window. The transaction data includes at least one of the following: transaction time and transaction location. Based on the transaction data, the live face images of the N users are divided to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background areas, where the background area is a non-face area in the live face image. Then, the interception data of the risky user corresponding to the live face image in the image cluster is determined. The interception data is used to indicate the interception of face recognition attacks related to the risky user within a second time window. In this way, potential attack images can be quickly and accurately discovered and identified from a large amount of daily transaction data, targeting both time and / or space dimensions. From the perspective of post-transaction handling, this approach fully utilizes the homogeneous background characteristics of illegal facial recognition attacks, combined with the clustering of attack image features that naturally distinguish them from the discrete features of normal user live-check facial images. By clustering live-check facial images, the live-check facial images of high-risk users and their corresponding interception data can be identified. This achieves the interception of facial recognition attacks related to high-risk users, flexibly controls the risks of facial recognition services, and, while ensuring the success rate of identity verification for normal users, intercepts facial recognition attacks from high-risk users as much as possible after a transaction occurs, avoiding potential or missed facial recognition attacks and improving the accuracy of facial recognition attack interception.
[0113] The interception device 40 for face recognition attacks in the embodiments of this application will be described in detail below.
[0114] In some embodiments of this application, the segmentation module 402 can be specifically used to extract features from the background regions of the live face images of N users based on transaction data, so as to obtain the background feature data of the live face image of each user.
[0115] Based on the background feature data of each user's biopsy face image, the biopsy face images of N users are clustered to obtain at least one image cluster.
[0116] In some embodiments of this application, the partitioning module 402 can be specifically used to cluster the biopsy face images of N users based on the similarity between the background feature data of the biopsy face images of at least two users using an unsupervised clustering algorithm, to obtain at least one initial image cluster.
[0117] At least one image cluster is determined from at least one initial image cluster; wherein the image cluster satisfies at least one of the following conditions: the number of biopsy face images in the image cluster is greater than or equal to the number of reference images, and the range of the average centroid distance interval of the image cluster belongs to the range of the reference average centroid distance interval.
[0118] In some embodiments of this application, the face recognition attack interception device 40 in this application embodiment may further include a comparison module, which is used to perform pixel consistency comparison on the background regions of the live face images of at least two users in the image cluster when the transaction data includes the transaction time, so as to obtain pixel comparison data of at least two background regions.
[0119] The determination module 403 can also be used to determine at least two users corresponding to the pixel comparison data as risk users when the data value of the pixel comparison data is greater than or equal to a first preset threshold.
[0120] In some embodiments of this application, the comparison module in the embodiments of this application can be used to perform pixel consistency comparison on the background regions of at least two users' live face images in an image cluster through a comparison algorithm to obtain pixel comparison data of at least two background regions; wherein, the comparison algorithm includes at least one of the following: normalized difference square algorithm, normalized correlation coefficient evaluation algorithm.
[0121] In some embodiments of this application, the face recognition attack interception device 40 in this application embodiment may further include a matching module, which is used to perform correlation matching on the background areas of at least two users' live face images in the image cluster based on the reference geographic attribute data of the transaction location when the transaction data includes the transaction location, to obtain content matching data.
[0122] The determination module 403 can also be used to determine at least two users corresponding to the content matching data as risk users when the data value of the content matching data is less than or equal to a second preset threshold.
[0123] In some embodiments of this application, the reference geographic attribute data includes at least one of the following: weather condition attribute data, indoor and outdoor environment attribute data, background complexity attribute data, scene function attribute data, and administrative division attribute data.
[0124] In some embodiments of this application, the face recognition attack interception device 40 in this application embodiment may further include a comparison module, which is used to perform pixel consistency comparison on the background regions of the live face images of at least two users in the image cluster when the transaction data includes the transaction time and transaction location, so as to obtain pixel comparison data of at least two background regions.
[0125] Furthermore, the face recognition attack interception device 40 in this application embodiment may also include a matching module, which is used to perform correlation matching on the background areas of at least two users' live face images in the image cluster based on reference geographic attribute data of the transaction location, to obtain content matching data.
[0126] The determination module 403 can also be used to determine the risk score data of each of at least two users based on pixel comparison data, a first risk confidence level corresponding to the pixel comparison data, content matching data, and a second risk confidence level corresponding to the content matching data.
[0127] The determination module 403 can also be used to determine the user corresponding to the risk score data as a risk user when the risk score data is greater than or equal to a third preset threshold.
[0128] In some embodiments of this application, the face recognition attack interception device 40 in this application embodiment may further include a verification module, which is used to verify the background area of the live face images of at least two users in the image cluster according to the risk verification data, and obtain verification result data; wherein, the risk verification data includes at least one of the following data: reference geographical attribute data of adjacent transaction locations of the transaction location, and live face images of adjacent transaction times of the transaction time.
[0129] The determination module 403 can also be used to determine the user corresponding to the content matching data as a risk user when the data value of the re-verification result data is higher than or equal to the fourth preset threshold.
[0130] In some embodiments of this application, the face recognition attack interception device 40 in this application embodiment may further include a sending module, which is used to send an interception instruction to the detection terminal corresponding to the risk user based on the interception data;
[0131] The detection end includes at least one of the following: the electronic device of the risky user, and the face detection system used to examine the live-shot face image of the risky user;
[0132] The interception instruction includes at least one of the following: a first interception instruction and a second interception instruction; the first interception instruction is used to instruct the electronic device of the risky user to increase the difficulty level of liveness detection for the risky user; the second interception instruction is used to instruct the face detection system to increase the hack detection interception level of the liveness detection face image of the risky user.
[0133] This application also provides a computer device. A detailed description is provided with reference to Figure 5.
[0134] Figure 5 is a schematic diagram of the structure of a computer device provided in one embodiment of this application.
[0135] As shown in Figure 5, the computer device may include at least one of the following as described in the embodiments of this application: an electronic device, a server. The computer device may include a processor 501 and a memory 502 storing computer program instructions.
[0136] Specifically, the processor 501 may include a central processing unit (CPU), an application-specific integrated circuit (ASTC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0137] Memory 502 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 502 may include removable or non-removable (or fixed) media. Where appropriate, memory 502 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 502 is non-volatile solid-state memory. In a particular embodiment, memory 502 includes solid-state storage (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0138] The processor 501 reads and executes computer program instructions stored in the memory 502 to implement any of the methods for intercepting face recognition attacks in the above embodiments.
[0139] In one example, the computer device may also include a communication interface 503 and a bus 510. As shown in Figure 5, the processor 501, memory 502, and communication interface 503 are connected via the bus 510 and communicate with each other.
[0140] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0141] Bus 510 includes hardware, software, or both, that couples components of a flow control device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard System (ETSA) bus, a Front Side Bus (FSB), an HyperTransport (HT) interconnect, an Industry Standard System (TSA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel System (MCA) bus, a Peripheral Component Interconnect (PCT) bus, a PCT-Express (PCT-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 510 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0142] The detection device for face recognition attacks can execute the face recognition attack interception method in the embodiments of this application, thereby realizing the face recognition attack interception method and device described in conjunction with Figures 1 to 4.
[0143] Furthermore, in conjunction with the methods for intercepting face recognition attacks described in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the face recognition attack interception methods described in the above embodiments.
[0144] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0145] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASTCs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0146] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0147] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for intercepting facial recognition attacks, comprising: Obtain transaction data and live face images of N users within a first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1. Based on the transaction data, the live face images of the N users are divided to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background regions, where the background regions are non-face regions in the live face images. The interception data for risky users corresponding to the biopsied face images in the image cluster is determined. The interception data is used to indicate the interception of face recognition attacks related to the risky users within a second time window.
2. The interception method according to claim 1, wherein, The step of dividing the live face images of the N users based on the transaction data to obtain at least one image cluster includes: Based on the transaction data, feature extraction is performed on the background regions in the live face images of the N users to obtain background feature data of the live face image of each user; Based on the background feature data of the biopsy face images of each user, the biopsy face images of the N users are clustered to obtain at least one image cluster.
3. The interception method according to claim 2, wherein, The process of clustering the biopsy face images of the N users based on the background feature data of each user's biopsy face image to obtain at least one image cluster includes: Using an unsupervised clustering algorithm, based on the similarity between the background feature data of at least two of the users' live face images, the live face images of the N users are clustered to obtain at least one initial image cluster; The at least one image cluster is determined from the at least one initial image cluster; wherein the image cluster satisfies at least one of the following conditions: the number of biopsy face images in the image cluster is greater than or equal to the number of reference images, and the range of the average centroid distance interval of the image cluster belongs to the range of the reference average centroid distance interval.
4. The interception method according to any one of claims 1 to 3, wherein, The transaction data includes the transaction time; before determining the interception data of the risky user corresponding to the live face image in the image cluster, the method further includes: Pixel consistency comparison is performed on the background regions of at least two users' biopsy face images in the image cluster to obtain pixel comparison data of at least two background regions; If the data value of the pixel comparison data is greater than or equal to a first preset threshold, at least two users corresponding to the pixel comparison data are identified as the risk users.
5. The interception method according to claim 4, wherein, The step of performing pixel consistency comparison on the background regions of at least two users' live face images in the image cluster to obtain pixel comparison data for at least two background regions includes: By using a comparison algorithm, pixel consistency comparison is performed on the background regions of at least two users' live face images in the image cluster to obtain pixel comparison data of at least two background regions; wherein, the comparison algorithm includes at least one of the following: normalized difference square algorithm, normalized correlation coefficient evaluation algorithm.
6. The interception method according to any one of claims 1 to 3, wherein, The transaction data includes the transaction location; before determining the interception data of the risky user corresponding to the live-detected face image in the image cluster, the method further includes: Based on the reference geographic attribute data of the transaction location, correlation matching is performed on the background regions of the live face images of at least two users in the image cluster to obtain content matching data; If the data value of the content matching data is less than or equal to a second preset threshold, at least two users corresponding to the content matching data are identified as the risk users.
7. The interception method according to claim 6, wherein, The reference geographic attribute data includes at least one of the following: weather condition attribute data, indoor and outdoor environment attribute data, background complexity attribute data, scene function attribute data, and administrative division attribute data.
8. The interception method according to any one of claims 1 to 3, wherein, The transaction data includes the transaction time and the transaction location; before determining the interception data of the risky user corresponding to the live-detected face image in the image cluster, the method further includes: Pixel consistency comparison is performed on the background regions of the biopsy face images of at least two users in the image cluster to obtain pixel comparison data of at least two background regions; and, based on the reference geographic attribute data of the transaction location, correlation matching is performed on the background regions of the biopsy face images of at least two users in the image cluster to obtain content matching data. Based on the pixel comparison data, the first risk confidence level corresponding to the pixel comparison data, the content matching data, and the second risk confidence level corresponding to the content matching data, the risk score data for each of the at least two users is determined; If the risk score data is greater than or equal to a third preset threshold, the user corresponding to the risk score data is identified as the risk user.
9. The interception method according to claim 4, 6 or 8, wherein, Before determining the interception data of the risky user corresponding to the live face image in the image cluster, the method further includes: Based on the risk verification data, the background areas of the live face images of at least two users in the image cluster are re-verified to obtain re-verification result data; wherein, the risk verification data includes at least one of the following: reference geographical attribute data of adjacent transaction locations of the transaction location, and live face images of adjacent transaction times of the transaction time; If the data value of the re-verification result is higher than or equal to the fourth preset threshold, the user corresponding to the content matching data is identified as the risk user.
10. The interception method according to claim 1, wherein, The method further includes: Based on the intercepted data, an interception command is sent to the detection terminal corresponding to the risky user; The detection end includes at least one of the following: the electronic device of the risk user, and a face detection system for verifying the live-shot face image of the risk user; The interception command includes at least one of the following: a first interception command and a second interception command; the first interception command is used to instruct the electronic device of the risky user to increase the difficulty level of live detection of the risky user; the second interception command is used to instruct the face detection system to increase the hack detection interception level of the live detection face image of the risky user.
11. A device for intercepting facial recognition attacks, comprising: The acquisition module is used to acquire transaction data and live face images of N users within a first time window. The transaction data includes at least one of the following: transaction time and transaction location, where N is an integer greater than 1. The segmentation module is used to segment the live face images of the N users based on the transaction data to obtain at least one image cluster. The image cluster includes live face images of at least two users with similar background regions, and the background regions are non-face regions in the live face images. The determination module is used to determine the interception data of the risky user corresponding to the live face image in the image cluster. The interception data is used to indicate the interception of face recognition attacks related to the risky user within a second time window.
12. An electronic device, the device comprising: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the steps of the method for intercepting face recognition attacks as described in any one of claims 1-10.
13. A storage medium storing computer program instructions, wherein the computer program instructions, when executed by a processor, implement the steps of the method for intercepting face recognition attacks as described in any one of claims 1-10.
14. A computer program product, characterized in that, The program product is stored in a storage medium and is executed by at least one processor to implement the steps of the method for intercepting face recognition attacks as described in any one of claims 1-10.