A face payment method, device and equipment
By capturing images and performing face detection in facial recognition payment devices, determining user intent, and automatically responding to payment operations, the problem of manual operation required in existing technologies is solved, realizing contactless facial recognition payment and improving convenience and security.
Patent Information
- Application Number
- CN202210126314.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-02-10
AI Technical Summary
Existing facial recognition payment methods require users to manually operate the facial recognition payment device, which is inconvenient and affects user hygiene and safety.
The system collects target images through facial recognition payment devices, performs face detection, determines whether the face detection data meets the automatic operation trigger conditions, determines the user's intended operation based on the processing progress information of pending payment orders, and directly responds to the user's intended operation to complete the payment without requiring manual input from the user.
It enables contactless facial recognition payment, avoiding manual operation by users and improving the convenience and hygiene of the operation.
Smart Images

Figure CN114549002B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of payment technology, and in particular to a facial recognition payment method, device, and equipment. Background Technology
[0002] With the development of computer and optical imaging technologies, facial recognition payment is becoming increasingly popular. Merchants can install facial recognition payment devices at their premises for customers to use when making offline purchases. Currently, users can manually operate the controls or buttons on the facial recognition payment device to initiate the payment process based on their personal preferences and input their payment intentions, thus completing the facial recognition payment.
[0003] Therefore, how to achieve contactless facial recognition payment based on user wishes, so as to avoid users having to manually operate the facial recognition payment device during the payment process, has become an urgent technical problem to be solved. Summary of the Invention
[0004] The embodiments of this specification provide a facial recognition payment method, device, and equipment that can realize contactless facial recognition payment based on user wishes, thereby avoiding the need for users to manually operate the facial recognition payment device during the facial recognition payment process.
[0005] To solve the above-mentioned technical problems, the embodiments in this specification are implemented as follows:
[0006] This specification provides an embodiment of a facial recognition payment method, including:
[0007] Facial recognition payment devices collect target images;
[0008] The target image is processed to obtain face detection data;
[0009] Determine whether the face detection data meets the first automatic operation trigger condition, and obtain the first determination result;
[0010] If the first judgment result indicates that the face detection data meets the first automatic operation triggering condition, then the user's intended operation is determined based on the processing progress information of the pending payment orders at the payment device; the user's intended operation is an operation that can be input by the user via touch on the payment device.
[0011] In response to the user's requested action, the pending payment order is processed.
[0012] This specification provides an embodiment of a facial recognition payment device, comprising:
[0013] The image acquisition module is used to acquire target images;
[0014] A face detection module is used to detect and process the target image to obtain face detection data;
[0015] The first judgment module is used to determine whether the face detection data meets the first automatic operation triggering condition and obtain the first judgment result;
[0016] The user intention determination module is used to determine the user's intended operation based on the processing progress information of the pending payment order at the payment device if the first judgment result indicates that the face detection data meets the first automatic operation triggering condition; the user's intended operation is an operation that can be input by the user via touch on the payment device.
[0017] The user intention execution module is used to process the pending payment order in response to the user intention operation.
[0018] This specification provides an embodiment of a facial recognition payment device, comprising:
[0019] At least one processor; and,
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0022] Facial recognition payment devices collect target images;
[0023] The target image is processed to obtain face detection data;
[0024] Determine whether the face detection data meets the first automatic operation trigger condition, and obtain the first determination result;
[0025] If the first judgment result indicates that the face detection data meets the first automatic operation triggering condition, then the user's intended operation is determined based on the processing progress information of the pending payment orders at the payment device; the user's intended operation is an operation that can be input by the user via touch on the payment device.
[0026] In response to the user's requested action, the pending payment order is processed.
[0027] At least one embodiment provided in this specification can achieve the following beneficial effects:
[0028] The facial recognition payment device can be pre-set with a first automatic operation trigger condition. If the face detection data of the target image collected by the facial recognition payment device meets the first automatic operation trigger condition, it indicates that the user intends to continue processing the pending payment order at the payment device. Therefore, based on the processing progress information of the pending payment order, the user's intention to input the pending payment order via touch can be determined; and the user's intention to input the pending payment order can be responded to directly without waiting for the user to manually input the user's intention. This ensures the credibility and accuracy of the user's intention and avoids the user having to manually operate the facial recognition payment device during the facial recognition payment process, realizing contactless facial recognition payment, which is convenient and fast. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a schematic diagram illustrating an application scenario of a facial recognition payment method as described in the embodiments of this specification.
[0031] Figure 2 A flowchart illustrating a facial recognition payment method provided in an embodiment of this specification;
[0032] Figure 3 This is a swimlane flowchart illustrating a facial recognition payment method provided in an embodiment of this specification.
[0033] Figure 4 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a facial recognition payment device;
[0034] Figure 5 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a facial recognition payment device. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of one or more embodiments of this specification.
[0036] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0037] In existing technologies, facial recognition payment is a novel payment method based on machine vision technology. However, current facial recognition payment methods typically still require users to touch the screen of the payment device. For example, when a user is about to use the facial recognition payment device to make a payment, if the screen is displaying a standby page, the user needs to touch the screen to activate the facial recognition interface; or, if the screen is displaying a payment confirmation interface, the user needs to touch the screen to confirm the payment order. Because facial recognition payment requires user contact with the device, it is not only inconvenient for users but also affects user hygiene and safety.
[0038] To address the shortcomings of existing technologies, this solution provides the following embodiments:
[0039] Figure 1 This is a schematic diagram illustrating an application scenario of a facial recognition payment method provided in the embodiments of this specification; for example... Figure 1 As shown, in an offline retail scenario, when user 110 pays using facial recognition, the facial recognition payment device 120 can capture images in real time. By inspecting and processing the captured images, facial detection data is obtained. Then, by determining whether the facial detection data meets the automatic operation trigger conditions, the device can ascertain user 110's operational intention and directly respond to it without requiring user 110 to manually execute the desired action. In practical applications, if the processing progress information of the pending payment order is "waiting to initiate payment order," and the automatic operation trigger conditions are met, the facial recognition payment device 120 can automatically display the facial recognition interface; if the pending payment order is in the payment confirmation stage, and the automatic operation trigger conditions are met, then the payment for the pending payment order will be processed.
[0040] In addition, the facial recognition payment device 120 can also assist in judging the user's intention to operate by recognizing the user's voice. For example, when the user's voice contains keywords such as "confirm payment", the automatic operation trigger conditions can be appropriately relaxed to improve the success rate of recognizing the user's intention to operate while ensuring the credibility of the determined user's intention to operate.
[0041] In the embodiments of this specification, the user 110's intended operation is determined by the above method, which can replace the existing facial recognition technology that requires the user to touch and input their intended operation at the facial recognition payment device, thereby realizing contactless facial recognition payment, which is convenient, fast, safe and hygienic.
[0042] Next, a facial recognition payment method provided in the embodiments of the specification will be described in detail with reference to the accompanying drawings:
[0043] Figure 2 This is a flowchart illustrating a facial recognition payment method provided in an embodiment of this specification. From the perspective of the implementing entity, Figure 2 The method shown can be executed by either the server of the facial recognition payment device or the facial recognition payment device itself. For example... Figure 2 As shown, the process may include the following steps:
[0044] Step 202: The facial recognition payment device collects the target image.
[0045] In the embodiments of this specification, the facial recognition payment device is a device that can be used to assist users in making facial recognition payments. The facial recognition payment device can be an existing cash register device with image acquisition and facial recognition payment application installed, or it can be a mobile phone, tablet or computer with facial recognition payment application installed, or it can also include devices with facial recognition payment capabilities such as smart vending machines, smart express cabinets, smart storage cabinets, and shared item rental terminals, etc., without specific limitations.
[0046] In practical applications, the facial recognition payment device can continuously acquire images through an image acquisition device, and the target image can be the latest image acquired by the facial recognition payment device in real time. The target image can be a color image, or an infrared image, a depth image, etc., without specific limitations.
[0047] Step 204: Perform detection processing on the target image to obtain face detection data.
[0048] In this embodiment of the specification, since it is necessary to determine the user's operational intention based on the user's facial information, at least face detection should be performed in step 204. If a user's face is detected, subsequent processing operations can proceed. If no user's face is detected, it indicates that there is currently no user who needs to operate the facial recognition payment device, and the process can proceed to the end step.
[0049] In practical applications, since more comprehensive facial information is needed to accurately identify a user's intention to perform an action, additional detection and processing operations can be performed on the target image as required. These operations include face depth detection, face sharpness detection, and selected face consistency detection to obtain face detection data. This face detection data can then be used to reflect whether the user intends to perform the action.
[0050] Step 206: Determine whether the face detection data meets the first automatic operation trigger condition, and obtain the first judgment result.
[0051] In the embodiments of this specification, the first automatic operation triggering condition can refer to the conditions that the facial recognition payment device must meet to automatically respond to a touch operation when no manual touch operation is detected from the user. In practical applications, the first automatic operation triggering condition may include one or more of the following: distance condition between the user's face and the facial recognition payment device (face depth condition), face image quality condition, and continuous consistency condition of selected faces. These will be explained in detail in subsequent embodiments and will not be elaborated here.
[0052] Step 208: If the first judgment result indicates that the face detection data meets the first automatic operation triggering condition, then the user's intended operation is determined according to the processing progress information of the pending payment order at the payment device; the user's intended operation is an operation that can be input by the user via touch on the payment device.
[0053] In the embodiments of this specification, the user intention operation can refer to the next operation that needs to be manually performed by the user, determined based on the processing progress information of the pending payment order. User intention operations are typically used to clarify the user's intention to continue the payment process; for example, initiating facial recognition payment or confirming the payment.
[0054] The processing progress information for pending payment orders may include, but is not limited to: waiting to initiate a payment order, waiting for facial recognition verification, waiting for confirmation of the payment account and payment order, and payment order processing completion. In practical applications, if the processing progress information for a pending payment order is "waiting to initiate a payment order," the determined user intention operation can be the operation to initiate facial recognition payment. Conversely, if the processing progress information for a pending payment order is "waiting for confirmation of the payment account and payment order," the determined user intention operation can be the operation to confirm payment.
[0055] In the embodiments described in this specification, once the facial recognition payment device determines that the first automatic operation triggering condition has been met, it can be considered to have received the user's intended operation, thus eliminating the need for manual execution by the user. In other words, the user's intended operation can include not only touch input at the facial recognition payment device but also contactless input. Specifically, the user can perform a specified action within the image acquisition range of the facial recognition payment device to input their intended operation. This specified action may include remaining stationary while facing the facial recognition payment device. Of course, the specified action may also include facial movements such as nodding, smiling, or opening the mouth wide, without specific limitations.
[0056] In practical applications, if the first judgment result indicates that the face detection data does not meet the first automatic operation triggering condition, it usually indicates that the user does not have the intention to continue processing the pending payment order, and thus can jump to the end step.
[0057] Step 210: In response to the user's intended action, process the pending payment order.
[0058] Figure 2 The method described above involves facial recognition payment devices determining whether the facial detection data meets the first automatic operation trigger condition to identify whether the user intends to advance the processing progress of the pending payment order. If so, the device can combine the processing progress information of the pending payment order to determine and directly respond to the user's intention operation that can be manually input by the user, thus processing the pending payment order without requiring the user to manually input the information. This ensures the credibility and accuracy of the user's intention while avoiding the user having to manually operate the facial recognition payment device during the facial recognition payment process, achieving contactless facial recognition payment that is convenient and fast.
[0059] based on Figure 2 In addition to the method described in the embodiments of this specification, some specific implementation schemes of the method are also provided, which will be described below.
[0060] In practical applications, judging a user's operational intention solely based on a target image may lead to inaccurate identification of that intention. Therefore... Figure 2The method described above, before determining the user's intended action, may also include:
[0061] The facial recognition payment device collects sound data; the time interval between the sound data collection time and the target image collection time is less than or equal to a first threshold.
[0062] The sound data is processed by speech recognition to obtain the user's speech recognition result.
[0063] Determine whether the user's voice recognition result meets the second automatic operation triggering condition to obtain a second determination result; the second automatic operation triggering condition is that the user's voice recognition result contains a preset keyword corresponding to the user's intended operation to be performed; the user's intended operation to be performed is an operation that can be performed by the user through touch input on the payment device for the pending payment order.
[0064] Correspondingly, determining whether the face detection data meets the first automatic operation trigger condition may specifically include:
[0065] If the second judgment result indicates that the user's voice recognition result does not meet the second automatic operation triggering condition, then it is determined whether the face detection data meets the third automatic operation triggering condition; the third automatic operation triggering condition may include: a first face depth condition, a first face image quality condition, and a selected face continuity and consistency condition.
[0066] If the second judgment result indicates that the user's voice recognition result meets the second automatic operation triggering condition, then it is determined whether the face detection data meets the fourth automatic operation triggering condition; the fourth automatic operation triggering condition may include: the second face depth condition and the second face image quality condition.
[0067] In the embodiments described in this specification, during the process of acquiring the target image, the facial recognition payment device may also acquire audio data. If the time interval between the acquisition time of the audio data and the acquisition time of the target image is small, it is highly likely that the audio data and the target image come from the same user. Therefore, the user's intention can be identified by combining audio data acquired at a time interval less than or equal to a first threshold with the target image, thereby improving the accuracy of the user intention identification result. In practical applications, the first threshold can be set according to the actual situation and is not specifically limited thereto.
[0068] In the embodiments of this specification, the speech recognition processing can be used to convert the vocabulary content of human speech in the sound data collected by the facial recognition payment device into computer-readable input (i.e., user speech recognition results). The user intention operation to be executed can refer to the user intention operation waiting for user touch input, determined according to the processing progress of the pending payment order; the keywords can be phrases used to execute the user intention operation to be executed. For example, if the processing progress information of the pending payment order is waiting for confirmation of the payment account and payment order, then the user intention operation to be executed can be a payment confirmation operation, and the keywords can be "confirm payment", "allow payment", "agree to payment", etc.
[0069] In this embodiment of the specification, if the voice recognition result fails to meet the second automatic operation triggering condition, the facial recognition payment device can perform single-modal recognition of the user's operation intention based on the face detection data according to steps 206-208. In this case, the face detection data needs to meet the third automatic operation triggering condition, namely the first face depth condition, the first face image quality condition, and the selected face continuity and consistency condition. In practical applications, if the face detection data does not meet any of the first face depth condition, the first face image quality condition, and the selected face continuity and consistency condition, it can be considered that the face detection data does not meet the third automatic operation triggering condition, which usually indicates that the user does not have the intention to continue processing the pending payment order, and thus the process can jump to the end step.
[0070] If the voice recognition result meets the second automatic operation trigger condition, the facial recognition payment device can simultaneously combine the voice recognition result and face detection data to perform multimodal recognition of the user's operation intention. In this case, the face detection data needs to meet the fourth automatic operation trigger condition, namely the second face depth condition and the second face image quality condition. In practical applications, if the face detection data does not meet either the second face depth condition or the second face image quality condition, it can be considered that the face detection data does not meet the fourth automatic operation trigger condition. This usually indicates that the user does not intend to continue processing the pending payment order, and thus the process can proceed to the end step.
[0071] In practical applications, when performing image-based unimodal judgment of user intent, since the user does not express their intent verbally, the requirements for the third automatic operation triggering condition are generally more stringent than those for the fourth automatic operation triggering condition. That is, the first face depth condition can be more stringent than the second face depth condition, and the first face image quality condition can be more stringent than the second face image quality condition. Furthermore, the number of third automatic operation triggering conditions can be greater; for example, the third automatic operation triggering condition can also include a condition for the continuous consistency of selected faces. In other words, the fourth automatic operation triggering condition for multimodal judgment of user intent is relatively lenient compared to the third automatic operation triggering condition for image-based unimodal judgment.
[0072] In practical applications, facial recognition payment devices can perform parallel processing of target image and audio data acquisition and processing. That is, the device can first determine whether the face detection data meets the third and fourth automatic operation trigger conditions, without needing to generate a second judgment result indicating whether the user's voice recognition result contains preset keywords. Subsequently, it only needs to select the judgment result corresponding to the third or fourth automatic operation trigger condition based on the second judgment result, thus improving the efficiency of facial recognition payments.
[0073] Of course, after generating a second judgment result indicating whether the user's voice recognition result contains preset keywords, the facial recognition payment device can then execute the step of judging whether the face detection data meets the third or fourth automatic operation trigger condition. Since only one of the third or fourth automatic operation trigger conditions needs to be judged, the amount of resources required for the judgment operation can be reduced.
[0074] In addition, facial recognition payment devices can also avoid collecting voice data and instead use face detection data to make image-based single-modal judgments of user intent, thereby reducing the implementation cost of this solution.
[0075] In the embodiments described in this specification, the facial recognition payment device combines voice with voice to assist in judging the user's operational intention. This can improve the accuracy of the determined user's operational intention while ensuring contactless payment, thereby enhancing the security of facial recognition payment.
[0076] In the embodiments of this specification, since the user will be close to the facial recognition payment device when making a facial recognition payment, the user's intention can be determined by the distance between the user and the facial recognition payment device.
[0077] Based on this, the detection processing of the target image to obtain face detection data may specifically include:
[0078] Face detection is performed on the target image to obtain the predicted region of the user's face in the target image.
[0079] A face depth detection process is performed on the predicted region of the user's face in the target image to obtain the user's face depth value; the user's face depth value can be used to characterize the distance between the user's face and the face-scanning payment device.
[0080] Correspondingly, determining whether the face detection data meets the third automatic operation triggering condition may specifically include:
[0081] Determine whether the user's face depth value meets the first face depth condition, wherein the first face depth condition is that the user's face depth value is less than or equal to the second threshold;
[0082] Correspondingly, determining whether the face detection data meets the fourth automatic operation trigger condition may specifically include:
[0083] Determine whether the user's face depth value meets the second face depth condition, wherein the second face depth condition is that the user's face depth value is less than or equal to the third threshold.
[0084] In this embodiment of the specification, the face detection can be used to determine the location information of the predicted region of the user's face in the target image. The location information of the predicted region of the user's face may include the coordinate information, size information, etc. of the predicted region of the user's face in the target image.
[0085] The aforementioned face depth detection can be used to determine the depth value of the user's face. There are various ways to implement face depth detection. For example, it can be achieved by acquiring target images carrying depth information using depth cameras such as structured light depth cameras or binocular vision cameras, and then using this depth information to perform face depth detection. Alternatively, it can be achieved using monocular depth estimation techniques based on deep learning.
[0086] In the embodiments of this specification, the third automatic operation triggering condition may include a first face depth condition. In practical applications, when judging whether the face detection data meets the third automatic operation triggering condition, if the user's face depth value does not meet the first face depth condition, it usually indicates that the user does not have the intention to continue processing the pending payment order, and thus the process can jump to the end step.
[0087] In the embodiments of this specification, the fourth automatic operation triggering condition may include a second face depth condition. In practical applications, when judging whether the face detection data meets the fourth automatic operation triggering condition, if the user's face depth value does not meet the second face depth condition, it usually indicates that the user does not have the intention to continue processing the pending payment order, and thus the process can jump to the end step. In practical applications, methods such as measuring the user's face depth value by acquiring depth images through a depth camera, or estimating the user's face depth value using monocular depth estimation technology based on deep learning, both rely on extremely high computing power. This not only places high demands on the performance of the facial recognition payment device but also consumes a lot of computing time, affecting the operating efficiency of facial recognition payment and thus hindering the widespread application of facial recognition payment devices.
[0088] In the embodiments of this specification, since the camera has the imaging characteristic of near objects appearing larger and distant objects appearing smaller, when the camera focal length is constant, objects that are closer to the camera will occupy a larger proportion in the formed image; therefore, the depth of the user's face (i.e., the distance between the user's face and the facial recognition payment device) can be estimated based on the area proportion of the predicted region of the user's face in the target image.
[0089] Based on this, face depth detection processing is performed on the predicted region of the user's face in the target image to obtain the user's face depth value, which may specifically include:
[0090] Calculate the area ratio of the predicted region of the user's face in the target image.
[0091] The user's face depth value is determined based on the area ratio.
[0092] In the embodiments of this specification, the area ratio of the predicted region of the user's face in the target image is the quotient of the area of the predicted region of the user's face and the area of the target image.
[0093] The facial recognition payment device has a pre-defined correspondence between the preset face area ratio and the preset face depth value. Therefore, after determining the area ratio of the user's face, the preset face depth value corresponding to the area ratio of the user's face can be used as the user's face depth value.
[0094] In practical applications, the relationship between the preset face area ratio and the preset face depth value may be affected by factors such as the user's face location in the target image (e.g., image edge, image center), height, gender, and age. Therefore, when estimating the user's face depth value using the face area ratio, the ratio can be corrected based on the user's face location in the target image, height, gender, and age to improve the accuracy of the face depth detection results.
[0095] In the embodiments of this specification, the depth value of the user's face is estimated by calculating the area ratio of the predicted region of the user's face in the target image. Compared with the method of acquiring depth images or executing monocular depth estimation algorithms, the calculation process is relatively simple, which helps to reduce the cost of facial recognition payment devices, improve the operating efficiency of facial recognition payment, and facilitate the promotion of facial recognition payment.
[0096] In the embodiments described in this specification, since the first face depth condition is usually more stringent than the second face depth condition, the second threshold should generally be less than the third threshold. For example, the second threshold can be set to 70 cm and the third threshold can be set to 90 cm. In practical applications, the second and third thresholds can be arbitrarily set according to the actual situation, and no specific limitation is made therein.
[0097] In practical applications, as illustrated in this embodiment, users intending to pay via facial recognition are more likely to approach the facial recognition payment device. Therefore, the facial depth value can reflect the user's willingness to proceed with the payment order. Based on this, the user's intention to pay can be accurately determined by judging whether the facial depth condition is met. Furthermore, by selecting different thresholds to set the facial depth condition based on whether the user's voice recognition result meets the second automatic operation trigger condition, the success rate of contactless facial recognition payment can be improved while ensuring accurate recognition of the user's intention.
[0098] In practical applications, the embodiments of this specification can also identify the user's operational intentions based on the detection results of the face image quality.
[0099] Based on this, the detection processing of the target image to obtain face detection data may specifically include:
[0100] Face detection is performed on the target image to obtain the predicted region of the user's face in the target image.
[0101] A face image quality detection process is performed on the predicted region of the user's face in the target image to obtain a user face image quality score; the user face image quality score can be used to characterize the accuracy of determining the user's intended operation based on the face image in the predicted region of the user's face.
[0102] Correspondingly, determining whether the face detection data meets the third automatic operation triggering condition may specifically include:
[0103] Determine whether the quality score of the user's face image meets the first face image quality condition, wherein the first face image quality condition is that the quality score of the user's face image is greater than or equal to the fourth threshold.
[0104] Correspondingly, determining whether the face detection data meets the fourth automatic operation trigger condition may specifically include:
[0105] Determine whether the quality score of the user's face image meets the second face image quality condition, wherein the second face image quality condition is that the quality score of the user's face image is greater than or equal to the fifth threshold.
[0106] In the embodiments of this specification, face image quality detection processing can refer to image quality evaluation of the predicted region of a user's face, and may include, but is not limited to, one or more of the following: sharpness detection, face pose information detection, cheating behavior detection, and expression detection.
[0107] The sharpness detection can be used to detect defocus blur and motion blur in the predicted region of a user's face. Many factors can cause image blurring during image acquisition, transmission, and processing. Incorrect focusing during target image acquisition can cause defocus blur, and the relative movement of the face and camera can cause motion blur in a certain direction. Sharpness detection can evaluate the sharpness of the predicted region of a user's face based on pixel-based techniques, transform domain techniques, and image gradient techniques. The sharpness information obtained through sharpness detection can not only characterize whether the predicted region of the user's face is sharp, but also whether the user's face exhibits a motion trend.
[0108] The facial pose information detection can be used to estimate the facial pose angles in the pitch, yaw, and roll directions based on the facial image within the predicted region of the user's face. Under normal circumstances, an intentional user's face will be directly facing the screen of the facial recognition payment device; therefore, the user's intention to operate can be determined based on the facial pose information.
[0109] The cheating detection function can be used to detect whether a user's face is obscured by obstacles such as limbs or masks. An obscured face usually indicates that the user does not intend to perform the action.
[0110] Correspondingly, the user's facial image quality score can be calculated from one or more of the following: sharpness information, facial pose information, cheating behavior information, expression detection information, and liveness detection information; or, it can be calculated by multiplying the above information by the corresponding weighting coefficients, without specific limitations. However, generally speaking, the higher the user's facial image quality score, the more accurate the determination of the user's operational intention based on the user's facial image will be.
[0111] In the embodiments described in this specification, since the quality conditions for the first face image are usually more stringent than those for the second face image, the fourth threshold can typically be set to a value greater than the fifth threshold. In practical applications, the fourth and fifth thresholds can be arbitrarily set according to the actual situation, and no specific limitations are imposed on them.
[0112] In the embodiments of this specification, the third automatic operation triggering condition may include the first face image quality condition. In practical applications, when judging whether the face detection data meets the third automatic operation triggering condition, if the user's face image quality score does not meet the first face image quality condition, it usually indicates that the user does not have the intention to continue processing the pending payment order, and thus the process can jump to the end step.
[0113] In the embodiments of this specification, the fourth automatic operation triggering condition may include a second face image quality condition. In practical applications, when judging whether the face detection data meets the fourth automatic operation triggering condition, if the user's face image quality score does not meet the second face image quality condition, it usually indicates that the user does not have the intention to continue processing the pending payment order, and thus the process can jump to the end step.
[0114] In the embodiments of this specification, by setting the first face image quality condition and the second face image quality condition, when a user advances the processing progress of an order to be paid by scanning their face, it is possible to accurately determine whether the user has the intention to operate, thereby helping to protect the user's property security.
[0115] In the examples in this manual, face image quality detection processing can be performed on face images based on deep learning technology.
[0116] Specifically, the step of performing face image quality detection processing on the predicted region of the user's face in the target image to obtain a user face image quality score may include:
[0117] The face image in the predicted region of the user's face is input into the face image quality detection model to obtain the user face image quality score of the face image output by the face image quality detection model. The face image quality detection model is obtained by training a deep learning model in advance using face image samples carrying face image quality labels. The face image quality labels are label data determined based on at least one of the face image sample's sharpness information, face image sample's face pose information, and face image sample's cheating behavior information.
[0118] In the embodiments of this specification, the face image sample may include a user face image pre-labeled with a face image quality score. The user face image serving as the face image sample may be obtained by extracting the face region from images collected in application scenarios such as shopping malls, supermarkets, and retail stores. The face image quality score label carried by the face image sample may be generated based on at least one of the following: face image sample sharpness information, face pose information, and cheating behavior information. The face pose information may include face pose angle values in directions such as pitch, yaw, and roll.
[0119] The training process for a face image quality detection model can be as follows: Face image samples carrying quality score labels are used as input to a deep learning model, and the output is the predicted quality score of the user's face image. When the error between the predicted quality score and the labeled quality score is less than a preset value, it indicates that the prediction result of the trained face image quality detection model is relatively accurate, and model training can be stopped and the model can be put into use. In practical applications, the face image quality detection model can specifically be a deep learning model built based on a CNN convolutional neural network.
[0120] In this embodiment of the specification, a user face image quality score is generated by using a trained face image quality detection model. This helps to improve the accuracy of the obtained user face image quality score, which in turn helps to improve the accuracy of the determined user's operation intention, thereby ensuring the security of face payment.
[0121] In the embodiments of this specification, when a pedestrian passes by the facial recognition payment device, the user's face can also be detected from the target image. However, the pedestrian usually does not have the intention to make facial recognition payment. Therefore, the facial recognition payment device can detect whether the selected user in multiple consecutive frames of images is the same user, so as to improve the accuracy of the identified user's intention to operate.
[0122] Based on this, the detection processing of the target image to obtain face detection data may specifically include:
[0123] Face detection is performed on the target image to obtain the predicted region of the user's face in the target image.
[0124] Based on the location information of the predicted region of the user's face in the target image, it is determined whether the user's face is the selected face in the target image, and a third determination result is obtained.
[0125] If the third determination result indicates that the user's face is the selected face in the target image, then the user's face feature data is extracted from the predicted region of the user's face.
[0126] Correspondingly, determining whether the face detection data meets the third automatic operation triggering condition may specifically include:
[0127] Determine whether the user's facial feature data meets the condition of continuous consistency of the selected face. The condition of continuous consistency of the selected face is that the user's facial feature data of the selected face is consistent in each image of the multi-frame continuous images. The multi-frame continuous images are images continuously collected by the facial recognition payment device with a number greater than or equal to a sixth threshold. The multi-frame continuous images may include the target image.
[0128] In this embodiment, the target image may refer to the current frame image captured by the facial recognition payment device. The face detection can be multi-target detection, thereby determining the predicted regions of each user's face contained in the target image. In practical applications, a selected face can be chosen from the target image based on factors such as whether the predicted region of each user's face is located at the center of the target image and the size of the predicted region of each user's face. This selected face may refer to the face of the user currently intending to make a facial recognition payment at the facial recognition payment device. Typically, there is only one selected face in the target image.
[0129] In the embodiments of this specification, the process of determining the continuous consistency condition of selected faces can be as follows: the user face feature data of the selected face in the current frame image is compared with the user face feature data of each selected face in the previous N (N is the difference between the sixth threshold and 1) frames of the target image. If they are all consistent, it can be said that the user face feature data satisfies the continuous consistency condition of selected faces.
[0130] Alternatively, after processing each frame preceding the target image, the number of consecutive images that match the selected face in that frame is updated. For example, if the current number of images is 2, it means that the selected face is the same in the previous two frames of the target image. In this case, it is only necessary to compare whether the user's facial feature data of the selected face in the target image matches the user's facial feature data of the selected face in the previous frame. If they match, the number of images can be updated to 3; if they do not match, the number of images can be updated to 1. Subsequently, it is only necessary to determine whether the updated number of images is greater than the sixth threshold, thereby reducing the number of times user facial feature data is compared, reducing the computational load on the device, and improving the efficiency of facial recognition payment.
[0131] In the embodiments of this specification, the sixth threshold can be a positive integer greater than 1, and the sixth threshold can be set according to actual needs, without specific limitations. For example, the sixth threshold can be set to 8 frames, then the condition for continuous consistency of selected faces is that the user facial feature data of the selected faces in at least 8 consecutive frames, including the current frame, are consistent.
[0132] In this embodiment of the specification, the third automatic operation triggering condition may include a selected face continuity and consistency condition. In practical applications, when judging whether the face detection data meets the third automatic operation triggering condition, if the user's facial feature data does not meet the selected face continuity and consistency condition, it usually indicates that the user does not intend to continue processing the pending payment order, and thus the process can jump to the end step. In this embodiment of the specification, by setting a selected face continuity and consistency condition, the possibility of users without operational intention falsely triggering the facial recognition payment device can be reduced, thereby improving the accuracy of identifying the user's operational intention.
[0133] In the embodiments described in this specification, there can be various types of user-intentioned operations.
[0134] If the user's intended action is to activate facial recognition payment, then the process of responding to the user's intended action by processing the pending payment order may specifically include:
[0135] In response to the wake-up facial recognition payment operation, a facial recognition interface is displayed; the facial recognition interface can be used to assist users in performing facial recognition operations during the payment process for the pending payment order.
[0136] In the examples provided in this manual, waking up the facial recognition payment device refers to switching the device from standby mode to facial recognition payment mode, allowing the user to begin facial recognition payment. In practical applications, when the facial recognition payment device is in standby mode, it can either have its screen off or display a standby page; in response to the wake-up facial recognition payment operation, the device can jump from the off screen or standby page to the facial recognition interface.
[0137] In the embodiments of this specification, the user's intended operation may further include a payment confirmation operation, which can be used to instruct the deduction of the amount to be paid from the account to be deducted.
[0138] Based on this, before the facial recognition payment device acquires the target image, it may further include:
[0139] The system displays a payment confirmation interface that includes the account to be deducted and the amount to be paid; the account to be deducted is the user account determined by facial recognition of the user, and the amount to be paid is the transaction amount of the pending payment order at the payment device.
[0140] Correspondingly, the process of handling the pending payment order in response to the user's intended operation may specifically include:
[0141] In response to the payment confirmation operation, the amount to be paid is deducted from the account to be deducted for the pending payment order.
[0142] In the embodiments described in this specification, after a user scans their face using a facial recognition interface, the facial recognition payment device can determine the user's payment account as the account to be deducted based on the user's facial information, and display a payment confirmation interface that includes the account to be deducted and the amount to be paid. Of course, the payment confirmation interface may also display information such as the merchant and products in the pending payment order, and there are no specific limitations on this.
[0143] The payment confirmation operation refers to the user's confirmation of agreement to have the amount deducted from the account after verifying the relevant information of the pending payment order and the account to be deducted at the payment device. This allows the facial recognition payment device to deduct the amount from the account to complete the payment for the pending payment order.
[0144] Figure 3 This is a swimlane flowchart illustrating the facial recognition payment method provided in the embodiments of this specification. From the perspective of the implementing entity, Figure 3 The entities that can perform the method shown can include users and facial recognition payment devices.
[0145] like Figure 3 As shown in the embodiments of this specification, users can enter the image acquisition area of the facial recognition payment device, and users can also make voice commands.
[0146] Facial recognition payment devices can collect target image and audio data. Subsequently, the device can perform face detection on the target image; if the target image does not contain a face, the process can proceed to the end. Figure 3 (Not shown in the image). If a face image exists in the target image, face detection data can be obtained. Furthermore, the facial recognition payment device can also perform speech recognition processing on the sound data to obtain the user's speech recognition result; determine whether the user's speech recognition result meets the second automatic operation triggering condition to obtain a second judgment result.
[0147] If the voice recognition result fails to meet the second automatic operation triggering condition, the facial recognition payment device can perform single-modal recognition of the user's operation intention based on the face detection data. In this case, it is necessary to determine whether the face detection data meets the third automatic operation triggering condition, that is, whether it simultaneously meets the first face depth condition, the first face image quality condition, and the selected face continuity and consistency condition.
[0148] If the voice recognition result meets the second automatic operation triggering condition, the facial recognition payment device can simultaneously combine the voice recognition result and face detection data to perform multimodal recognition of the user's operation intention. In this case, it is necessary to determine whether the face detection data meets the fourth automatic operation triggering condition, that is, whether it simultaneously meets the second face depth condition and the second face image quality condition.
[0149] If the voice recognition result does not meet the second automatic operation trigger condition but the face detection data meets the third automatic operation trigger condition, or if the voice recognition result meets the second automatic operation trigger condition and the face detection data meets the fourth automatic operation trigger condition, then it can be indicated that the user has the intention to continue processing the pending payment order at the face recognition payment device. Therefore, based on the processing progress information of the pending payment orders at the payment device, the user's intended operation can be determined; in response to the user's intended operation, the pending payment order is processed, and after processing is completed, the process can proceed to the end step.
[0150] In practical applications, if the payment order to be processed is in the payment initiation stage, the user's intention operation can be to activate the facial recognition payment operation. In response to the user's intention operation, the facial recognition payment device can automatically display the facial recognition interface without requiring the user to manually touch the facial recognition payment device to activate the facial recognition interface.
[0151] If the pending payment order is in the payment confirmation stage, the user's intended action can be to confirm the payment for the account to be deducted and the amount to be paid displayed on the payment confirmation interface. In response to the user's intended action, the facial recognition payment device can deduct the amount to be paid from the account to be deducted for the pending payment order, thus eliminating the need for the user to manually touch the facial recognition payment device to click the confirm payment control.
[0152] In addition, if the face detection data does not meet any of the following conditions: first face depth condition, first face image quality condition, selected face continuity and consistency condition, second face depth condition, and second face image quality condition, it can usually indicate that the user does not have the intention to continue processing the pending payment order. Therefore, the process can be skipped to the end step to ensure the accuracy of the determined user's operational intention.
[0153] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods. Figure 4 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a facial recognition payment device. Figure 4 As shown, the device may include:
[0154] The image acquisition module 402 can be used to acquire target images.
[0155] The face detection module 404 can be used to detect and process the target image to obtain face detection data.
[0156] The first judgment module 406 can be used to determine whether the face detection data meets the first automatic operation triggering condition and obtain the first judgment result.
[0157] The user intention determination module 408 can be used to determine the user's intention operation based on the processing progress information of the pending payment order at the payment device if the first judgment result indicates that the face detection data meets the first automatic operation triggering condition; the user intention operation is an operation that can be input by the user via touch on the payment device.
[0158] The user intention execution module 410 can be used to process the pending payment order in response to the user intention operation.
[0159] based on Figure 4 The embodiments of this specification also provide some specific implementations of the device, which will be described below.
[0160] Optional, Figure 4 The device may further include:
[0161] The sound acquisition module can be used to acquire sound data; the time interval between the acquisition time of the sound data and the acquisition time of the target image is less than or equal to a first threshold.
[0162] The speech recognition module can be used to perform speech recognition processing on the sound data to obtain the user's speech recognition result.
[0163] The second judgment module can be used to determine whether the user's voice recognition result meets the second automatic operation triggering condition, and obtain the second judgment result; the second automatic operation triggering condition is that the user's voice recognition result contains a preset keyword corresponding to the user's intention operation to be executed; the user's intention operation to be executed is an operation that can be performed by the user through touch input on the payment device for the pending payment order.
[0164] The first determination module 406 may include:
[0165] The first judgment submodule can be used to determine whether the face detection data meets the third automatic operation triggering condition if the second judgment result indicates that the user voice recognition result does not meet the second automatic operation triggering condition; the third automatic operation triggering condition may include: a first face depth condition, a first face image quality condition, and a selected face continuity and consistency condition.
[0166] The second judgment submodule can be used to determine whether the face detection data meets the fourth automatic operation triggering condition if the second judgment result indicates that the user voice recognition result meets the second automatic operation triggering condition; the fourth automatic operation triggering condition may include: the second face depth condition and the second face image quality condition.
[0167] Optional, Figure 4 The device may further include:
[0168] The face detection submodule can be used to perform face detection on the target image to obtain the predicted region of the user's face in the target image.
[0169] The depth detection submodule can be used to perform face depth detection processing on the predicted region of the user's face in the target image to obtain the user's face depth value; the user's face depth value can be used to characterize the distance between the user's face and the facial recognition payment device.
[0170] Correspondingly, the first judgment submodule can be used to determine whether the user's face depth value meets the first face depth condition, wherein the first face depth condition is that the user's face depth value is greater than or equal to the second threshold.
[0171] Correspondingly, the second judgment submodule can be used to determine whether the user's face depth value meets the second face depth condition, which is that the user's face depth value is greater than or equal to a third threshold.
[0172] Optionally, the depth detection submodule can be used to calculate the area ratio of the predicted region of the user's face in the target image; and determine the depth value of the user's face based on the area ratio.
[0173] Optionally, the face detection module 404 may include:
[0174] The face detection submodule can be used to perform face detection on the target image to obtain the predicted region of the user's face in the target image.
[0175] The image quality detection submodule can be used to perform face image quality detection processing on the predicted region of the user's face in the target image to obtain a user face image quality score; the user face image quality score can be used to characterize the accuracy of determining the user's intended operation based on the face image in the predicted region of the user's face.
[0176] Correspondingly, the first judgment submodule can be used to determine whether the quality score of the user's face image meets the first face image quality condition, wherein the first face image quality condition is that the quality score of the user's face image is greater than or equal to the fourth threshold.
[0177] Correspondingly, the second judgment submodule can be used to determine whether the user's face image quality score meets the second face image quality condition, which is that the user's face image quality score is greater than or equal to the fifth threshold.
[0178] Optionally, the image quality detection submodule can be specifically used for:
[0179] The face image in the predicted region of the user's face is input into the face image quality detection model to obtain the user face image quality score of the face image output by the face image quality detection model. The face image quality detection model is obtained by training a deep learning model in advance using face image samples carrying face image quality labels. The face image quality labels are label data determined based on at least one of the face image sample's sharpness information, face image sample's face pose information, and face image sample's cheating behavior information.
[0180] Optionally, the face detection module 404 may include:
[0181] The face detection submodule can be used to perform face detection on the target image to obtain the predicted region of the user's face in the target image.
[0182] The location information determination submodule can be used to determine whether the user's face is a selected face in the target image based on the location information of the predicted region of the user's face in the target image, and obtain a third determination result.
[0183] The face feature extraction submodule can be used to extract user face feature data from the predicted region of the user face if the third judgment result indicates that the user face is the selected face in the target image.
[0184] Correspondingly, the first judgment submodule can be used to determine whether the user's facial feature data meets the condition of continuous consistency of the selected face. The condition of continuous consistency of the selected face is that the user's facial feature data of the selected face is consistent in each image of the multi-frame continuous images. The multi-frame continuous images are images continuously collected by the face-scanning payment device with a number greater than or equal to a sixth threshold. The multi-frame continuous images may include the target image.
[0185] Optionally, the user's intended action may include: activating the facial recognition payment function.
[0186] The user intention execution module 410 can be used to respond to the wake-up face payment operation and display a face recognition interface; the face recognition interface can be used to assist the user in performing face recognition operations during the payment process of the pending payment order.
[0187] Correspondingly, Figure 4 The device may further include:
[0188] The display module can be used to display a payment confirmation interface that includes the account to be deducted and the amount to be paid; the account to be deducted is the user account determined by facial recognition of the user, and the amount to be paid is the transaction amount of the pending payment order at the payment device.
[0189] Correspondingly, the user's intended action may include: confirming payment; the confirming payment action may be used to instruct the deduction of the amount to be paid from the account to be deducted.
[0190] Correspondingly, the user intention execution module 410 can be specifically used to deduct the amount to be paid from the account to be deducted in response to the payment confirmation operation and for the pending payment order.
[0191] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0192] Figure 5 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a facial recognition payment device. Figure 5 As shown, device 500 may include:
[0193] At least one processor 510; and a memory 530 communicatively connected to the at least one processor; wherein the memory 530 stores instructions 520 executable by the at least one processor 510, the instructions being executed by the at least one processor 510 to enable the at least one processor 510 to:
[0194] Facial recognition payment devices collect target images.
[0195] The target image is processed to obtain face detection data.
[0196] Determine whether the face detection data meets the first automatic operation trigger condition, and obtain the first determination result.
[0197] If the first judgment result indicates that the face detection data meets the first automatic operation triggering condition, then the user's intended operation is determined based on the processing progress information of the pending payment orders at the payment device; the user's intended operation is an operation that can be input by the user via touch on the payment device.
[0198] In response to the user's requested action, the pending payment order is processed.
[0199] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for... Figure 5 As the device shown is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0200] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0201] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0202] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0203] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0204] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0205] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine that can be used to implement the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0206] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0207] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that can be used to implement a process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0208] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0209] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0210] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0211] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0212] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0213] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0214] The above description is merely an embodiment of this application and should not be construed as limiting the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A face payment method, comprising: a face payment device acquires a target image; detecting processing is performed on the target image to obtain face detection data; the face payment device acquires sound data; a time interval between a time when the sound data is acquired and a time when the target image is acquired is less than or equal to a first threshold; voice recognition processing is performed on the sound data to obtain a user voice recognition result; it is judged whether the user voice recognition result satisfies a second automatic operation triggering condition to obtain a second judgment result; the second automatic operation triggering condition is that a preset keyword corresponding to a to-be-executed user will operation is contained in the user voice recognition result; the to-be-executed user will operation is an operation input by a user for a to-be-processed payment order at the face payment device; if the second judgment result indicates that the user voice recognition result does not satisfy the second automatic operation triggering condition, it is judged whether the face detection data satisfies a third automatic operation triggering condition; if the second judgment result indicates that the user voice recognition result satisfies the second automatic operation triggering condition, it is judged whether the face detection data satisfies a fourth automatic operation triggering condition; the third automatic operation triggering condition is stricter than the fourth automatic operation triggering condition; if the face detection data satisfies the third automatic operation triggering condition or the face detection data satisfies the fourth automatic operation triggering condition, a user will operation is determined according to processing progress information of a to-be-processed payment order at the face payment device; the user will operation is an operation input by a user for the face payment device; in response to the user will operation, the to-be-processed payment order is processed.
2. The method of claim 1, before the determining a user will operation, further comprising: it is judged whether the face detection data satisfies a first automatic operation triggering condition, specifically comprising: if the second judgment result indicates that the user voice recognition result does not satisfy the second automatic operation triggering condition, it is judged whether the face detection data satisfies the third automatic operation triggering condition; the third automatic operation triggering condition includes a first face depth condition, a first face image quality condition, and a selected face continuous consistency condition; if the second judgment result indicates that the user voice recognition result satisfies the second automatic operation triggering condition, it is judged whether the face detection data satisfies the fourth automatic operation triggering condition; the fourth automatic operation triggering condition includes a second face depth condition and a second face image quality condition.
3. The method of claim 2, the detecting processing on the target image to obtain face detection data, specifically comprising: face detection is performed on the target image to obtain a predicted region of a user face in the target image; face depth detection processing is performed on the predicted region of the user face in the target image to obtain a user face depth value; the user face depth value is used to represent a distance between the user face and the face payment device; the judging whether the face detection data satisfies the third automatic operation triggering condition, specifically comprising: determining whether the user face depth value meets a second face depth condition, the second face depth condition being that the user face depth value is less than or equal to a third threshold value. The determining whether the face detection data meets the fourth automatic operation triggering condition specifically includes: determining whether the user face depth value meets a second face depth condition, the second face depth condition being that the user face depth value is less than or equal to a third threshold value.
4. The method of claim 3, wherein the face depth detection processing on the predicted region of the user face in the target image to obtain the user face depth value specifically includes: calculating an area proportion of the predicted region of the user face in the target image; determining the user face depth value according to the area proportion.
5. The method of claim 2, wherein the detection processing on the target image to obtain the face detection data specifically includes: performing face detection on the target image to obtain a predicted region of a user face in the target image; performing face image quality detection processing on the predicted region of the user face in the target image to obtain a user face image quality score; the user face image quality score is used to represent an accuracy degree of determining a user willing operation based on a face image in the predicted region of the user face; The determining whether the face detection data meets the third automatic operation triggering condition specifically includes: determining whether the user face image quality score meets a first face image quality condition, the first face image quality condition being that the user face image quality score is greater than or equal to a fourth threshold value; The determining whether the face detection data meets the fourth automatic operation triggering condition specifically includes: determining whether the user face image quality score meets a second face image quality condition, the second face image quality condition being that the user face image quality score is greater than or equal to a fifth threshold value.
6. The method of claim 5, wherein the face image quality detection processing on the predicted region of the user face in the target image to obtain the user face image quality score specifically includes: inputting a face image in the predicted region of the user face into a face image quality detection model to obtain a user face image quality score of the face image output by the face image quality detection model; the face image quality detection model is obtained by pre-training a deep learning model using face image samples carrying face image quality labels; the face image quality labels are label data determined according to at least one of clarity information of the face image samples, face posture information of the face image samples, and cheating behavior information of the face image samples.
7. The method of claim 2, wherein the detection processing on the target image to obtain the face detection data specifically includes: performing face detection on the target image to obtain a predicted region of a user face in the target image; determining whether the user face is a selected face in the target image according to position information of the predicted region of the user face in the target image to obtain a third determination result; if the third determination result indicates that the user face is a selected face in the target image, extracting user face feature data of the user face from a prediction region of the user face; the third automatic operation trigger condition is that the user face feature data of the selected face in each of a plurality of continuous images is consistent, the plurality of continuous images are images continuously collected by the face payment device and the number of the images is greater than or equal to a sixth threshold, and the plurality of continuous images include the target image. waking up a face payment operation; 8. The method of any of claims 1-7, the user intent operation comprising: the processing of the to-be-processed payment order in response to the user intention operation specifically includes: in response to the wake-up face payment operation, displaying a face recognition interface; the face recognition interface is used to assist a user in performing a face recognition operation in a payment process for the to-be-processed payment order.
9. The method of any one of claims 1-7, before the face payment device collects the target image, further comprising: displaying a payment confirmation interface including a to-be-deducted account and a to-be-paid amount, the to-be-deducted account being a user account determined by performing a face recognition operation on the user, and the to-be-paid amount being a transaction amount of a to-be-processed payment order at the face payment device; the user intention operation includes a confirmation payment operation, and the confirmation payment operation is used to instruct to deduct the to-be-paid amount from the to-be-deducted account; the processing of the to-be-processed payment order in response to the user intention operation specifically includes: in response to the confirmation payment operation, deducting the to-be-paid amount from the to-be-deducted account for the to-be-processed payment order.
10. A face payment device, comprising: a graph collection module configured to collect a target image by a face payment device; a face detection module configured to perform a detection operation on the target image to obtain face detection data; a sound collection module configured to collect sound data, wherein a time interval between a collection time of the sound data and a collection time of the target image is less than or equal to a first threshold; a voice recognition module configured to perform a voice recognition operation on the sound data to obtain a user voice recognition result; a second determination module configured to determine whether the user voice recognition result meets a second automatic operation trigger condition to obtain a second determination result, wherein the second automatic operation trigger condition is that the user voice recognition result includes a preset keyword corresponding to a to-be-executed user intention operation; the to-be-executed user intention operation is an operation input by a user for a to-be-processed payment order at the face payment device; a first determination submodule configured to determine whether the face detection data meets a third automatic operation trigger condition if the second determination result indicates that the user voice recognition result does not meet the second automatic operation trigger condition. The second determining submodule is configured to determine whether the face detection data satisfies a fourth automatic operation triggering condition if the second determining result indicates that the user voice recognition result satisfies the second automatic operation triggering condition. The user intention determining module is configured to determine a user intention operation according to processing progress information of a to-be-processed payment order at the face payment device if the face detection data satisfies a third automatic operation triggering condition or the face detection data satisfies a fourth automatic operation triggering condition, wherein the user intention operation is an operation input by the user for the face payment device. The user intention executing module is configured to process the to-be-processed payment order in response to the user intention operation.
11. The apparatus of claim 10, further comprising: The first determining module comprises: The first determining submodule is configured to determine whether the face detection data satisfies a third automatic operation triggering condition if the second determining result indicates that the user voice recognition result does not satisfy the second automatic operation triggering condition, wherein the third automatic operation triggering condition comprises a first face depth condition, a first face image quality condition, and a selected face consistency condition. The second determining submodule is configured to determine whether the face detection data satisfies a fourth automatic operation triggering condition if the second determining result indicates that the user voice recognition result satisfies the second automatic operation triggering condition, wherein the fourth automatic operation triggering condition comprises a second face depth condition and a second face image quality condition.
12. The apparatus of claim 11, wherein the face detection module comprises: The face detection submodule is configured to perform face detection on the target image to obtain a predicted region of a user face in the target image. The depth detection submodule is configured to perform face depth detection processing on the predicted region of the user face in the target image to obtain a user face depth value, wherein the user face depth value is used to represent a distance between the user face and the face payment device. The first determining submodule is specifically configured to determine whether the user face depth value satisfies a first face depth condition, wherein the first face depth condition is that the user face depth value is less than or equal to a second threshold value. The second determining submodule is specifically configured to determine whether the user face depth value satisfies a second face depth condition, wherein the second face depth condition is that the user face depth value is less than or equal to a third threshold value.
13. The apparatus of claim 12, wherein the depth detection submodule is specifically configured to: calculate an area proportion of the predicted region of the user face in the target image; and determine the user face depth value according to the area proportion.
14. The apparatus of claim 11, wherein the face detection module comprises: The face detection submodule is configured to perform face detection on the target image to obtain a predicted region of a user face in the target image. an image quality detection submodule, configured to perform face image quality detection processing on the predicted region of the user face in the target image, to obtain a user face image quality score; the user face image quality score is used to represent an accuracy degree of determining a user intention operation based on a face image in the predicted region of the user face; the first judgment submodule is specifically configured to judge whether the user face image quality score meets a first face image quality condition, the first face image quality condition being that the user face image quality score is greater than or equal to a fourth threshold value; the second judgment submodule is specifically configured to judge whether the user face image quality score meets a second face image quality condition, the second face image quality condition being that the user face image quality score is greater than or equal to a fifth threshold value.
15. The apparatus of claim 14, wherein the image quality detection submodule is specifically configured to: input the face image in the predicted region of the user face into a face image quality detection model, to obtain a user face image quality score of the face image output by the face image quality detection model; the face image quality detection model is obtained by pre-training a deep learning model using face image samples carrying face image quality labels, the face image quality label being label data determined according to at least one of clarity information of the face image sample, face posture information of the face image sample, and cheating behavior information of the face image sample.
16. The apparatus of claim 11, wherein the face detection module comprises: a face detection submodule, configured to perform face detection on the target image, to obtain a predicted region of a user face in the target image; a position information judgment submodule, configured to judge whether the user face is a selected face in the target image according to position information of the predicted region of the user face in the target image, to obtain a third judgment result; a face feature extraction submodule, configured to, if the third judgment result indicates that the user face is a selected face in the target image, extract user face feature data of the user face from the predicted region of the user face; the first judgment submodule is specifically configured to: judge whether the user face feature data meets a selected face continuous consistency condition, the selected face continuous consistency condition being that user face feature data of a selected face in each image of a plurality of continuous images is consistent, the plurality of continuous images being images continuously collected by the face payment device, the number of the images being greater than or equal to a sixth threshold value, and the plurality of continuous images including the target image.
17. The apparatus of any of claims 10-16, the user intent operation comprising: wake up a face payment operation; the user intention execution module is specifically configured to, in response to the wake-up face payment operation, display a face recognition interface; the face recognition interface is used to assist a user to perform a face recognition operation in a payment process for the to-be-processed payment order.
18. The apparatus of any one of claims 10-16, further comprising: The display module is configured to display a payment confirmation interface including a to-be-deducted account and a to-be-paid amount, the to-be-deducted account being a user account determined by performing face recognition on the user, and the to-be-paid amount being a transaction amount of a to-be-processed payment order at the face payment device; The user intention operation includes a confirmation payment operation, and the confirmation payment operation is configured to instruct to deduct the to-be-paid amount from the to-be-deducted account; The user intention execution module is specifically configured to, in response to the confirmation payment operation, deduct the to-be-paid amount from the to-be-deducted account for the to-be-processed payment order.
19. A face payment device, comprising: at least one processor; and, a memory in communication connection with the at least one processor; wherein the memory stores instructions executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: the face payment device acquires a target image; detects the target image to obtain face detection data; the face payment device acquires sound data; the time interval between the acquisition time of the sound data and the acquisition time of the target image is less than or equal to a first threshold value; performs speech recognition processing on the sound data to obtain a user speech recognition result; determine whether the user speech recognition result meets a second automatic operation triggering condition to obtain a second determination result; the second automatic operation triggering condition is that the user speech recognition result contains a preset keyword corresponding to a to-be-executed user intention operation; the to-be-executed user intention operation is an operation input by the user for a to-be-processed payment order at the face payment device; if the second determination result indicates that the user speech recognition result does not meet the second automatic operation triggering condition, determine whether the face detection data meets a third automatic operation triggering condition; if the second determination result indicates that the user speech recognition result meets the second automatic operation triggering condition, determine whether the face detection data meets a fourth automatic operation triggering condition; the third automatic operation triggering condition is stricter than the fourth automatic operation triggering condition; if the face detection data meets the third automatic operation triggering condition, or the face detection data meets the fourth automatic operation triggering condition, determine a user intention operation according to the processing progress information of the to-be-processed payment order at the face payment device; the user intention operation is an operation input by the user for the face payment device; in response to the user intention operation, process the to-be-processed payment order.
Citation Information
Patent Citations
Face-scanning payment method and device and electronic equipment
CN111292092A
Face-scanning payment method, device and equipment
CN111553706A
Fee payment method of face recognition terminal, control device and terminal
CN111915310A
Payment processing method and device
CN113240428A
Payment mode recommendation processing method, device, equipment and system
CN113947400A