Biometric payment processing method, device, electronic device, and computer program

The biometric payment method enhances contactless payment accuracy and security by detecting the movement speed of a target part in successive images to confirm payment intent, addressing the challenges of misidentification and device risks in existing systems.

JP2025529840AActive Publication Date: 2025-09-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025510371
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-02
Filing Date
2023-09-27
Publication Date
2025-09-09
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

Biometric payment technologies face challenges in accurately identifying the intention to make contactless payments, particularly due to the risk of misidentification without a direct signal indicating payment intent, and there is a risk of device damage or contamination in contact-based methods.

Method used

A biometric payment processing method that involves acquiring images of a living organism, detecting a target part linked to the payment function, determining the movement speed of the target part in successive images, and executing a payment operation if the speed is below a threshold, thereby confirming the user's intent to pay.

Benefits of technology

This approach improves the accuracy of identifying payment intent in contactless transactions and enhances security by ensuring that the user's intention to pay is clearly indicated through controlled movement, reducing errors and risks associated with misidentification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529840000001_ABST
    Figure 2025529840000001_ABST
Patent Text Reader

Abstract

The present application provides a biometric payment processing method, device, electronic device, computer-readable storage medium, and computer program product, the biometric payment processing method including acquiring image data obtained by collecting images of a living organism, performing a target site detection process on an image in the image data, where the target site is a site on the living organism linked to a biometric payment function, determining a movement speed corresponding to the target site in the multiple images in response to detecting the target site from multiple images collected consecutively, and performing a payment operation based on the target site in response to the movement speed being less than a speed threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application is based on and claims priority from a Chinese patent application bearing application number 202211362513.8 and filed on November 2, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the technical field of the Internet, and in particular to a biometric payment processing method, device, electronic device, computer-readable storage medium, and computer program product. [Background technology]

[0003] As biometric payment technology becomes more widespread, more and more users are beginning to use it (e.g., facial recognition payment, fingerprint payment, etc.). Depending on the payment method, biometric payment can be divided into contact payment and contactless payment, and contact payment identifies a trigger signal through contact. For example, such technical measures usually require the user to click to confirm, and require the device (e.g., facial recognition payment terminal) to provide an interactive screen. At the same time, because it is contact-based, there is a risk that the device screen may be damaged or dirty when used in public places.

[0004] Contactless payment can solve the above problems, but there is a risk of misidentification because there is no signal indicating the intention to make a payment (for example, an electrical signal generated by contact between a finger and a payment terminal) like in contactless payment.

[0005] The related art does not provide an effective solution for how to improve the identification of the intention to make contactless payments. Summary of the Invention [Problem to be solved by the invention]

[0006] Embodiments of the present application provide a biometric payment processing method, device, electronic device, computer-readable storage medium, and computer program product that can improve the accuracy of identifying the intention to make contactless payments and increase the security of payments. [Means for solving the problem]

[0007] The technical means of the embodiments of the present application are realized as follows.

[0008] An embodiment of the present application provides a biometric payment processing method executed by an electronic device, the biometric payment processing method including: acquiring image data obtained by collecting images of a living organism; performing a detection process for a target part, which is a part of the living organism linked to a biometric payment function, on an image in the image data; determining a movement speed corresponding to the target part in the multiple images in response to detecting the target part from the multiple images collected successively; and executing a payment operation based on the target part in response to the movement speed being less than a speed threshold.

[0009] An embodiment of the present application provides a biometric payment processing device including: an acquisition module configured to acquire image data obtained by collecting images of a living organism; a detection module configured to perform a detection process for a target part, which is a part of the living organism linked to a biometric payment function, on an image in the image data; a determination module configured to determine a movement speed corresponding to the target part in the plurality of images collected in response to detecting the target part from the plurality of images collected successively; and a payment module configured to execute a payment operation based on the target part in response to the movement speed being less than a speed threshold.

[0010] An embodiment of the present application provides an electronic device including a memory for storing executable commands and a processor for implementing a biometric payment method provided in an embodiment of the present application when the executable commands stored in the memory are executed.

[0011] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement a biometric payment processing method provided in an embodiment of the present application.

[0012] An embodiment of the present application provides a computer program or a computer program product including computer-executable instructions that, when executed by a processor, implements the biometric payment processing method provided in the embodiment of the present application. [Effects of the Invention]

[0013] The embodiments of the present application have the following beneficial effects:

[0014] When images of a living creature are collected and a target part of the creature is detected in the collected images, the intention to make a payment is confirmed according to the movement speed of the target part in the multiple images, and if the movement speed of the target part is detected to be lower than the speed threshold, i.e., if the living creature clearly stops, it can be assumed that the living creature currently has a clear intention to make a payment, and a payment operation based on the target part is performed. This improves the accuracy of identifying the intention to make a payment in contactless payments and improves the safety of payments. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is a schematic diagram of the architecture of a biometric payment processing system 100 provided in an embodiment of the present application. [Figure 2] FIG. 2 is a structural schematic diagram of an electronic device 500 provided in an embodiment of the present application. [Figure 3] FIG. 3 is a flowchart of a biometric payment processing method provided in an embodiment of the present application. [Figure 4] FIG. 4 is a flowchart of a biometric payment processing method provided in an embodiment of the present application. [Figure 5] FIG. 5 is a flowchart of a biometric payment processing method provided in an embodiment of the present application. [Figure 6]FIG. 6 is a schematic diagram of a palm authentication collection and payment function model provided in an embodiment of the present application. [Figure 7] FIG. 7 is a schematic diagram of a target detection system provided in an embodiment of the present application. [Figure 8] FIG. 8 is a schematic diagram of the principle of grid division provided in the embodiment of the present application. [Figure 9] FIG. 9 is a structural schematic diagram of the YOLO model provided in the examples of the present application. [Figure 10] FIG. 10 is a schematic diagram of the location of the palm box in an image of the palm provided in the examples of the present application. [Figure 11] FIG. 11 is a schematic diagram of an area image of the area where the palm is located, provided in an embodiment of the present application. [Figure 12] FIG. 12 is a schematic diagram of key points on the palm of the hand provided in the examples of the present application. [Figure 13] FIG. 13 is a structural schematic diagram of the keypoint detection model provided in the embodiment of the present application. [Figure 14] FIG. 14 is a schematic diagram of the principle of keypoint detection by calling the keypoint detection model provided in the embodiment of the present application. [Figure 15] FIG. 15 is a schematic diagram of the principle of calculating the key point movement speed of the palm provided in the embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0016] In order to clarify the objectives, technical means and advantages of the present application, the present application will be described in more detail below in conjunction with the accompanying drawings. The described embodiments are not considered as limitations on the present application. All other embodiments that can be obtained by those skilled in the art without any creative effort fall within the scope of protection of the present application.

[0017] In the following description, references to "some embodiments" refer to a subset of all possible embodiments, but "some embodiments" may be the same subset of all possible embodiments or different subsets, and may be combined with each other if not inconsistent.

[0018] In addition, in the embodiments of the present application, when data related to user information, etc. (e.g., user fingerprints, palm prints, etc.) is implemented in a specific product or technology, permission or consent from the user must be obtained, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant country or region.

[0019] In the following description, the related terms "first / second / ..." are used merely to distinguish between similar objects and do not represent a particular order of the objects. Note that "first / second / ..." can be used interchangeably to refer to a particular order or order of precedence, where permitted. Therefore, the embodiments of the present application described herein may be performed in an order other than that shown or described.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used herein are for the purpose of describing the examples of the present application only and are not intended to limit the present application.

[0021] Before describing the embodiments of the present application in more detail, the nouns and terms used in the embodiments of the present application will be explained. The following interpretations of the nouns and terms used in the embodiments of the present application will be applied.

[0022] 1) Respond: Used to express a condition or state on which an operation to be executed depends, and when the dependent condition or state is satisfied, the operation or operations to be executed may be in real time or may have a set lag. Unless otherwise specified, there is no restriction on the order in which multiple operations are executed.

[0023] 2) Biometric payment: refers to a method of authenticating a user's personal identity and approving a payment using biometric identification technology. Biometric identification technology is a close combination of computers and technological means such as optics, acoustics, biosensors, and biostatistics, which uses physiological characteristics unique to living organisms (e.g., fingerprints, irises, palm prints, etc.) and behavioral characteristics (e.g., handwriting, voice, walking style, etc.) to authenticate an individual's identity.

[0024] 3) Object Detection: Also known as object extraction, it usually refers to detecting the location and corresponding category of an object (e.g., a target of interest) in an image.

[0025] 4) Contactless payment: This refers to completing a payment operation in a situation where direct contact with a sensor is required. For example, the user needs to indicate their intention to pay by the electrical signal generated when their finger touches the payment terminal, for example, by pressing the payment confirmation button displayed on the payment terminal screen. When the payment terminal receives the user's operation to press the payment confirmation button, it will carry out the subsequent payment operation.

[0026] 5) Contactless payment: Completing a payment operation without direct contact with a sensor. For example, without direct contact with a sensor, characteristics that distinguish a living organism (e.g., a human body) from other living organisms, such as fingerprint characteristics, palm characteristics (including at least a palm print, and may also include internal structures such as veins, bones, and soft tissues), facial characteristics, etc., are identified, and the identity information of the living organism is authenticated based on the identified biometric characteristics, and the subsequent payment operation is performed after successful authentication.

[0027] The present embodiments provide a biometric payment processing method, device, electronic device, computer-readable storage medium, and computer program product that can improve the accuracy of identifying payment intention for contactless payment. Exemplary applications of the electronic device provided in the embodiments of the present application are described below. The electronic device provided in the embodiments of the present application may be implemented as a terminal device, a server, or a terminal device and a server in cooperation with each other.

[0028] Hereinafter, the biometric payment processing method provided in the embodiment of the present application will be described as an example that is implemented by a server and a terminal device.

[0029] Please refer to Fig. 1, which is a schematic architecture diagram of a biometric payment processing system 100 provided in an embodiment of the present application. To support applications that improve the accuracy of identifying payment intention for contactless payments, as shown in Fig. 1, the biometric payment processing system 100 includes a server 200, a network 300, and a terminal device 400, where the network 300 may be a local area network, a wide area network, or a combination of both.

[0030] In some embodiments, a client 410 runs on the terminal device 400 (e.g., a facial recognition payment terminal, a palm recognition payment terminal, etc.). The client 410 may be a dedicated payment client or a real-time communication client running in a payment client model. In response to a biometric payment command issued by the payment user to make an electronic payment, the client 410 calls an image sensor built into the terminal device 400 or an external image sensor to periodically collect images of the payment user and obtain image data. The terminal device 400 can then transmit the collected image data to the server 200 via the network 300. After receiving the image data transmitted from the terminal device 400, the server 200 can perform a target region detection process (i.e., a region associated with the payment user's biometric payment function, e.g., the payment user's palm) on the image in the image data. If the payment user's palm is detected in multiple consecutively collected images, the server 200 can further determine the corresponding movement speed of the payment user's palm in the multiple images. When server 200 determines that the moving speed of the payer's palm is lower than the speed threshold (in this case, the payer is deemed to have a clear intention to pay), it can send a payment notification to terminal device 400, causing terminal device 400 to perform a payment operation based on the target location. For example, terminal device 400 can notify a store cash register system connected to terminal device 400 to perform an appropriate debit operation on the payer's account.

[0031] In some other embodiments, the biometric payment processing method provided in the embodiments of the present application may be performed independently by a terminal device. For example, taking the terminal device 400 shown in FIG. 1 as an example, after acquiring image data from an image collection of a payment person, the terminal device 400 uses its own computing power to detect a target portion in the image data and calculate the corresponding movement speed of the target portion in the multiple images. After that, if the terminal device 400 determines that the movement speed of the target portion is lower than the speed threshold, it can notify the connected store cash register system to perform a corresponding debit operation on the payment person's account.

[0032] In some embodiments, the embodiments of the present application may be implemented using cloud technology, which is a type of hosting technology that integrates a set of resources, such as hardware, software, and networks, within a wide area network or a local area network to realize data calculation, storage, processing, and sharing.

[0033] Cloud technology is a collective term for network technology, information technology, integration technology, management platform technology, and application technology based on the application of the cloud computing business model. It creates a resource pool that can be used as needed, providing flexibility and convenience. Cloud computing technology is an important pillar. The background services of technical network systems require large amounts of computing and storage resources.

[0034] 1 may be a standalone physical server, a server cluster, or a distributed system consisting of multiple physical servers. It may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain services, security services, content delivery networks (CDNs), big data, and artificial intelligence platforms. The terminal device 400 may be, but is not limited to, a smartphone, tablet computer, notebook computer, desktop computer, smart speaker, smart watch, face recognition payment terminal, palm recognition payment terminal, etc. The terminal device 400 and the server 200 may be directly or indirectly connected via wired or wireless communication, and this is not limited in the embodiments of the present application.

[0035] In some other embodiments, the terminal device 400 may further execute a computer program to implement the biometric payment processing method provided in the embodiments of the present application. For example, the computer program may be a native program or software module in an operating system, and may be the client 410. The client may be a native application (APP), i.e., a program that can be executed by being installed in an operating system, such as a payment APP. The client may also be an applet, i.e., a program that can be executed simply by downloading it to a browser environment. Furthermore, the client may be a payment applet that can be incorporated into any APP. In short, the computer program may be any type of application program, module, or plug-in.

[0036] In some embodiments, the biometric payment processing method provided in the embodiments of the present application may be implemented in combination with blockchain technology. Blockchain is a new application model of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanisms, and encryption algorithms. A blockchain is essentially a distributed database, a series of data blocks associated and generated using encryption methods, where each data block contains information about a batch of network transactions and is used to authenticate the validity of the information (e.g., prevent counterfeiting) and generate the next block. A blockchain may include a blockchain bottom-layer platform, a platform product service layer, and an application service layer.

[0037] For example, the terminal device detects a target part (e.g., palm) of a living being (e.g., user A) from multiple consecutively collected images, then matches the palmprint features of the detected palm with multiple authorized palmprint features stored in the blockchain, and determines the identity information corresponding to the authorized palmprint features as the current identity information of the living being. In this way, the security of payment can be further improved based on the tamper-proof property of the blockchain.

[0038] The structure of the electronic device provided in the embodiment of the present application will be further described below. Taking the electronic device as a terminal device as an example, reference is made to FIG. 2. FIG. 2 is a structural diagram of an electronic device 500 provided in the embodiment of the present application. The electronic device 500 shown in FIG. 2 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The components in the electronic device 500 are coupled via a bus system 540. It can be seen that the bus system 540 is used to realize communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, the various buses in FIG. 2 will be referred to as the bus system 540.

[0039] The processor 510 may be an integrated circuit chip having signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or a programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc., where the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0040] The user interface 530 includes one or more output devices 531 capable of presenting media content, including one or more speakers and / or one or more visual displays, and the user interface 530 further includes one or more input devices 532, including user interface members that support user input, such as a keyboard, a mouse, a microphone, a touch panel display, a camera, other input buttons, and widgets.

[0041] Memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Memory 550 optionally includes one or more storage devices that are physically remote from processor 510.

[0042] Memory 550 may include volatile or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). Memory 550 as described in embodiments herein is intended to include any suitable type of memory.

[0043] In some embodiments, memory 550 may store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as illustratively described below.

[0044] The operating system 551 includes system programs, such as a frame layer, a core library layer, a driver layer, etc., for processing various basic system services and executing hardware-related tasks, and is used to realize various basic services and process hardware-based tasks.

[0045] The network communications module 552 is used to connect to other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include Bluetooth, Wireless Compatibility (WiFi), and Universal Serial Bus (USB), among others.

[0046] The presentation module 553 is used to present information via one or more output devices 531 (e.g., displays, speakers, etc.) associated with the user interface 530 (e.g., the user interface for operating peripherals and displaying content and information).

[0047] The input processing module 554 is used to detect one or more user inputs or interactions from one of the one or more input devices 532 and to interpret the detected inputs or interactions.

[0048] In some embodiments, the devices provided in the embodiments of the present application may be implemented in a software manner. FIG. 2 shows a biometric payment processing device 555 stored in a memory 550. The biometric payment processing device 555 may be software in the form of a program, plug-in, or the like, and includes the following software modules: an acquisition module 5551, a detection module 5552, a determination module 5553, a payment module 5554, a cropping module 5555, an enlargement module 5556, and a division module 5557. These modules are logical and may be arbitrarily combined or further divided according to the functions to be implemented. Note that FIG. 2 only shows all of the above modules at once for clarity, and this does not exclude an implementation in which the biometric payment processing device 555 includes only the acquisition module 5551, the detection module 5552, the determination module 5553, and the payment module 5554. The functions of each module are described below.

[0049] Hereinafter, the biometric payment processing method provided in the embodiment of the present application will be specifically described in combination with exemplary applications and implementations of the terminal device provided in the embodiment of the present application.

[0050] Please refer to Fig. 3. Fig. 3 is a flowchart of the biometric payment processing method provided in the embodiment of the present application. The description will be made in conjunction with the steps shown in Fig. 3.

[0051] 3 may be executed by various types of computer programs that can be executed by terminal devices, and may be, for example, the above-mentioned operating systems, software modules, scripts, applets, etc., and are not limited to clients. Therefore, the following client examples are not considered to be limitations on the embodiments of the present application. Also, for ease of understanding, hereinafter, no specific distinction is made between terminal devices and clients executed by terminal devices.

[0052] In step 101, image data is acquired.

[0053] Here, the image data may be obtained by collecting images of a living being (for example, a human body, that is, a person making a payment, for example, user A who needs to perform a payment operation).

[0054] In some embodiments, when a terminal device (e.g., a facial recognition payment terminal or a palm recognition payment terminal) receives a payment request triggered by a payment person, it can call an image sensor built into the terminal device or an external image sensor to periodically collect images of the payment person and obtain image data.

[0055] For example, taking a palm authentication payment scenario as an example, when the palm authentication payment terminal receives a payment request triggered by a payment person, for example, when the palm authentication payment terminal detects using a distance sensor that the distance between the payment person (e.g., user A) and the palm authentication payment terminal is less than a distance threshold (e.g., 20 cm), the built-in image collection device (e.g., a camera) can periodically collect images of the payment person.

[0056] Taking a facial recognition payment scenario as an example, the facial recognition payment terminal can save resources by normally being in a sleep state (for example, the screen of the facial recognition payment terminal can be in a screen-off state). Also, when the built-in distance sensor of the facial recognition payment terminal detects that the distance between the payer (i.e., the user who needs to make a payment, for example, User A) and the facial recognition payment terminal is less than a distance threshold (for example, 20 cm), the facial recognition payment terminal automatically switches from a sleep state to an operating state and calls an internal (or external) image collection device (for example, a camera) to periodically collect images of User A who is close to the facial recognition payment terminal.

[0057] In step 102, a target portion detection process is performed on the image in the image data.

[0058] Here, the target part refers to a part of a living organism that is linked to the biometric payment function, and for example, types of target parts include the palm, fingers, wrist, face, etc. of the living organism.

[0059] For example, if the target site is the palm of a living being (i.e., the payment person), the payment person must first authenticate the palm before using palm authentication payment. For example, the palm authentication payment terminal first collects an image of the payment person's palm, extracts the payment person's palmprint features from the collected image, and stores the biometric features accepted by the payment person in a database corresponding to the palm authentication payment terminal for subsequent network authentication. Alternatively, the extracted palmprint features of the payment person can be stored locally in the palm authentication payment terminal for subsequent offline authentication.

[0060] In some embodiments, the terminal device can invoke a target detection model to perform target feature detection processing on images in the image data, where the target detection model is trained based on sample images and sample features labeled for the sample images.

[0061] A target detection model according to an embodiment of the present invention will be described below.

[0062] In some embodiments, target detection models can be classified into two types: one-stage target detection models (i.e., end-to-end target detection algorithms such as YOLO and SSD) and two-stage target detection models (i.e., target detection algorithms based on region proposals such as R-CNN, SPP-net, and Fast R-CNN). In two-stage target detection models, a series of candidate boxes are first generated as samples by an algorithm, and then sample classification is performed using a convolutional neural network. In contrast, one-stage target detection models do not require the generation of candidate boxes, and instead directly transform the target frame localization problem into a regression problem. The differences between the two methods result in differences in performance, with two-stage target detection models being superior in detection accuracy and localization accuracy, and one-stage target detection models being superior in algorithm speed.

[0063] Below, an example will be described in which a two-stage target detection model is called to perform target detection on an image.

[0064] For example, the terminal device can call the above target detection model in the following manner to perform target portion detection processing on an image in the image data: for the image in the image data, call the target detection model to determine multiple candidate boxes in the image and corresponding confidence scores for each candidate box, where the confidence scores are used to represent the probability that the candidate box contains the target portion; classify each candidate box based on the confidence scores to determine whether it contains the target portion; and perform regression processing on the target candidate boxes determined to contain the target portion to obtain modified positions of the target candidate boxes.

[0065] Below, we continue to explain the loss function in the training process of the target detection model.

[0066] Embodiments of the present application can train target detection models using a variety of different types of loss functions, including, for example, a regression loss function, a two-class classification loss function, a hinge loss, a multi-class classification loss function, and a multi-class classification cross-entropy loss.

[0067] Illustratively, multi-class classification cross-entropy loss is an extension of binary cross-entropy loss, where we consider an input vector X i and the corresponding one-hot encoded target vector Y i The losses for and are as follows:

[0068]

number

[0069] where p ij denotes the embedding representation of the coding features of the sample image, and y ij indicates the location of the labeled candidate boxes in the sample image based on the target detection model.

[0070] For example, the hinge loss is used in a support vector machine that mainly has class labels (e.g., including 1 and 0, where 1 represents success, i.e., the candidate box contains the target region, and 0 represents failure, i.e., the candidate box does not contain the target region). For example, the calculation formula for the hinge loss for data (x, y) is as follows:

[0071]

number

[0072] where y denotes the labeled sample region for the sample image, f(x) denotes the location of the labeled candidate box in the sample image based on the target detection model and the category of the object contained in the candidate box, and the hinge loss simplifies the mathematical operation of the support vector machine and maximizes the loss.

[0073] In addition, the embodiments of the present application can perform target portion detection processing on an image by calling any one of the above types of target detection models. For example, to improve payment efficiency and reduce the payment user's waiting time, a one-stage target detection model (i.e., an end-to-end target detection model) can be called to perform target portion detection processing on an image. On the other hand, to improve identification accuracy, a two-stage target detection model can be called to perform target portion detection processing on an image. The embodiments of the present application are not specifically limited thereto.

[0074] In some other embodiments, before performing target portion detection processing on the image in the image data, the terminal device may first further perform liveness detection processing and quality detection processing on the image in the image data (e.g., perform quality score detection processing on the image), and only after passing the liveness detection processing and the quality detection processing on the image in the image data, perform target portion detection processing on the image in the image data. In this way, by determining whether the detected image is derived from a real living organism (i.e., a living organism), it is possible to prevent illegal acquisition of identity information by showing a photo, mask, etc. in front of the terminal device and improve the security of payment.

[0075] In step 103, in response to detecting the target feature in the plurality of successively acquired images, a corresponding moving speed of the target feature in the plurality of images is determined.

[0076] In some embodiments, step 103 shown in Figure 3 is realized by steps 1031 to 1033 shown in Figure 4. The steps shown in Figure 4 will be described in combination.

[0077] In step 1031, in response to detecting a target region from a plurality of consecutively acquired images, a keypoint detection process is performed on the target region to obtain a plurality of keypoints included in the target region.

[0078] In some embodiments, the terminal device may perform keypoint detection processing on the target site by invoking a keypoint detection model to perform keypoint detection processing on the target site and obtain a plurality of keypoints contained in the target site, where the keypoint detection model is obtained by training based on the sample site of the sample organism and the keypoints labeled for the sample site.

[0079] In image processing, a keypoint is essentially a feature, an abstract description of a fixed region or physical relationship in space, describing a combination or context within a certain range. It not only represents a single point or a single location, but also the combinational relationship between the context and the surrounding region. Keypoint detection is primarily based on two methods: point regression (e.g., Coordinate) and point classification (e.g., Heatmap). Both methods can locate the location and relationship of points in an image. For example, Coordinate directly uses keypoint coordinates as the final goal that the network needs to regress, providing direct location information for each coordinate point. Heatmap, on the other hand, represents each type of coordinate as a probability graph, assigning a probability to the location of each pixel in the image and indicating the probability that the point belongs to the corresponding category of keypoint. Here, the closer a pixel is to the keypoint, the closer the probability is to 1; the farther a pixel is from the keypoint, the closer the probability is to 0.

[0080] For example, taking the point regression method as an example, the keypoint detection model may include a plurality of cascaded convolutional layers and a plurality of cascaded fully connected layers, and a first convolutional layer in the plurality of cascaded convolutional layers performs convolutional processing on feature information corresponding to the target region; the convolutional result output from the first convolutional layer is input to a subsequent cascaded convolutional layer, and the subsequent cascaded convolutional layers continue to perform convolutional processing up to the last convolutional layer; and a first fully connected layer in the plurality of cascaded fully connected layers: By performing fully-connected processing on the convolution results output from the last convolution layer, inputting the fully-connected results output from the first fully-connected layer into a subsequent cascaded fully-connected layer, and continuing the fully-connected processing through the subsequent cascaded fully-connected layer up to the last fully-connected layer, and determining the multiple points in the image to which the multiple coordinates output from the last fully-connected layer respectively correspond as multiple keypoints included in the target region, it is possible to call the above-mentioned keypoint detection model to perform keypoint detection processing on the target region and obtain the multiple keypoints included in the target region.

[0081] For example, if the target location is the palm of the payer's hand, as the payer's palm approaches the palm authentication payment terminal, the size of the palm box may be inaccurate at different distances due to image distortion caused by the angle, which will affect the subsequent calculation of the palm's movement speed. In light of this, embodiments of the present application call a keypoint detection model to perform keypoint detection processing on the area of ​​the image where the palm is located, and obtain several accurate palm keypoints that are relatively fixed on the palm. The palm's movement speed can then be calculated based on the palm keypoints, thereby avoiding the problem of the palm box size being inaccurate due to posture and large errors in the calculated movement speed.

[0082] In some other embodiments, before calling a keypoint detection model to perform keypoint detection processing on the target region, a region image of the region where the target region exists may be cut out from the image, and an enlargement process may be performed on the region image (e.g., the cut-out region image may be enlarged by a set multiple), and then the keypoint detection model may be called to further perform keypoint detection processing on the enlarged region image.

[0083] For example, if the target site is the palm of a living organism, in order to further improve the accuracy of prediction, before calling the keypoint detection model and performing keypoint detection processing on the palm, first cut out an area image of the area where the palm is located from the image, enlarge the cut-out image area by a set multiple, for example, enlarge the area image by three times, and then input the enlarged area image into the keypoint detection model, thereby making it possible to more accurately predict the location of the palm keypoints from the area image.

[0084] In step 1032, the movement speed corresponding to each keypoint in the multiple images is determined.

[0085] In some embodiments, step 1032 can be implemented by selecting, for each keypoint, a first image and a second image from the plurality of images, determining a time difference between the collection time of the first image and the collection time of the second image, determining a distance between a first coordinate, which is the coordinate of the keypoint in the first image, and a second coordinate, which is the coordinate of the keypoint in the second image, and determining a result of dividing the distance and the time difference as a corresponding movement speed of the keypoint in the plurality of images.

[0086] For example, the first and second images can be selected from the plurality of images by selecting the first and second images whose quality parameters (e.g., clarity, resolution, area occupied by the target region in the entire image, etc.) are greater than a quality parameter threshold. Here, the time difference between the acquisition time of the first image and the acquisition time of the second image is greater than a time difference threshold (e.g., 1 second). For example, the plurality of images can be sorted in descending order according to the quality parameter, and some images whose quality parameters are greater than the quality parameter threshold can be selected from the results sorted in descending order. Then, from the selected images, the image with the earliest acquisition time can be designated as the first image, and the image with the latest acquisition time can be designated as the second image. In this way, it is possible to avoid a situation in which the time difference between the acquisition times of two images is too small, resulting in a large error in calculating the movement speed of subsequent keypoints.

[0087] Of course, two images may be randomly selected from the plurality of images whose quality parameters are greater than the quality parameter threshold value to be used as the first and second images, provided that the time difference between the acquisition times of these two images is greater than the time difference threshold value. However, the embodiments of the present application are not specifically limited thereto.

[0088] For example, if the target part is the palm of a living creature, a keypoint detection model can be invoked to perform keypoint detection processing on the palm to obtain multiple keypoints contained in the palm. For example, assuming that the palm includes four keypoints, keypoint A, keypoint B, keypoint C, and keypoint D, keypoint A may be the boundary position between the thumb and index finger, keypoint B may be the boundary position between the index finger and middle finger, keypoint C may be the boundary position between the middle finger and ring finger, and keypoint D may be the boundary position between the ring finger and little finger. After obtaining the four keypoints contained in the palm, the corresponding movement speeds can be determined for each keypoint in multiple images (for example, assuming that the palm includes 10 images, image 1 to image 10).

[0089] Taking keypoint A among the four keypoints mentioned above as an example, first, multiple images are sorted in descending order according to their sharpness, and then some images whose sharpness is greater than a sharpness threshold are selected from the results of sorting in descending order (for example, assume that the results include four images: image 2, image 4, image 7, and image 10). The image collected at the earliest time among the selected images (for example, assume that image 2) is designated as the first image, and the image collected at the latest time (for example, assume that image 10) is designated as the second image. Next, the time difference between the collection time of image 10 and the collection time of image 2 is calculated, and the distance between the coordinates of keypoint A in image 10 (i.e., the second coordinates) and the coordinates of keypoint A in image 2 (i.e., the first coordinates) is calculated. Finally, the result of dividing the distance by the time difference is determined as the movement speed corresponding to keypoint A in the multiple images.

[0090] In addition, the calculation method of the movement speed corresponding to other key points (i.e., the above key points B, C, and D) in multiple images is similar to the calculation method of the movement speed of key point A, and since the calculation method of the movement speed of key point A can be referred to, it will not be described again in the embodiments of the present application.

[0091] In step 1033, the moving speeds corresponding to the target region in the multiple images are determined based on the moving speeds corresponding to the multiple key points, respectively.

[0092] In some embodiments, step 1033 can be implemented by determining an average movement speed of multiple movement speeds that correspond one-to-one to multiple key points, and determining the average movement speed as the movement speed that the target feature corresponds to in multiple images.

[0093] For example, if the target part is the palm of an animal, it is assumed that in step 1031, a keypoint detection model is called to perform keypoint detection processing on the palm, and then four keypoints, namely keypoint A, keypoint B, keypoint C, and keypoint D, contained in the palm are obtained. In step 1032, the movement speeds corresponding to these four keypoints in multiple images are calculated as V A , V B , V C , V D By calculating the average value of these four movement speeds (i.e., the average movement speed), the movement speed of the palm corresponding to the multiple images can be determined. In this way, the accuracy of the movement speed of the palm can be improved.

[0094] The terminal device may first authenticate the identity information of the living creature before determining the corresponding movement speed of the target part in the multiple images. Authentication methods for the identity information of the living creature include network authentication and offline authentication. Network authentication refers to the terminal device detecting the target part (e.g., palm) from the multiple images and then sending an identity identification request for the palm to the server, causing the server to extract palmprint features from the palm photo sent from the terminal device and match them with multiple authorized palmprint features stored in a database to determine the identity information corresponding to the authorized palmprint features in the matching. Offline authentication refers to the terminal device performing authentication locally. For example, the terminal device may detect the target part (e.g., palm) from the multiple images, extract palmprint features from the palm, and match them with multiple authorized palmprint features stored locally in the terminal device to determine the identity information corresponding to the authorized palmprint features in the matching as the identity information of the current living creature.

[0095] In step 104, in response to the moving speed being less than the speed threshold, a settlement operation is performed based on the target location.

[0096] In some embodiments, if the terminal device detects that the moving speed of the target location of the living being (i.e., the person making the payment) is less than the speed threshold, i.e., if the living being has a clear stopping motion, it can be assumed that the person making the payment now has a clear intention to make a payment, and a payment operation based on the target location can be performed, for example, by notifying the store cash register system connected to the terminal device to perform a corresponding debit operation on the person's account.

[0097] For example, if the target site is the palm of a living organism, when the palm authentication payment terminal detects that the movement speed of the payer's palm is lower than the speed threshold, i.e., when there is a clear stopping motion of the payer's palm, it indicates that the payer currently has a clear intention to make a payment, and the palm authentication payment terminal can notify the connected retail store cash register system to perform a debit operation on the payer's account. In this way, the accuracy of identifying palm authentication payments is improved, authentication errors caused by the payer's palm simply passing by can be avoided, and payment security can be improved.

[0098] In some other embodiments, to further improve the accuracy of identifying contactless payments, the terminal device can further perform the following processes: dividing the multiple images into multiple image groups according to a set frame interval; determining the movement speed corresponding to the target portion in each image group; and performing a payment operation based on the target portion in response to the multiple movement speeds corresponding to the target portion in the multiple image groups being all smaller than a speed threshold.

[0099] For example, if the target part is the palm of a living organism, in order to further improve the accuracy of identifying palm authentication payments, the palm authentication payment terminal may divide multiple images (i.e., multiple images including the palm) into multiple image groups according to a set frame interval (e.g., 5 frames), determine the corresponding movement speed of the palm in each image group, and perform a palm-based payment operation if the movement speeds of the palm in each image group are all less than a speed threshold.

[0100] In the embodiments of the present application, the width of the frame intervals dividing the multiple images can be set according to the actual situation. For example, to save terminal device resources, the width of the frame intervals can be set somewhat large, for example, 10 frames. On the other hand, to further improve the accuracy of calculating the moving speed of the target portion, the width of the frame intervals can be set somewhat small, for example, 5 frames. This is not specifically limited in the embodiments of the present application.

[0101] In some embodiments, following the example above, if the terminal device detects that the movement speed of any one image group of the target part is greater than a speed threshold, it can cancel the payment operation based on the target part. For example, if the target part is the palm of a living organism, if the palm authentication payment terminal detects that the movement speed of any one image group of the payment person's palm is greater than the speed threshold, it is possible that the payment person's palm simply passed through the palm authentication payment terminal and does not actually intend to make a palm authentication payment. By canceling the palm-based payment operation, it is possible to avoid authentication errors and improve payment security.

[0102] In some other embodiments, reference is made to Fig. 5. Fig. 5 is a flowchart of a biometric payment processing method provided in an embodiment of the present application. As shown in Fig. 5, after performing step 103 shown in Fig. 3, step 105 shown in Fig. 5 may be further performed. Step 105 shown in Fig. 5 will be described in combination.

[0103] In step 105, in response to the moving speed being greater than the speed threshold, the settlement operation based on the target location is cancelled.

[0104] In some embodiments, if the terminal device detects that the moving speed of the target portion is greater than the speed threshold, it is considered that the living being (i.e., the person making the payment) does not currently have a clear intention to make a payment, and the person's target portion may have simply passed through the terminal device, and the payment operation based on the target portion can be canceled to avoid authentication errors and ensure the security of the payment.

[0105] The biometric payment processing method provided in the embodiments of the present application collects images of a living organism, and when a target part of the organism is detected in the collected multiple images, the intention to make a payment is confirmed according to the corresponding movement speed of the target part in the multiple images.If it is detected that the movement speed of the target part is lower than a speed threshold, it is considered that the organism has a clear intention to make a payment, and for the first time a payment operation based on the target part can be performed, thereby improving the accuracy of identifying the intention to make a contactless payment and increasing the security of the payment.

[0106] Hereinafter, an exemplary application of the embodiment of the present application in one practical application scenario will be described using palm authentication payment as an example.

[0107] In related art, the user usually confirms the intention to make a payment by clicking a payment confirmation button, but this method requires the user to click confirmation, so it is necessary to provide an interactive screen for the payment device. Furthermore, since this method involves contact, there is a problem that the screen of the payment device in a public place may be damaged or stained by public operation.

[0108] In view of this, an embodiment of the present application provides a biometric payment processing method, which detects key points on the user's palm, predicts the movement speed of the key points on the palm in the XY plane, and determines the stopping point of the palm movement by judging the speed change curve, thereby confirming the user's intention to make a payment. That is, the method provided in the embodiment of the present application requires the user to confirm the user's intention to make a payment during the payment process to avoid authentication errors, that is, ensure that the user's palm shows a relatively clear intention to stop during the payment process, so as to avoid being identified and paid by casually passing by the user's palm.

[0109] The biometric payment processing method provided in the embodiment of the present application will be specifically described below.

[0110] For example, refer to FIG. 6. FIG. 6 is a schematic diagram of a palm authentication collection and payment function model provided in an embodiment of the present application. As shown in FIG. 6, the palm authentication payment processing method provided in an embodiment of the present application includes two stages: collection and payment. Here, in the collection stage, a user extends his / her hand to collect palm prints, and a palm authentication payment device (i.e., palm authentication device, hereinafter referred to as palm authentication device) stores the collected palm print in a database as the user's authorized biometric characteristics. In the payment stage, the palm authentication device reads the user's palm print and then matches the read palm print with the palm print stored in the database. If the two match, it indicates that the user's identity information has been successfully authenticated, and the subsequent payment operation can be performed. It can also be seen from FIG. 6 that the user extends his / her hand to collect palm prints, preferably biometric data, and then identification is performed, but the user does not need to perform any contactless operations throughout the entire process.

[0111] The biometric payment processing method provided in the embodiments of the present application predicts the movement speed of the user's palm by detecting the position of the user's palm and, more accurately, changes in the position of palm key points. If the movement speed of the user's palm is lower than a set speed threshold within any consecutive period, it is assumed that the user has a clear intention to make a payment, rather than simply having their palm pass by and interfere, thereby realizing contactless confirmation of the user's intention to make a payment.

[0112] In some embodiments, the biometric payment processing method provided in the embodiments of the present application mainly includes three steps: "1. Palm detection", "2. Palm keypoint detection", and "3. Palm movement speed calculation". Each step will be described in detail below.

[0113] 1. Palm detection In some embodiments, palm position detection can be performed using a target detection system, such as a YOLO model-based detection system, as shown in Figure 7. During the training phase of the YOLO model, sample data containing palm box position information can be obtained through pre-orientation, and then sent to a YOLO model waiting to be trained to obtain a trained YOLO model. The palm authentication terminal then performs real-time inference using the trained YOLO model, ultimately obtaining a palm detection box in the image.

[0114] For example, as shown in FIG. 7, after invoking the YOLO model to perform a target detection process on image 700, a candidate box 704 surrounding target 701 (e.g., a human body) and a corresponding confidence score for candidate box 704, a candidate box 705 surrounding target 702 (e.g., a small dog) and a corresponding confidence score for candidate box 705, and a candidate box 706 surrounding target 703 (e.g., a large dog) and a corresponding confidence score for candidate box 706 can be labeled in image 100.

[0115] The YOLO model provided in the embodiments of the present application will be described below.

[0116] In some embodiments, as shown in FIG. 7, the YOLO model first divides the input image into S*S grids. Then, for each grid, B candidate boxes (or bounding boxes) are predicted. Each candidate box contains five predicted values: x, y, w, h, and confidence. Here, x and y are the center coordinates of the candidate box, aligned with the grid cell (i.e., offset value relative to the current grid cell), and range from 0 to 1. w and h are the width and height of the candidate box, respectively. w and h may be normalized, for example, by dividing them by the width and height of the image, respectively. In this way, the final w and h are also within the range of 0 to 1. Furthermore, as shown in FIG. 8, each grid can predict the probabilities of C hypothetical categories. For example, if S=7, B=2, and C=20 (i.e., assuming there are 20 categories), the final tensor may be 7*7*30.

[0117] For example, as shown in FIG. 8, to detect whether a target 801 (e.g., a small dog) exists in an image 800, the image 800 is first divided into 49 grids, then the probability that each grid contains the target 801 is predicted, and finally a regression process is performed to finally obtain a candidate box 802 containing the target 801.

[0118] Each candidate box corresponds to a confidence score, and if there is no object in the grid cell, the corresponding confidence score is 0; if there is an object, the corresponding confidence score may be equal to the IOU (Intersection Over Union) value between the predicted box and the ground truth (i.e., the actual box).

[0119] The structure of the YOLO model provided in the examples of the present application will be described below.

[0120] For example, refer to FIG. 9. FIG. 9 is a structural schematic diagram of the YOLO model provided in the embodiment of the present application. As shown in FIG. 9, the YOLO model provided in the embodiment of the present application may include multiple convolutional layers 901 (e.g., six convolutional layers) and multiple fully connected layers 902 (e.g., two fully connected layers). Here, the convolutional layers 901 are mainly used to extract features, and the fully connected layers 902 are mainly used to predict category probabilities and coordinates. The final output of the YOLO model is a 7*7*30 tensor, where 7*7 is the number of grid cells.

[0121] In some embodiments, as shown in Figure 10, the palm authentication terminal can obtain the position of the palm box 1001 in the palm image 1000 after calling the above YOLO model to perform detection on the image 1000. Here, x and y are the pixel position of the upper left corner of the palm box 1001, and w and h are the width and height of the palm box 1001, respectively.

[0122] 2. Palm keypoint detection In some embodiments, as shown in FIG. 11, after detecting the image using the YOLO model, the palm recognition terminal can obtain an area image 1100 of the area where the user's palm is located, thereby effectively avoiding interference from unrelated content.

[0123] However, when the user's palm approaches the palm authentication terminal, the size of the palm box will be inaccurate at different distances due to image distortion caused by angles, etc. Based on this, the embodiment of the present application invokes a keypoint detection model (e.g., a palm keypoint detection algorithm) to obtain several accurate palm keypoint positions that are relatively fixed on the palm, thereby avoiding the problem of inaccurate palm box size caused by posture interference.

[0124] For example, refer to Fig. 12. Fig. 12 is a schematic diagram of palm keypoints provided in an embodiment of the present application. As shown in Fig. 12, the palm authentication terminal calls the keypoint detection model to perform keypoint detection on the area image of the palm area shown in Fig. 11, and then can obtain multiple palm keypoints included in the user's palm, such as palm keypoint 1201, palm keypoint 1202, palm keypoint 1203, and palm keypoint 1204.

[0125] The keypoint detection model provided in the embodiment of the present application will be described below.

[0126] In some embodiments, palm keypoints may be predicted using the DeepPose model. The main idea of ​​the DeepPose model is to transform the keypoint detection algorithm into a pure learning prediction problem without considering the complex anatomical issues of human body postures. By artificially labeling a large amount of human body keypoint data in various postures and using a deep neural network (DNN) to learn from the sample data, a more universal end-to-end keypoint detection algorithm can be realized.

[0127] For example, please refer to Fig. 13. Fig. 13 is a structural diagram of a keypoint detection model provided in an embodiment of the present application. As shown in Fig. 13, after a series of convolution processes are performed on an input image 1300 with a size of 220*220, the coordinates (x i ,y i ) can be obtained.

[0128] Furthermore, since the size of the input target is uncertain and the keypoint detection model itself only accepts input images of size 220*220, scaling of an image that is too large may result in errors in the prediction of the final target position. In view of this, as shown in Figure 14, the final prediction accuracy can be improved by cropping the area around the predicted point position in image 1400, enlarging the cropped image area 1401, and then continuing to make even more accurate predictions.

[0129] It should be noted that the palm keypoint prediction itself can be simplified into a single regression problem, which will not be described again in the embodiments of this application. The core idea is to achieve end-to-end return network prediction results by matching nonlinear regression problems based on DNN networks.

[0130] 3. Calculation of palm movement speed In some embodiments, a user's intention to make a payment may be determined according to the moving speed of the palm in the XY plane within a certain period of time. For example, if the palm authentication terminal determines that the moving speed of the user's palm is less than a set speed threshold, the payment is confirmed; otherwise, the payment is canceled. For example, during the payment process, if the moving speed of the user's palm in the XY plane (i.e., horizontal plane) within a set time (e.g., 3 seconds, i.e., 75 frames) is all less than a speed threshold (temporarily designated as α), the palm authentication terminal can confirm the payment.

[0131] For example, refer to FIG. 15. FIG. 15 is a schematic diagram of the principle of calculating the movement speed of palm keypoints provided in an embodiment of the present application. As shown in FIG. 15, the movement speed of the palm can be estimated using the movement speed of the palm keypoints. For example, by acquiring palm photos taken by the palm authentication terminal for five frames at a time (i.e., every 200 milliseconds) from 0 to 3 seconds, the movement speed of the palm in each of the five frames can be calculated. For example, taking palm keypoint 1501 shown in FIG. 15 as an example, if the coordinates in palm photo 1500 collected at 0 seconds for palm keypoint 1501 on the XY plane are assumed to be (X1, Y1) and the coordinates in palm photo 1505 collected at 200 milliseconds for palm keypoint 1501 are assumed to be (X2, Y2), then the movement speed V of palm keypoint 1501 from 0 seconds to 200 milliseconds can be calculated as follows: A becomes as follows:

[0132]

number

[0133] Here, ΔH indicates the distance between the coordinates (X1, Y1) of palm keypoint 1501 in palm photograph 1500 and the coordinates (X2, Y2) in palm photograph 1505, and Δt indicates the time difference between the collection time of palm photograph 1505 and the collection time of palm photograph 1500 (i.e., 200 milliseconds).

[0134] Similarly, the movement speeds of the palm keypoints 1502, 1503, and 1504 shown in Fig. 15 within 200 milliseconds from 0 seconds are respectively expressed as V B , V C and V D Next, the average value of the movement speeds of these four palm key points may be calculated as the palm movement speed within 0 to 200 milliseconds. That is, the palm movement speed V1 within 0 to 200 milliseconds is calculated as follows:

[0135]

number

[0136] Then, the palm movement speed V1 is compared with the set speed threshold α, and if V1≦α, the palm authentication terminal can confirm the payment; if not, it indicates that the user's palm has simply passed through the palm authentication device, that is, the palm authentication device can cancel the payment.

[0137] For the same reason, the palm authentication device may calculate the corresponding movement speed for each of the five frames of the palm within 0 to 3 seconds as V1, V2, V3, V4, .... In order to avoid authentication errors, the palm authentication device will only confirm the payment if all of the above movement speeds are smaller than the set speed threshold α; otherwise, the palm authentication device can cancel the payment.

[0138] It should be noted that the above five frames are not fixed, and the frame interval may be flexibly adjusted based on the actual situation, or a fixed time interval may be used, which is not specifically limited in the embodiments of the present application.

[0139] In some other embodiments, the movement of the palm can be determined not only based on the movement speed of the palm keypoints in the XY plane (i.e., horizontal plane), but also by determining the speed in the Z direction (i.e., vertical direction), it can be determined whether the palm is in the process of approaching or has clearly stopped, thereby determining the user's intention to make a payment.

[0140] The biometric payment processing method provided in the embodiments of the present application detects palm keypoints during the process of a user extending their hand to make a palm authentication payment, predicts the movement speed of the palm keypoints on the XY plane, and determines the speed change curve to determine the stopping point of the palm movement and confirm the user's intention to make a payment. That is, it ensures that the user only confirms the payment when the user's palm shows a relatively clear intention to stop during the payment process, thereby avoiding authentication errors caused by the palm passing by casually and increasing the user's sense of security when making a palm authentication payment.

[0141]

[0047] The following continues to describe an exemplary structure of the biometric payment processing device 555 provided in the embodiments of the present application, which is implemented as a software module. In some embodiments, as shown in FIG. 2, the software modules in the biometric payment processing device 555 stored in the memory 550 may include an acquisition module 5551, a detection module 5552, a determination module 5553, and a payment module 5554.

[0142] The acquisition module 5551 is configured to acquire image data obtained by collecting images of a living organism. The detection module 5552 is configured to perform a target location detection process on an image in the image data, where the target location is a location on the living organism associated with a biometric payment function. The determination module 5553 is configured to determine a movement speed corresponding to the target location in the multiple images in response to detecting the target location from multiple images collected consecutively. The payment module 5554 is configured to perform a payment operation based on the target location in response to the movement speed being less than a speed threshold.

[0143] In some embodiments, the detection module 5552 is further configured to perform a keypoint detection process on the target region to obtain a plurality of keypoints included in the target region, and the determination module 5553 is further configured to determine a corresponding movement speed for each keypoint in the plurality of images, and determine a corresponding movement speed for the target region in the plurality of images based on the movement speeds corresponding to the plurality of keypoints, respectively.

[0144] In some embodiments, the determination module 5553 is further configured to perform the following processes for each keypoint: sorting a first image and a second image from the plurality of images; determining a time difference between the acquisition time of the first image and the acquisition time of the second image; determining a distance between a first coordinate, which is the coordinate of the keypoint in the first image, and a second coordinate, which is the coordinate of the keypoint in the second image; and determining a result of dividing the distance and the time difference as a corresponding movement speed of the keypoint in the plurality of images.

[0145] In some embodiments, the determination module 5553 is further configured to sort out a first image and a second image from the plurality of images, the quality parameter being greater than a quality parameter threshold, wherein a time difference between an acquisition time of the first image and an acquisition time of the second image is greater than a time difference threshold.

[0146] In some embodiments, the determination module 5553 is further configured to determine an average movement speed of the plurality of movement speeds that correspond one-to-one to the plurality of key points, and determine the average movement speed as the movement speed to which the target portion corresponds in the plurality of images.

[0147] In some embodiments, the detection module 5552 is further configured to invoke a keypoint detection model to perform a keypoint detection process on the target site to obtain a plurality of keypoints contained in the target site, where the keypoint detection model is obtained by training based on the sample site of the sample organism and the keypoints labeled for the sample site.

[0148] In some embodiments, the biometric payment processing device 555 further includes a cropping module 5555 and an enlargement module 5556. The cropping module 5555 is configured to crop a region image of a region where the target region exists from the image before the detection module 5552 invokes a keypoint detection model to perform keypoint detection processing on the target region. The enlargement module 5556 is configured to perform enlargement processing on the region image. The detection module 5552 is further configured to invoke a keypoint detection model to perform keypoint detection processing on the enlarged region image.

[0149] In some embodiments, the keypoint detection model includes a plurality of cascaded convolution layers and a plurality of cascaded fully connected layers, and the detection module 5552 further includes: a first convolution layer in the plurality of cascaded convolution layers performs a convolution process on the feature information corresponding to the target region; the convolution result output from the first convolution layer is input to a subsequent cascaded convolution layer; and the subsequent cascaded convolution layers continue to perform convolution processes up to the last convolution layer. The first fully connected layer of the plurality of cascaded fully connected layers performs fully connected processing on the convolution result output from the last convolution layer, and the fully connected result output from the first fully connected layer is input to a subsequent cascaded fully connected layer, and the subsequent cascaded fully connected layers continue to perform fully connected processing up to the last fully connected layer, so that a plurality of points in the image corresponding to the plurality of coordinates output from the last fully connected layer are determined as a plurality of key points included in the target region.

[0150] In some embodiments, the biometric payment processing device 555 further includes a division module 5557 configured to divide the plurality of images into a plurality of image groups according to a set frame interval, a determination module 5553 configured to determine a movement speed corresponding to the target portion in each image group, and a payment module 5554 configured to perform a payment operation based on the target portion in response to the plurality of movement speeds corresponding to the target portion in the plurality of image groups being all less than a speed threshold.

[0151] In some embodiments, the settlement module 5554 is further configured to cancel a settlement operation based on the target site in response to the speed of movement of the target site in any one image group being greater than a speed threshold.

[0152] In some embodiments, the detection module 5552 is further configured to invoke a target detection model to perform target feature detection processing on images in the image data, where the target detection model is trained based on sample images and labeled sample features for the sample images.

[0153] In some embodiments, the detection module 5552 is further configured to: invoke a target detection model on an image in the image data to determine a plurality of candidate boxes in the image and a corresponding confidence score for each candidate box; classify each candidate box as containing or not containing a target feature based on the confidence scores; and perform a regression process on the target candidate boxes determined to contain the target feature to obtain revised positions of the target candidate boxes, where the confidence score is used to represent the probability that the candidate box contains the target feature.

[0154] In some embodiments, target site types include palm, fingers, wrist, and face.

[0155] The description of the device in the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects to the method embodiment, so the description will be omitted. Details of the technology not mentioned in the biometric payment processing device provided in the embodiment of the present application can be understood based on the description of any one of Figures 3 to 5.

[0156] An embodiment of the present application provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium, wherein a processor of a computing device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, thereby causing the computing device to perform the biometric payment processing method of the embodiment of the present application.

[0157] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, cause the processor to perform a biometric payment processing method provided in an embodiment of the present application, for example, the biometric payment processing method shown in any one of Figures 3 to 5.

[0158] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM, or may be any device including one or any combination of the above memories.

[0159] In some embodiments, the executable commands may be in the form of a program, software, software module, script, or code, organized according to any type of programming language (including compiled or interpreted, or declarative or procedural), and may be organized according to any type, including as a separate program, module, assembly, subroutine, or other unit suitable for use in a computing environment.

[0160] As an example, the executable commands may be configured to be executed by a single electronic device, or may be configured to be executed by multiple electronic devices located at a single location, or may be configured to be executed by multiple electronic devices distributed across multiple locations and interconnected via a communications network.

[0161] The above content is merely an embodiment of the present application and does not limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the scope of protection of the present application.

Claims

1. A biometric payment processing method executed by an electronic device, comprising: acquiring image data obtained by collecting images of the living organism; performing a detection process for a target part, which is a part of the living being linked to a biometric payment function, on an image in the image data; responsive to detecting the target feature from a plurality of the successively acquired images, determining a corresponding speed of movement of the target feature in the plurality of images; and executing a payment operation based on the target location in response to the moving speed being less than a speed threshold.

2. Determining the corresponding movement speed of the target portion in the plurality of images includes: performing a keypoint detection process on the target portion to obtain a plurality of keypoints included in the target portion; determining a corresponding movement speed for each of said key points in a plurality of said images; The biometric payment processing method according to claim 1 , further comprising determining corresponding movement speeds of the target portion in the plurality of images based on the movement speeds respectively corresponding to the plurality of key points.

3. Determining a corresponding movement speed for each of the key points in the plurality of images includes: For each of said keypoints, selecting a first image and a second image from the plurality of images and determining a time difference between an acquisition time of the first image and an acquisition time of the second image; determining a distance between first coordinates of the key point in the first image and second coordinates of the key point in the second image; The biometric payment processing method according to claim 2 , further comprising: determining a result of dividing the distance by the time difference as the movement speed corresponding to the key point in the plurality of images.

4. Selecting a first image and a second image from the plurality of images includes:

4. The biometric payment processing method of claim 3, further comprising: selecting a first image and a second image from the plurality of images, the first image and the second image each having a quality parameter greater than a quality parameter threshold; and a time difference between the capture time of the first image and the capture time of the second image being greater than a time difference threshold.

5. Determining corresponding movement speeds of the target portion in the plurality of images based on movement speeds corresponding to the plurality of key points, determining an average moving speed of a plurality of moving speeds corresponding one-to-one to the plurality of key points; The biometric payment processing method according to claim 2 , further comprising determining the average movement speed as the movement speed corresponding to the target portion in the plurality of images.

6. performing a keypoint detection process on the target portion to obtain a plurality of keypoints included in the target portion; 3. The biometric payment processing method of claim 2, further comprising: calling a keypoint detection model to perform a keypoint detection process on the target site to obtain a plurality of keypoints contained in the target site; and the keypoint detection model is obtained by training based on a sample site of a sample organism and keypoints labeled for the sample site.

7. before invoking the keypoint detection model to perform keypoint detection processing on the target region; Further, the method includes: cutting out a region image of a region where the target portion exists from the image; and performing an enlargement process on the region image; Invoking a keypoint detection model to perform keypoint detection processing on the target region includes: The biometric payment processing method according to claim 6 , further comprising calling a keypoint detection model and performing keypoint detection processing on the enlarged area image.

8. the keypoint detection model includes a plurality of cascaded convolutional layers and a plurality of cascaded fully connected layers; Invoking a keypoint detection model to perform a keypoint detection process on the target portion and acquiring a plurality of keypoints included in the target portion, performing a convolution process on feature information corresponding to the target region by a first convolution layer in the plurality of cascaded convolution layers; The convolution result output from the first convolution layer is input to a subsequent cascaded convolution layer, and convolution processing is continuously performed by the subsequent cascaded convolution layer up to the last convolution layer; performing a fully connected process on the convolution result output from the last convolution layer by a first fully connected layer in the plurality of cascaded fully connected layers; The fully connected result output from the first fully connected layer is input to a subsequent cascade-connected fully connected layer, and the subsequent cascade-connected fully connected layer continues to perform fully connected processing up to the final fully connected layer. and determining, as a plurality of key points included in the target area, a plurality of points in the image corresponding to the plurality of coordinates output from the final fully connected layer.

9. Determining the corresponding movement speed of the target portion in the plurality of images includes: Dividing the plurality of images into a plurality of image groups according to a set frame interval; determining a corresponding movement speed of the target portion in each of the image groups; Executing a settlement operation based on the target location in response to the moving speed being less than a speed threshold includes: The biometric payment processing method according to claim 1, further comprising: executing a payment operation based on the target portion in response to the movement speeds corresponding to the target portion in the image groups being all smaller than a speed threshold.

10. The biometric payment processing method of claim 9 , further comprising canceling a payment operation based on the target portion in response to the moving speed in the image group of any one of the target portions being greater than the speed threshold.

11. The target portion detection process is performed on the image in the image data.

2. The biometric payment processing method of claim 1, further comprising: calling a target detection model to perform target feature detection processing on images in the image data, wherein the target detection model is obtained by training based on sample images and sample features labeled for the sample images.

12. Calling a target detection model and performing a target portion detection process on an image in the image data includes: Invoking a target detection model for an image in the image data; determining a plurality of candidate boxes in the image and a confidence score corresponding to each of the candidate boxes, the confidence score being used to represent a probability that the candidate box contains the target feature; performing a classification process for each of the candidate boxes to determine whether or not it includes the target portion based on the confidence score; The biometric payment processing method according to claim 11, further comprising: performing a regression process on the target candidate box determined to include the target portion, and acquiring a corrected position of the target candidate box.

13. The biometric payment processing method according to claim 1 , wherein the target area types include palm, fingers, wrist, and face.

14. an acquisition module configured to acquire image data obtained by acquiring images of the living organism; A detection module configured to perform a detection process for a target part, which is a part of the living thing associated with a biometric payment function, on an image in the image data; a determination module configured to, in response to detecting the target feature in a plurality of the consecutively acquired images, determine a corresponding moving speed of the target feature in the plurality of the images; a payment module configured to execute a payment operation based on the target location in response to the moving speed being less than a speed threshold.

15. a memory for storing executable commands; and a processor for implementing the biometric payment processing method of any one of claims 1 to 13 when executing executable commands stored in the memory.

16. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the biometric payment processing method of any one of claims 1 to 13.

17. 14. A computer program product comprising a computer program or computer executable instructions which, when executed by a processor, implements the biometric payment processing method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Collation device and collation method

    JP2016035675A

  • Moving body detection device, moving body detection system and moving body detection method

    JP2018124689A

  • Target object motion recognition method, device, and electronic device

    JP2021531589A

  • Touchless fingerprint payment system

    US20180307886A1

  • Action recognition method and apparatus, and human-machine interaction method and apparatus

    US20210271892A1