Payment processing method and device based on facial recognition, equipment and medium

By performing feature analysis of audio data, the number of people in the environment is judged, and facial recognition payment is directly carried out in a single-person scenario, the problem of increasing time-consuming facial recognition payment system in complex environments is solved, and payment efficiency and user experience are improved.

CN120047146APending Publication Date: 2025-05-27SHENZHEN TENCENT COMP SYST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311607670.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In complex environments, the facial recognition payment system needs to load algorithm modules such as multi-person detection and attention mechanisms, resulting in increased operation time and reduced payment efficiency and business experience.

Method used

By performing feature analysis on the audio data, the number of people in the current environment is judged. If it is a single person, the facial recognition model of the single person scene is used for payment operations to avoid loading large-scale storage modules such as multi-person detection.

Benefits of technology

The load time-consuming of the facial payment process is optimized, payment efficiency and user experience are improved, and the accuracy and comprehensiveness of information expression are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047146A_ABST
    Figure CN120047146A_ABST
Patent Text Reader

Abstract

The invention provides a payment processing method and device based on facial recognition, equipment and a medium, relates to the technical field of artificial intelligence, can be applied to scenes such as cloud technology, artificial intelligence, smart traffic and aided driving, and comprises the steps of responding to a facial recognition payment request of a target object; acquiring image data acquired by image acquisition equipment and audio data corresponding to the image data; performing feature analysis on the audio data according to the number of personnel to obtain an audio analysis result; if the audio analysis result indicates that the number of persons corresponding to the audio data is a single person, performing face recognition on the image data based on a face recognition model corresponding to a single person scene to obtain target face data of the target object; and performing payment operation corresponding to the facial recognition payment request based on the target facial data. According to the invention, the redundant time consumption of payment identification can be effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a payment processing method, apparatus, device and medium based on facial recognition. Background Art

[0002] With the development of artificial intelligence, intelligent recognition payment based on facial features has become a universal business model. Compared with traditional electronic payment methods, intelligent recognition of facial features does not require carrying mobile devices or physical objects, and realizes efficient payment. However, in actual scenarios, due to the complexity of the environment, the payment terminal may collect the faces of multiple objects at one time. Therefore, before each recognition payment, facial frame detection and facial screening are performed by default through algorithms such as multi-person detection and attention mechanism to avoid misidentification of objects. Although this method effectively improves the accuracy and security of facial payment, it is necessary to load algorithm modules such as multi-person detection and attention before each payment, which occupy a lot of memory, greatly increasing the operation time and reducing payment efficiency and business experience. Summary of the invention

[0003] The present application provides a payment processing method, apparatus, device and medium based on facial recognition, which can significantly improve the accuracy and comprehensiveness of information expression in payment processing based on facial recognition.

[0004] In one aspect, the present application provides a payment processing method based on facial recognition, the method comprising:

[0005] In response to a facial recognition payment request of a target object, acquiring image data captured by an image capture device and audio data corresponding to the image data;

[0006] Performing feature analysis on the audio data based on the number of people to obtain an audio analysis result;

[0007] If the audio analysis result indicates that the number of people corresponding to the audio data is a single person, performing facial recognition on the image data based on a facial recognition model corresponding to a single-person scene to obtain target facial data of the target object;

[0008] A payment operation corresponding to the facial recognition payment request is performed based on the target facial data.

[0009] Another aspect provides a payment processing device based on facial recognition, the device comprising:

[0010] Data acquisition module: used for acquiring image data acquired by the image acquisition device and audio data corresponding to the image data in response to the facial recognition payment request of the target object;

[0011] Feature analysis module: used to perform feature analysis on the audio data based on the number of people to obtain audio analysis results;

[0012] Facial recognition module: for performing facial recognition on the image data based on a facial recognition model corresponding to a single-person scene to obtain target facial data of the target object if the audio analysis result indicates that the number of people corresponding to the audio data is a single person;

[0013] Payment module: used to perform the payment operation corresponding to the facial recognition payment request based on the target facial data.

[0014] In a possible implementation manner, the audio analysis result includes sound source separation data; the feature analysis module includes a sound source separation submodule: configured to perform sound source separation detection on the audio data to obtain the sound source separation data, wherein the sound source separation data includes at least one sound source signal corresponding to the audio data;

[0015] Correspondingly, the device further comprises a judgment module: configured to determine, if the sound source separation data comprises one sound source signal, that the audio analysis result indicates that the number of persons corresponding to the audio data is a single person.

[0016] In a possible implementation manner, the sound source separation submodule includes:

[0017] Framing unit: used for performing framing processing on the audio data to obtain multiple audio frames;

[0018] A feature extraction unit: configured to extract energy features from the plurality of audio frames to obtain audio energy features;

[0019] A sound source separation unit is used to perform sound source separation corresponding to the audio energy feature based on a sound source separation model to obtain the sound source separation data; the sound source separation model is obtained by performing constraint training for sound source separation on a preset first deep neural network with sample audio energy features as input and sample sound source signals corresponding to the sample audio energy features as expected output, and the sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to more than one sample sound sources.

[0020] In a possible implementation manner, the audio analysis result includes noise index data; and the feature analysis module includes:

[0021] Noise determination submodule: used to determine the noise signal in the audio data;

[0022] Amplitude calculation submodule: used to calculate the noise field amplitude of the noise signal to obtain noise index data, wherein the noise index data is used to indicate the intensity of the environmental noise corresponding to the audio data;

[0023] The judgment module is also used for: after performing feature analysis on the audio data for the number of people and obtaining an audio analysis result, if the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining that the audio analysis result indicates that the number of people corresponding to the audio data is a single person.

[0024] In a possible implementation manner, the noise determination submodule includes:

[0025] Framing unit: used for performing framing processing on the audio data to obtain multiple audio frames;

[0026] A feature extraction unit: configured to extract energy features from the plurality of audio frames to obtain audio energy features;

[0027] A noise estimation unit is used to perform noise estimation corresponding to the audio energy feature based on a noise estimation model to obtain the noise signal; the noise estimation model is obtained by constrained training of a preset second deep neural network for noise estimation with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output, and the sample audio energy features are obtained by energy feature extraction based on audio data of a sample sound source or mixed audio data corresponding to more than one sample sound sources.

[0028] In a possible implementation manner, the terminal is provided with a plurality of audio acquisition devices arranged in an array, and the data acquisition module includes:

[0029] Initial audio acquisition submodule: used to acquire multiple initial audio data collected by multiple audio collection devices;

[0030] The sound source localization submodule is used to localize the sound source of the plurality of initial audio data to obtain the sound source position information of each of the initial audio data;

[0031] Audio filtering submodule: used to determine, based on the sound source position information, from the multiple initial audio data, an audio signal whose sound source is located within a first spatial range of the terminal to obtain the audio data, wherein the first spatial range is located in front of the display screen of the terminal and within the field of view of the image acquisition device.

[0032] In a possible implementation manner, the device further includes:

[0033] A face detection module: configured to perform face detection on the image data based on a multi-person detection model to obtain multiple face images if the audio analysis result indicates that the number of persons corresponding to the audio data is more than one;

[0034] A face screening module: used for screening out a target face image of the target object from the multiple face images;

[0035] The facial recognition module is also used to perform facial recognition on the target facial image based on the facial recognition model to obtain the target facial data.

[0036] In a possible implementation manner, the device further includes:

[0037] The instruction sound source positioning module is used for, before responding to the facial recognition payment request of the target object, responding to the facial payment voice instruction, performing sound source positioning on the instruction audio data carried by the facial payment voice instruction, and obtaining the instruction sound source position information, wherein the instruction sound source position information is used to indicate the relative position of the issuing object of the instruction audio data relative to the image acquisition device;

[0038] Payment request generation module: used to generate a facial recognition payment request corresponding to the facial payment voice instruction if the instruction sound source location information indicates that the issuing object is located within a second spatial range, and the second spatial range is located in front of the display screen of the terminal and the distance between the second spatial range and the display screen is within a preset distance range.

[0039] In a possible implementation, the face screening module includes:

[0040] A face position acquisition submodule: used to acquire face position information corresponding to each of the plurality of face images, wherein the face position information is used to represent the position of the object to which the face image belongs relative to the image acquisition device;

[0041] The facial image matching submodule is used to determine the facial image whose facial position information matches the instruction sound source position information among the multiple facial images as the target facial image.

[0042] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the facial recognition-based payment processing method as described above.

[0043] On the other hand, a computer-readable storage medium is provided, in which at least one instruction or at least one program is stored, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the facial recognition-based payment processing method as described above.

[0044] On the other hand, a server is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the facial recognition-based payment processing method as described above.

[0045] On the other hand, a terminal is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the facial recognition-based payment processing method as described above.

[0046] On the other hand, a computer program product or a computer program is provided, which includes computer instructions, and when the computer instructions are executed by a processor, the payment processing method based on facial recognition as described above is implemented.

[0047] The payment processing method, apparatus, device, storage medium, server, terminal, computer program and computer program product based on facial recognition provided by this application have the following technical effects:

[0048] The technical solution of the present application responds to the facial recognition payment request of the target object, obtains the image data collected by the image acquisition device and the audio data corresponding to the image data, and then performs feature analysis on the audio data for the number of people to obtain the audio analysis result. If the audio analysis result indicates that the number of people corresponding to the audio data is a single person, the image data is subjected to facial recognition based on the facial recognition model corresponding to the single-person scene to obtain the target facial data of the target object, thereby performing the payment operation corresponding to the facial recognition payment request based on the target facial data. In this way, before loading the multi-person detection module, feature analysis is first performed based on the audio data to obtain an audio analysis result that can indicate the number of people in front of the current terminal. When it indicates that the current environment is a single-person scene, facial recognition is directly performed without triggering the loading of related large-memory business modules such as multi-person detection, thereby optimizing the load time of the facial payment process, improving payment efficiency, and improving product experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application;

[0051] Figure 2 It is a flowchart of a payment processing method based on facial recognition provided by an embodiment of the present application;

[0052] Figure 3 It is a flowchart of another payment processing method based on facial recognition provided by an embodiment of the present application;

[0053] Figure 4 It is a flowchart of another payment processing method based on facial recognition provided by an embodiment of the present application;

[0054] Figure 5 It is a flowchart of another payment processing method based on facial recognition provided by an embodiment of the present application;

[0055] Figure 6 It is a schematic diagram of a payment processing device based on facial recognition provided in an embodiment of the present application;

[0056] Figure 7 This is a hardware structure block diagram of an electronic device that executes a payment processing method based on facial recognition provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0058] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server comprising a series of steps or submodules is not necessarily limited to those steps or submodules clearly listed, but may include other steps or submodules that are not clearly listed or inherent to these processes, methods, products or devices.

[0059] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0060] 3D camera: Similar to traditional cameras, it adds living body-related software and hardware, including depth cameras and infrared cameras, to ensure information security.

[0061] SN: string serial number, which can uniquely identify a device ID.

[0062] SQLite: is a lightweight database and a relational database management system that complies with ACID.

[0063] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0064] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics and other technologies. Among them, the pre-trained model is also called a large model or a basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0065] Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development of deep learning, which integrates the above technologies.

[0066] Deep learning: The concept of deep learning originates from the study of artificial neural networks. A multi-layer perceptron with multiple hidden layers is a deep learning structure. Deep learning discovers distributed feature representations of data by combining low-level features to form more abstract high-level representations of attribute categories or features.

[0067] Computer Vision (CV) is a science that studies how to make machines "see". To put it more concretely, it refers to machine vision that uses cameras and computers to replace human eyes to identify, detect and measure targets, and further performs image processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0068] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless cars, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0069] See also Figure 1 , Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present application, such as Figure 1 As shown, the application environment may include at least a terminal 01 and a server 02. In practical applications, the terminal 01 and the server 02 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0070] The server 02 in the embodiment of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0071] Specifically, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology can be applied to various fields, such as medical cloud, cloud Internet of Things, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calls, and cloud social networking. Cloud technology is based on the cloud computing business model application. It distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space, and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" are infinitely expandable in the eyes of users, and can be obtained at any time, used on demand, expanded at any time, and paid for by use. As a basic capability provider for cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service)) platform will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose to use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0072] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS layer, and the SaaS (Software as a Service) layer can be deployed on the PaaS layer, or SaaS can be directly deployed on IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is a variety of business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.

[0073] Specifically, the server 02 mentioned above may include a physical device, which may specifically include a network communication submodule, a processor, a memory, etc., and may also include software running in the physical device, which may specifically include an application program, etc.

[0074] Specifically, terminal 01 may include physical devices such as smart phones, desktop computers, tablet computers, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices, and may also include software running in physical devices, such as applications.

[0075] In the embodiment of the present application, the terminal 01 can be used to respond to the trigger operation related to facial recognition payment submitted by the target object, generate a facial recognition payment request, and then collect image data and audio data to perform feature analysis on the audio data for the number of people to obtain audio analysis results; based on the number of people corresponding to the audio data, determine whether to trigger and load a multi-person detection model to perform facial detection and target facial image screening of the target object, and send a related business payment request to the server 02 based on the target facial data finally obtained, so that the server 02 performs the corresponding payment processing in response to the business payment request. The server 02 provides payment services and data storage services, such as using a SQLite database to store facial data for payment comparison and matching.

[0076] Furthermore, it is understandable that Figure 1 What is shown is merely an application environment of a payment processing method based on facial recognition, and the application environment may include more or fewer nodes, and this application does not impose any limitation thereto.

[0077] The application environment involved in the embodiments of the present application, or the terminal 01 and server 02 in the application environment, etc., can be a distributed system formed by connecting a client and multiple nodes (any form of computing devices in the access network, such as servers and terminals) through network communication. The distributed system can be a blockchain system, which can provide the above-mentioned payment processing services based on facial recognition, model training services, and data storage services. In one embodiment, the application environment can run a dialogue system, which can store basic corpus information and include a role setting module, a question and answer module (QA module), an expression processing module, and a reply generation module (reference Figure 7 ).

[0078] The following is an introduction to the technical solution of this application based on the above application environment and dialogue system. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. Please refer to Figure 2 , Figure 2 This is a flowchart of a payment processing method based on facial recognition provided by an embodiment of the present application. This specification provides method operation steps such as the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many orders, and does not represent the only order of execution. When the actual system or server product is executed, it can be executed in the order of the method shown in the embodiment or the figure or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 2 As shown, the following steps S201-S207 may be included:

[0079] S201: In response to a facial recognition payment request from a target object, image data captured by an image capture device and audio data corresponding to the image data are acquired.

[0080] Specifically, terminal 01 can be provided with a display screen, an image acquisition device and an audio acquisition device. The display screen can be used to display a facial acquisition frame, etc., so as to display payment-related operation information and interfaces to the target object. The field of view of the image acquisition device is located in front of the display screen, and is used to collect image data directly in front of the display screen for payment processing based on facial recognition. There can be multiple audio acquisition devices, such as multiple audio acquisition devices can be arranged in an array, for audio data acquisition and corresponding sound source positioning.

[0081] In some embodiments, in order to enhance the accuracy and security of facial recognition of an object, the image acquisition device may include a 3D camera, and the acquired image data may include, in addition to color images (such as RGB images), images including depth information such as depth maps, and may also include infrared images, etc.

[0082] In the embodiment of the present application, the terminal may run a facial recognition application, which may specifically include a facial payment module, an algorithm auxiliary module, and an algorithm scheduling module. The facial payment module is used for facial acquisition and image screening, the algorithm auxiliary module runs a multi-person detection module and an attention detection module, etc., to achieve multi-person facial detection, and the algorithm scheduling module is used to perform facial recognition on the acquired facial image of the target object to obtain facial data for payment processing. Facial acquisition refers to calling an image acquisition device to perform image acquisition. It can be understood that what is acquired is image stream data, such as RGB image stream, depth image stream, and infrared image stream, etc. Image screening is used to filter out image data for subsequent facial recognition from the image stream data. Specifically, it can be based on a comprehensive evaluation of parameter indicators such as facial area size, facial angle, image contrast, image brightness and clarity, so as to determine the optimal image data containing facial images. The existing facial recognition payment scheme will introduce a multi-person detection module and an attention detection module before the image screening stage and the subsequent facial recognition stage. Through the multi-person detection algorithm and combined with the attention detection mechanism, it is determined whether there are facial image areas of multiple people in the current image. In actual scenarios, there are single-person payment scenarios, that is, there are no multiple faces in front of the terminal. The use of the above existing schemes will still load the multi-person detection module and the attention detection module. The algorithm parameters of the multi-person detection module and the attention detection module are large, and the model occupies a high memory. Loading and detection are time-consuming, resulting in redundant resource occupation, affecting the facial payment experience. The facial payment module of this application also runs a personnel number detection module, which is used to detect the number of people in front of the terminal.

[0083] It can be understood that there may be multiple objects in front of the terminal. The target object refers to the object that initiates facial payment for the current terminal. The facial recognition payment request refers to the request generated by the terminal based on the facial payment triggering operation of the target object. The facial payment triggering operation may include but is not limited to voice command triggering, button submission, etc.

[0084] Specifically, the audio data may be mixed audio captured by an audio capture device, including sound signals in the terminal's surrounding environment. The audio data corresponding to the image data refers to an audio stream that matches the image capture time of the image data, that is, audio that includes the image capture time and the time difference between the audio capture time and the image capture time is within a preset time difference range.

[0085] In some embodiments, the terminal is provided with a plurality of audio collection devices arranged in an array, such as a microphone array, preferably a linear microphone array, in which a plurality of microphones are arranged in a straight line at a certain interval. Figure 3 , S201 may include S301-S305:

[0086] S301: Acquire multiple initial audio data collected by multiple audio collection devices;

[0087] S303: Perform sound source positioning on the multiple initial audio data to obtain the sound source position information of each initial audio data;

[0088] S305: Determine, based on the sound source position information, from the multiple initial audio data an audio signal whose sound source is located within a first spatial range of the terminal to obtain audio data.

[0089] Specifically, the initial audio data is the data obtained by preprocessing the original audio collected by the audio acquisition device, and the preprocessing may include but is not limited to filtering, noise reduction, gain control, etc. Sound source localization refers to the technology of determining the location of the sound source by analyzing the propagation characteristics of sound waves and the received audio signals. It uses multiple audio acquisition devices (such as microphones) distributed in different spatial positions, and the received signals have different time delays and amplitude differences, so that the sound source position can be calculated. The sound source position information here is used to characterize the relative position between the sound source of the initial audio data and the audio acquisition device, and is used to determine the direction of the sound source of the initial audio data relative to the terminal.

[0090] Specifically, a sound source localization algorithm is used to process the multiple initial audio data obtained by preprocessing to determine the location of the sound source and obtain the sound source location information. The sound source localization algorithm used may include but is not limited to a beamforming algorithm, a covariance matching algorithm, and a least squares method, etc., which are not specifically limited here. Furthermore, a sound source localization tracking algorithm may also be used to track the sound source to maintain the positional relationship between the detection audio acquisition device and the sound source. The sound source localization tracking algorithm may include but is not limited to a Kalman filter, a particle filter, and an extended Kalman filter, etc., which are not specifically limited here.

[0091] Specifically, the first spatial range is located in front of the display screen of the terminal and within the field of view of the image acquisition device, that is, the first spatial range is limited to only using the initial audio data generated by the sound source directly in front of the display screen (or image acquisition device). It can be understood that the image acquisition device is facing the front of the terminal display screen, and the above scheme accurately filters out the audio data from the front of the terminal to reduce interference from stray sound sources, reduce the difficulty of sound source separation, and improve the efficiency of subsequent personnel location.

[0092] In some embodiments, the sound source location information can be used to directly filter the audio data, and the sound source direction corresponding to the sound source location information can be matched with the camera direction to determine the angle difference between the sound source direction and the camera direction. If the angle difference is less than a certain angle threshold, the sound source is considered to be located directly in front of the image acquisition device, i.e., the display screen. Otherwise, it is not within the first spatial range. In other embodiments, a sound source separation process can be used in combination with the aforementioned sound source location information to perform sound source separation on multiple initial audio data to separate the sound signal directly in front and obtain audio data. The sound source separation algorithm used can be, but is not limited to, blind source separation, non-blind source separation, and separation based on a sound source separation model.

[0093] S203: Perform feature analysis on the audio data based on the number of people to obtain an audio analysis result.

[0094] Specifically, the audio analysis result is used to characterize the sound source index characteristics of the audio data, and may include but is not limited to at least one of sound source separation data and noise index data. The number of people refers to how many people each sound source forming the audio data comes from.

[0095] Specifically, feature analysis is to analyze the audio energy features of audio data to obtain audio analysis results. Audio energy features are used to characterize the energy of each frame of audio data in the time domain, that is, to describe the characteristics of audio energy, including but not limited to short-time energy, short-time average amplitude, short-time power spectrum, etc.

[0096] In some embodiments, S203 includes S401: performing sound source separation detection on the audio data to obtain sound source separation data.

[0097] Specifically, sound source separation detection refers to separating different sound sources in audio data to separate different sound source signals from a piece of audio data. Accordingly, the sound source separation data includes at least one sound source signal corresponding to the audio data. It can be understood that if there is only a single sound source signal, it indicates that the current audio data comes from the same person, which is a single-person scene. If there are more than one sound source signal, it is determined that the audio data comes from multiple people, indicating that a multi-sound source multi-person scene is hit.

[0098] In some cases, when the sound source separation process is used in the aforementioned S305 to separate the sound signal directly in front and obtain audio data, the sound source separation data in S401 can reuse the result obtained by the sound source separation process in S305, thereby reducing the amount of data processing. In other cases, the sound source separation process is not used in the aforementioned S305, and the sound source separation algorithm used in the sound source separation detection in S401 may include blind source separation, non-blind source separation, and separation based on a sound source separation model. The above-mentioned sound separation data is obtained by performing sound source separation processing on the audio energy features, and the audio energy features are obtained by performing frame processing on the audio data to obtain multiple audio frames, and extracting energy features from the multiple audio frames.

[0099] Accordingly, in one embodiment, S401 includes: performing frame processing on the audio data to obtain multiple audio frames; performing energy feature extraction on the multiple audio frames to obtain audio energy features, that is, obtaining energy features of each audio frame, such as short-time energy, short-time average amplitude, short-time power spectrum, etc.; performing sound source separation corresponding to the audio energy features based on the sound source separation model to obtain sound source separation data. In this way, sound source separation is performed through the intelligent model to improve the accuracy of sound source separation.

[0100] Specifically, the sound source separation model mentioned above is obtained by constraining the sound source separation of the preset first deep neural network with the sample audio energy feature as input and the sample sound source signal corresponding to the sample audio energy feature as the expected output. The sample audio energy feature is obtained by extracting the energy feature based on the audio data of a sample sound source or the mixed audio data corresponding to more than one sample sound source. The audio data of the sample sound source can be human voice audio, and the mixed audio data can refer to audio including multiple human voices. The sample sound source signal is used as a training label, which refers to the sound signal emitted by a single sample sound source, and is used for supervised training of the preset deep neural network. It can be understood that the sample audio energy feature here is similar to the method of obtaining the aforementioned audio energy feature, and will not be elaborated.

[0101] In one embodiment, an audio data set containing multiple human voices is obtained, and the voices of different human voices are separated separately as audio data of sample sound sources, and then the audio data of the sample sound source is converted into a computer-processable feature vector, and the methods used include but are not limited to short-time Fourier transform (STFT), Mel-frequency cepstral coefficients (MFCC) and energy feature algorithms. The initial model is constructed using a preset deep neural network (such as a convolutional neural network (CNN), etc.), and then the initial model is trained. During the training process, optimizers such as Adam or SGD can be used for optimization, and the learning rate is set based on the needs and the loss is calculated using an existing loss function, such as cross entropy. The audio data to be separated is feature processed and input into the trained sound source separation model, and the sound source signals of different human voices predicted by the output model are output. In the case of long audio data, long audio data can be processed by sliding windows.

[0102] Accordingly, after S203, the method further includes S501-S503:

[0103] S501: If the sound source separation data includes a sound source signal, determining that the audio analysis result indicates that the number of people corresponding to the audio data is a single person;

[0104] S503: If the sound source separation data includes more than one sound source signal, determine that the audio analysis result indicates that the number of persons corresponding to the audio data is more than one person.

[0105] In this way, the energy characteristics are analyzed through the intelligent model, and the sound characteristics of different human voices are modeled and separated. The audio data is divided into different sound sources, and then the number of different sound source signals separated is the different people, that is, the number of people. If there is more than one, it means that there may be multiple people around, and related business modules such as the multi-person detection module need to be loaded. If there is equal to one, facial recognition and payment processing are directly performed to improve payment efficiency.

[0106] In some other embodiments, S203 includes S403-S405:

[0107] S403: Determine a noise signal in the audio data;

[0108] S405: Calculate the noise field amplitude of the noise signal to obtain noise index data, which is used to indicate the intensity of the environmental noise corresponding to the audio data, such as the noise decibel value.

[0109] Specifically, the noise signal refers to the noise audio separated from the audio data, and its acquisition can specifically include: performing frame processing on the audio data to obtain multiple audio frames; performing energy feature extraction on the multiple audio frames to obtain audio energy features; performing noise estimation corresponding to the audio energy features based on the noise estimation model to obtain the noise signal. It can be understood that the audio frames and audio energy features in the noise signal acquisition process and the sound source separation process can be reused with each other.

[0110] Specifically, the noise estimation model is obtained by constraining training of a preset second deep neural network for noise estimation with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output. The sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to more than one sample sound sources.

[0111] It can be understood that the sample noise signal is equivalent to a training label for supervised training. The sample audio energy features, the audio data of the sample sound source, and the mixed audio data here can be shared with the training data of the sound source separation model, or can be separately set training data. The second deep neural network can be, for example, a convolutional neural network. In this way, noise signals are generated through an intelligent model to assist in judging the noisiness of the environment, thereby improving the accuracy of multi-person scene judgment.

[0112] The extracted features are analyzed by the noise estimation model to obtain the energy distribution characteristics of the noise, such as average energy, peak energy, etc., so as to estimate the noise and generate a noise signal. Specifically, the noise field amplitude calculation refers to calculating the noise field decibel value of the noise signal using the noise intensity algorithm.

[0113] Accordingly, after S203, the method may further include S505-S507:

[0114] S505: If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining that the audio analysis result indicates that the number of people corresponding to the audio data is one person;

[0115] S507: If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, determine that the audio analysis result indicates that the number of persons corresponding to the audio data is more than one person.

[0116] It is understandable that when there are too many people around the terminal, high-decibel noise will lead to inadequate sound separation. By introducing noise signal detection and noise intensity index calculation as a separate judgment logic, if the decibel value exceeds the preset intensity threshold, it can be judged that the decibel value of the noise field is unqualified, indicating that there may be many people around, and the multi-person detection module needs to be loaded. If it is lower than the preset intensity threshold, it is determined to be a single-person scene, so as to improve the accuracy of judgment of multi-person and single-person scenes through noise detection.

[0117] In some embodiments, the sound source separation detection and the noise field amplitude calculation of the noise signal are performed in parallel. After the audio data is obtained, the audio data is subjected to sound source separation detection and noise signal estimation and corresponding noise amplitude calculation respectively, and the sound source separation result and noise index data are obtained respectively; in some cases, if the number of persons indicated by any one of the sound source separation results and the noise index data is more than one, S209 is triggered, and if the number of persons indicated by the sound source separation result and the noise index data is a single person, S205 is triggered; in other cases, if the number of persons indicated by any one of the sound source separation results and the noise index data is a single person, S205 is triggered, and if the number of persons indicated by the sound source separation result and the noise index data is more than one person, S209 is triggered. In this way, through parallel detection, the two methods assist each other to achieve accurate discrimination of multi-person and single-person scenes.

[0118] In other embodiments, after obtaining the audio data, a sound source separation test is first performed. If the sound source separation data includes more than one sound source signal, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is multiple people, and S209 is triggered. If the sound source separation data includes one sound source signal, the audio data is subjected to noise signal estimation and noise field amplitude calculation. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is single person, and S205 is triggered. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is multiple people, and S209 is triggered. In this way, noise detection is selectively triggered to ensure detection accuracy while reducing resource usage for detecting the number of people.

[0119] In some other embodiments, after obtaining the audio data, the noise signal estimation and noise field amplitude calculation are first performed. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than the preset intensity threshold, if the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than the preset intensity threshold, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is multiple people, and S209 is triggered. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than the preset intensity threshold, it is determined that the noise index result indicates a one-person scene. Then, the sound source separation result of the audio data is obtained. If the sound source separation data includes one sound source signal, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is single person, and S205 is triggered. If the sound source separation data includes more than one sound source signal, it is determined that the audio analysis result indicates that the number of people corresponding to the audio data is multiple people, and S209 is triggered. In this way, the sound source separation detection is selectively triggered to ensure the detection accuracy while reducing the resource occupation of the number of people detection.

[0120] It can be understood that audio energy features are used in both noise signal estimation and sound source separation detection. Therefore, in the above-mentioned various schemes, frame processing and energy feature extraction are only performed once, that is, frame processing is first performed to obtain multiple audio frames, and then energy features of the multiple audio frames are extracted to obtain audio energy features, and the audio energy features are used for noise signal estimation and sound source separation processing.

[0121] This solution is based on noise and sound source detection, pre-detection and identification of the complexity of people in the current scene. If the scene does not have too much noise or related human voice conversations, it is likely to be a single-person payment scene. At this time, the algorithm can directly perform face recognition without loading related capabilities for multi-person detection. If there is too much noise or redundant human voice conversations, a multi-person detection module is added to improve payment efficiency and reduce resource loss.

[0122] S205: If the audio analysis result indicates that the number of people corresponding to the audio data is a single person, facial recognition is performed on the image data based on a facial recognition model corresponding to a single-person scene to obtain target facial data of the target object.

[0123] In some embodiments, if the result based on noise signal detection is a single person or the result based on sound source separation detection is a single person, it is confirmed that the audio analysis result indicates that the number of people corresponding to the audio data is a single person, and then the algorithm scheduling module directly calls the facial recognition model applicable to single-person facial recognition to perform facial recognition to obtain the target facial data without loading the multi-person detection algorithm and the corresponding attention detection module, etc. The target facial data can be a feature vector that characterizes the facial image area of ​​the target object and has person-specificity. The facial recognition model can be constructed using existing recognition models and trained with facial recognition training data, and there is no limitation here. If the result based on noise signal detection is more than one person and the result based on sound source separation detection is more than one person, it is confirmed that the audio analysis result indicates that the number of people corresponding to the audio data is more than one, and then the multi-person detection algorithm is called accordingly to perform multi-person facial detection.

[0124] In some other embodiments, if the results obtained based on noise signal detection and sound source separation detection both indicate that the number of people is a single person, then the audio analysis result indicates that the number of people corresponding to the audio data is a single person, that is, the single-person scene is clarified, and then facial recognition is directly performed without loading a multi-person detection algorithm and a corresponding attention detection module, etc. If the result based on noise signal detection is more than one person or the result based on sound source separation detection is more than one person, and it is confirmed that the audio analysis result indicates that the number of people corresponding to the audio data is more than one, then the multi-person detection algorithm is called accordingly to perform multi-person facial detection. In this way, the XOR judgment of noise and sound source separation is realized, the accuracy of single-person scene recognition is improved, and the missed detection rate of multi-person scenes is reduced.

[0125] S207: Perform a payment operation corresponding to the facial recognition payment request based on the target facial data.

[0126] It is understandable that the facial recognition model can be run on the terminal, and then the terminal sends the obtained target facial data to the server for facial data matching operation, thereby realizing the payment operation. Alternatively, the facial recognition model runs on the server, and after the terminal determines the single-person scene, it sends the filtered image data to the server to input the facial recognition model and output the target facial data for facial data matching, thereby realizing the payment operation.

[0127] Based on the above technical solution, before loading the multi-person detection module, this application first performs feature analysis based on audio data to obtain an audio analysis result that can indicate the number of people in front of the current terminal. When it indicates that the current environment is a single-person scene, facial recognition is directly performed without triggering the loading of related large-memory business modules such as multi-person detection, thereby optimizing the load time of the facial payment process, improving payment efficiency, and improving product experience.

[0128] Specifically, after the facial recognition service receives the image data uploaded from the terminal, it extracts features from the current image data to obtain the target facial data, and compares the target facial data with the features stored in the database, and compares the feature data with the highest matching score with the facial data in the backend database to determine whether the payment request of the target object is compliant and the corresponding payment information, and then executes the payment under compliance. For example, after determining compliance, the downstream server returns the target object's final payment account information or payment code information, etc., for payment interaction verification.

[0129] Based on some or all of the above implementations, in the embodiments of this application, reference Figure 4 The method further includes S209-S213:

[0130] S209: If the audio analysis result indicates that the number of persons corresponding to the audio data is more than one, perform facial detection on the image data based on a multi-person detection model to obtain multiple facial images;

[0131] S211: Filtering out a target facial image of a target object from a plurality of facial images;

[0132] S213: Perform facial recognition on the target facial image based on the facial recognition model to obtain target facial data.

[0133] It is understandable that the multi-person detection model can adopt an existing detection model, which can be constructed based on a deep neural network and trained by using multi-person facial images for image detection. It can be run on a terminal to obtain multiple facial images directly at the terminal, or it can be run on a server. The terminal sends image data to the server and carries multi-person detection trigger information so that the server calls the multi-person detection model to perform facial detection on the image data. After obtaining multiple facial images, it is necessary to screen the images of the target object. The screening process can be operated at the terminal or implemented by the server. In this way, multi-person scene detection is achieved by predicting the number of people. When it is determined that there are multiple people, the multi-person detection module is loaded, and then the target facial image to be identified is obtained through facial screening, thereby improving the overall payment efficiency while ensuring payment security.

[0134] In some cases, facial images can be screened based on information such as the image area size and image area position corresponding to multiple facial images in the image data. It can be understood that the larger the image area size and the closer the image area position is to the image center, the higher the probability that it is the target facial data. The image area size, image area position and other information can be scored based on a voting scoring method, and then the scores can be added together to determine the facial image with the highest score as the target facial image.

[0135] Based on some or all of the above embodiments, in some embodiments, the facial payment scenario can be superimposed with voice instructions, that is, the facial recognition payment request can be triggered by a voice instruction. Figure 5 , before S201, the method further includes S101-S103:

[0136] S101: In response to a facial payment voice instruction, performing sound source positioning on instruction audio data carried by the facial payment voice instruction to obtain instruction sound source position information;

[0137] S103: If the instruction sound source location information indicates that the issuing object is located within the second space, a facial recognition payment request corresponding to the facial payment voice instruction is generated.

[0138] Specifically, after the audio acquisition device of the terminal acquires the original audio data and obtains the initial audio data, it can perform voice recognition with it. If the voice recognition result shows that the corresponding voice text hits the facial payment voice command, such as "face payment", etc., the position of the audio signal corresponding to the voice text in the initial audio data is located through the sound source localization algorithm to obtain the command sound source position information. The command sound source position indicates the relative position of the issuing object of the command audio data relative to the image acquisition device, and then according to the position information of the image acquisition device, the position information of the display screen and the command sound source position information, the relative position of the issuing object relative to the display screen is determined. If the determined relative position is within the second spatial range, it is determined that the command sound source position information indicates that the issuing object is within the second spatial range. Otherwise, it is determined that the position of the issuing object exceeds the second spatial range. The sound source localization algorithm here can be similar to the sound source localization processing in the previous text, which will not be repeated here.

[0139] If the instruction sound source location information indicates that the issuing object is within the second spatial range, the object is determined to be the target object for initiating the facial payment operation of the current terminal, thereby generating a facial recognition payment request, and the second spatial range is located in front of the display screen of the terminal and the distance between the display screen and the terminal is within the preset distance range. If the instruction sound source location information indicates that the issuing object is beyond the second spatial range, it means that the issuing object is not the target object, and then no facial recognition payment request is generated, but the sound signals around the terminal are continuously collected. In this way, in the case of multiple payment terminals, crosstalk of voice commands to other terminals is avoided, and the accuracy of voice triggering is improved.

[0140] Accordingly, in the scenario where the voice command triggers the facial payment operation, if it is determined based on the above audio analysis result that the current number of people is more than one person, then S211 selects the target facial image of the target object from the multiple facial images, including:

[0141] S601: Obtain facial position information corresponding to each of a plurality of facial images;

[0142] S603: Determine a facial image whose facial position information matches the instruction sound source position information among the multiple facial images as a target facial image.

[0143] Specifically, the facial position information is used to characterize the position of the object to which the facial image belongs relative to the image acquisition device, and can be obtained by calculating the relative position of the object and the image acquisition device based on the image area position of the facial image in the image data and the camera parameters for acquiring the image data. The image area position can be obtained based on the facial detection in S209.

[0144] Furthermore, the facial position information in each facial image is compared with the position information of the command sound source, and the facial image with the closest position is determined as the target facial image. In this way, the target facial image is determined in combination with the positioning result of the voice command, which improves the efficiency of target determination in multi-person scenes.

[0145] In some embodiments, in order to avoid comparison errors caused by changes in the position of the target object in the time interval between the issuance of the voice command and the detection of multiple faces, after determining that a facial recognition payment request has been generated, the terminal generates a prompt sound for re-entering the voice signal, such as "Repeat facial payment command", to prompt the target object to re-enter the voice command, and then separate the voice command source to update the command source location information as comparison data for the facial position information, thereby improving the accuracy of position comparison.

[0146] This application builds a built-in audio acquisition device in the payment terminal device, which continuously acquires the sound source directly in front of the device in silent state, and makes timbre and noise judgment on the sound source. If multiple timbres or high noise are hit, the multi-person detection and attention judgment module is additionally loaded when the facial recognition payment is started, otherwise it is not loaded, so as to optimize the time consumption caused by the algorithm load and improve the business experience.

[0147] The embodiment of the present application also provides a payment processing device 700 based on facial recognition, which is arranged in a terminal, such as Figure 6 As shown, Figure 6 A schematic structural diagram of a facial recognition-based payment processing device provided in an embodiment of the present application is shown, and the device may include the following modules.

[0148] Data acquisition module 10: used to obtain image data acquired by an image acquisition device and audio data corresponding to the image data in response to a facial recognition payment request of a target object;

[0149] Feature analysis module 20: used to perform feature analysis on the audio data according to the number of people, and obtain audio analysis results;

[0150] Facial recognition module 30: for performing facial recognition on the image data based on a facial recognition model corresponding to a single-person scene to obtain target facial data of the target object if the audio analysis result indicates that the number of people corresponding to the audio data is a single person;

[0151] Payment module 40: used to perform payment operations corresponding to facial recognition payment requests based on target facial data.

[0152] In some embodiments, the audio analysis result includes sound source separation data; the feature analysis module 20 includes a sound source separation submodule: used to perform sound source separation detection on the audio data to obtain sound source separation data, the sound source separation data including at least one sound source signal corresponding to the audio data;

[0153] Correspondingly, the device also includes a judgment module: after performing feature analysis on the audio data for the number of people and obtaining the audio analysis result, if the sound source separation data includes a sound source signal, determine that the audio analysis result indicates that the number of people corresponding to the audio data is a single person.

[0154] In some embodiments, the sound source separation submodule includes:

[0155] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;

[0156] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;

[0157] Sound source separation unit: used to perform sound source separation corresponding to the audio energy feature based on the sound source separation model to obtain sound source separation data; the sound source separation model is obtained by constraining the sound source separation training of a preset first deep neural network with the sample audio energy feature as input and the sample sound source signal corresponding to the sample audio energy feature as the expected output, and the sample audio energy feature is obtained by extracting energy features based on the audio data of a sample sound source or the mixed audio data corresponding to more than one sample sound sources.

[0158] In some embodiments, the audio analysis result includes noise index data; the feature analysis module 20 includes:

[0159] Noise determination submodule: used to determine the noise signal in the audio data;

[0160] Amplitude calculation submodule: used to calculate the noise field amplitude of the noise signal to obtain noise index data, which is used to indicate the environmental noise intensity corresponding to the audio data;

[0161] The judgment module is also used for: after performing feature analysis on the audio data for the number of people and obtaining the audio analysis result, if the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining that the audio analysis result indicates that the number of people corresponding to the audio data is a single person.

[0162] In some embodiments, the noise determination submodule includes:

[0163] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;

[0164] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;

[0165] Noise estimation unit: used to perform noise estimation corresponding to the audio energy feature based on the noise estimation model to obtain a noise signal; the noise estimation model is obtained by constrained training of noise estimation on a preset second deep neural network with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output, and the sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to more than one sample sound sources.

[0166] In some embodiments, the terminal is provided with a plurality of audio acquisition devices arranged in an array, and the data acquisition module 10 includes:

[0167] Initial audio acquisition submodule: used to acquire multiple initial audio data collected by multiple audio collection devices;

[0168] Sound source localization submodule: used to localize the sound source of multiple initial audio data to obtain the sound source position information of each initial audio data;

[0169] Audio filtering submodule: used to determine the audio signal whose sound source is located within the first spatial range of the terminal from multiple initial audio data based on the sound source position information, and obtain audio data. The first spatial range is located in front of the display screen of the terminal and within the field of view of the image acquisition device.

[0170] In some embodiments, the apparatus further comprises:

[0171] Facial detection module: used to perform facial detection on the image data based on a multi-person detection model to obtain multiple facial images if the audio analysis result indicates that the number of people corresponding to the audio data is more than one;

[0172] Face screening module: used to screen out a target face image of a target object from multiple face images;

[0173] The facial recognition module is also used to perform facial recognition on the target facial image based on the facial recognition model to obtain target facial data.

[0174] In some embodiments, the apparatus further comprises:

[0175] Instruction sound source localization module: used for performing sound source localization on the instruction audio data carried by the facial payment voice instruction in response to the facial payment voice instruction before responding to the facial recognition payment request of the target object, and obtaining instruction sound source location information, wherein the instruction sound source location information is used to indicate the relative position of the issuing object of the instruction audio data relative to the image acquisition device;

[0176] Payment request generation module: used to generate a facial recognition payment request corresponding to a facial payment voice command if the instruction sound source location information indicates that the issuing object is located within a second spatial range, and the second spatial range is located in front of the terminal's display screen and the distance between the display screen is within a preset distance range.

[0177] In some embodiments, the face screening module includes:

[0178] The facial position acquisition submodule is used to acquire facial position information corresponding to each of the multiple facial images, where the facial position information is used to represent the position of the object to which the facial image belongs relative to the image acquisition device;

[0179] Facial image matching submodule: used to determine the facial image whose facial position information in multiple facial images matches the position information of the command sound source as the target facial image.

[0180] It should be noted that the above device embodiment and method embodiment are based on the same implementation method.

[0181] An embodiment of the present application provides a device, which may be a terminal or a server, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a payment processing method based on facial recognition as provided in the above method embodiment.

[0182] The memory can be used to store software programs and modules. The processor executes various functional applications and abnormality detection by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0183] The method embodiments provided in the embodiments of the present application can be executed in electronic devices such as mobile terminals, computer terminals, servers or similar computing devices. Figure 7 1 is a hardware structure block diagram of an electronic device for a payment processing method based on facial recognition provided by an embodiment of the present application. Figure 7 As shown, the electronic device 900 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 910 (the processor 910 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 930 for storing data, and one or more storage media 920 (such as one or more mass storage devices) for storing application programs 923 or data 922. Among them, the memory 930 and the storage medium 920 can be short-term storage or permanent storage. The program stored in the storage medium 920 may include one or more modules, each of which may include a series of instruction operations in the electronic device. Furthermore, the central processing unit 910 can be configured to communicate with the storage medium 920 and execute a series of instruction operations in the storage medium 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input and output interfaces 940, and / or, one or more operating systems 921, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM, FreeBSDTM, etc.

[0184] The input / output interface 940 may be used to receive or send data via a network. The specific example of the network may include a wireless network provided by a communication provider of the electronic device 900. In one example, the input / output interface 940 includes a network adapter (Network Interface Controller, NIC), which may be connected to other network devices via a base station so as to communicate with the Internet. In one example, the input / output interface 940 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0185] It can be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 7 More or fewer components as shown, or with Figure 7 Different configurations are shown.

[0186] An embodiment of the present application also provides a computer-readable storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to an anomaly detection method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the anomaly detection method provided by the above method embodiment.

[0187] Optionally, in this embodiment, the storage medium may be located in at least one of the multiple network servers of the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0188] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.

[0189] The payment processing method, apparatus, device, storage medium, server, terminal and program product based on facial recognition provided by the above-mentioned application,

[0190] The technical solution of the present application responds to the facial recognition payment request of the target object, obtains the image data collected by the image acquisition device and the audio data corresponding to the image data, and then performs feature analysis on the audio data for the number of people to obtain the audio analysis result. If the audio analysis result indicates that the number of people corresponding to the audio data is a single person, the image data is subjected to facial recognition based on the facial recognition model corresponding to the single-person scene to obtain the target facial data of the target object, thereby performing the payment operation corresponding to the facial recognition payment request based on the target facial data. In this way, before loading the multi-person detection module, feature analysis is first performed based on the audio data to obtain an audio analysis result that can indicate the number of people in front of the current terminal. When it indicates that the current environment is a single-person scene, facial recognition is directly performed without triggering the loading of related large-memory business modules such as multi-person detection, thereby optimizing the load time of the facial payment process, improving payment efficiency, and improving product experience.

[0191] It should be noted that the above-mentioned sequence of the embodiments of the present application is for description only and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of the present application are described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0192] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0193] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0194] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A payment processing method based on facial recognition, characterized in that, the method includes: responding to a facial recognition payment request of a target object, acquiring image data collected by an image acquisition device and audio data corresponding to the image data; performing feature analysis on the audio data for the number of persons to obtain an audio analysis result; if the audio analysis result indicates that the number of persons corresponding to the audio data is one person, performing facial recognition on the image data based on a facial recognition model corresponding to a single-person scenario to obtain target facial data of the target object; performing a payment operation corresponding to the facial recognition payment request based on the target facial data.

2. The method according to claim 1, characterized in that, the audio analysis result includes sound source separation data; the performing feature analysis on the audio data for the number of persons to obtain an audio analysis result includes: performing sound source separation detection on the audio data to obtain the sound source separation data, the sound source separation data including at least one sound source signal corresponding to the audio data; after the performing feature analysis on the audio data for the number of persons to obtain an audio analysis result, the method further includes: if the sound source separation data includes one sound source signal, determining that the audio analysis result indicates that the number of persons corresponding to the audio data is one person.

3. The method according to claim 2, characterized in that, the performing sound source separation detection on the audio data to obtain the sound source separation data includes: performing frame processing on the audio data to obtain a plurality of audio frames; performing energy feature extraction on the plurality of audio frames to obtain audio energy features; performing sound source separation corresponding to the audio energy features based on a sound source separation model, the sound source separation model being obtained by performing constrained training on a preset first deep neural network for sound source separation with sample audio energy features as input and sample sound source signals corresponding to the sample audio energy features as expected outputs, the sample audio energy features being obtained by performing energy feature extraction on audio data of a sample sound source or mixed audio data corresponding to one or more sample sound sources.

4. The method according to claim 1, characterized in that, the audio analysis result includes noise index data; the performing feature analysis on the audio data for the number of persons to obtain an audio analysis result includes: determining a noise signal in the audio data; performing noise field amplitude calculation on the noise signal to obtain noise index data, the noise index data being used to indicate the environmental noise intensity corresponding to the audio data; after the performing feature analysis on the audio data for the number of persons to obtain an audio analysis result, the method further includes: if the noise index data indicates that the environmental noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining that the audio analysis result indicates that the number of persons corresponding to the audio data is one person.

5. The method according to claim 4, characterized in that, the determining the noise signal in the audio data includes; Perform frame processing on the audio data to obtain a plurality of audio frames; Extract energy features from the plurality of audio frames to obtain audio energy features; Perform noise estimation corresponding to the audio energy features based on a noise estimation model to obtain the noise signal; The noise estimation model is obtained by performing constrained training on a preset second deep neural network for noise estimation with sample audio energy features as the input and the sample noise signal corresponding to the sample audio energy features as the expected output. The sample audio energy features are obtained by performing energy feature extraction on the audio data of a sample sound source or the mixed audio data corresponding to one or more sample sound sources.

6. The method according to claim 1, wherein, The terminal is provided with a plurality of audio acquisition devices arranged in an array. The acquisition method of the audio data includes: Acquire a plurality of initial audio data acquired by the plurality of audio acquisition devices; Perform sound source localization on the plurality of initial audio data to obtain the sound source position information of each of the initial audio data; Based on the sound source position information, determine the audio signals whose sound sources are located within the first spatial range of the terminal from the plurality of initial audio data to obtain the audio data. The first spatial range is in front of the display screen of the terminal and within the field of view of the image acquisition device.

7. The method according to claim 1, wherein, The method further includes: If the audio analysis result indicates that the number of persons corresponding to the audio data is more than one, perform face detection on the image data based on a multi-person detection model to obtain a plurality of face images; Screen out the target face image of the target object from the plurality of face images; Perform face recognition on the target face image based on the face recognition model to obtain the target face data.

8. The method according to any one of claims 1-7, wherein, Before responding to the face recognition payment request of the target object, the method further includes: In response to a face payment voice command, perform sound source localization on the command audio data carried by the face payment voice command to obtain the command sound source position information, which is used to indicate the relative position of the issuing object of the command audio data relative to the image acquisition device; If the command sound source position information indicates that the issuing object is located within a second spatial range, generate a face recognition payment request corresponding to the face payment voice command. The second spatial range is in front of the display screen of the terminal and the distance from the display screen is within a preset distance range.

9. The method according to claim 8, wherein, Screening out the target face image of the target object from the plurality of face images includes: Obtain the face position information corresponding to each of the plurality of face images, where the face position information is used to characterize the position of the object to which the face image belongs relative to the image acquisition device; Determine the face image whose face position information in the plurality of face images matches the command sound source position information as the target face image.

10. A payment processing device based on face recognition, characterized in that, the device comprises: a data acquisition module: configured to acquire image data collected by an image acquisition device and audio data corresponding to the image data in response to a face recognition payment request of a target object; a feature analysis module: configured to perform feature analysis on the audio data for the number of persons to obtain an audio analysis result; a face recognition module: configured to, if the audio analysis result indicates that the number of persons corresponding to the audio data is one person, perform face recognition on the image data based on a face recognition model corresponding to a single-person scenario to obtain target face data of the target object; a payment module: configured to perform a payment operation corresponding to the face recognition payment request based on the target face data.

11. A computer-readable storage medium, characterized in that, at least one instruction or at least one segment of program is stored in the storage medium, and the at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the face recognition-based payment processing method according to any one of claims 1-9.

12. A computer device, characterized in that, the device comprises a processor and a memory, and at least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the face recognition-based payment processing method according to any one of claims 1-9.

Citation Information

Cited By

  • Payment processing method and apparatus based on facial recognition, and device and medium

    EP4769275A1