Payment processing method and apparatus based on facial recognition, and device and medium
By performing feature analysis on audio data to determine the number of objects, directly perform facial recognition without loading multi-object detection modules, it solves the problem of high time-consuming existing payment systems and improves payment efficiency and experience.
Patent Information
- Application Number
- PCT/CN2024/115928
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-08-30
- Publication Date
- 2025-06-05
AI Technical Summary
The existing facial recognition-based payment system needs to load algorithm modules such as multi-person detection and attention mechanism before each payment, resulting in increased operation time and reduced payment efficiency and business experience.
By performing feature analysis on the audio data, the number of objects in the current environment is judged. If it is judged as a single object, facial recognition will be performed directly to avoid loading the multi-object facial detection module.
Optimizes the load time-consuming of the facial payment process, and improves payment efficiency and product experience.
Smart Images

Figure CN2024115928_05062025_PF_FP_ABST
Abstract
Description
Facial recognition-based payment processing method, device, equipment, and medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 27, 2023, with application number 202311607670.5 and application name “Payment processing method, device, equipment and medium based on facial recognition”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a payment processing method, apparatus, device, and medium based on facial recognition. Background Art
[0003] With the development of artificial intelligence, intelligent facial recognition payment has become a universal business model. Compared to traditional electronic payment methods, intelligent facial recognition eliminates the need for mobile devices or physical objects, enabling efficient payment. However, in real-world scenarios, due to the complexity of the environment, payment terminals may capture the faces of multiple subjects at once. Therefore, before each payment is made, facial frame detection and facial screening are performed by default using algorithms such as multi-person detection and attention mechanisms to avoid misidentification. While this method can effectively improve the accuracy and security of facial payment, it requires the loading of memory-intensive algorithm modules such as multi-person detection and attention mechanisms before each payment, significantly increasing operation time and reducing payment efficiency and the business experience.
[0004] Summary of the Invention
[0005] This application provides a payment processing method, apparatus, device and medium based on facial recognition, which can significantly improve the payment efficiency in facial payment services.
[0006] In one aspect, the present application provides a payment processing method based on facial recognition, the method comprising:
[0007] In response to a facial recognition payment request from a target object, acquiring image data captured by an image capture device and audio data corresponding to the image data;
[0008] Performing feature analysis on the audio data based on the number of objects to obtain audio analysis results;
[0009] If the audio analysis result indicates that the number of objects corresponding to the audio data is a single object, performing facial recognition on the image data based on the facial recognition model to obtain target facial data of the target object;
[0010] Execute the payment operation corresponding to the facial recognition payment request based on the target facial data.
[0011] Another aspect provides a payment processing device based on facial recognition, the device comprising:
[0012] Data acquisition module: used to obtain image data captured by the image acquisition device and audio data corresponding to the image data in response to the facial recognition payment request of the target object;
[0013] Feature analysis module: used to perform feature analysis on audio data based on the number of objects and obtain audio analysis results;
[0014] Facial recognition module: for performing facial recognition on the image data based on the facial recognition model to obtain target facial data of the target object if the audio analysis result indicates that the number of objects corresponding to the audio data is only one;
[0015] Payment module: used to perform payment operations corresponding to facial recognition payment requests based on target facial data.
[0016] In a possible implementation, the feature analysis module includes a sound source separation submodule configured to perform sound source separation detection on the audio data to obtain sound source separation data, the sound source separation data including at least one sound source signal corresponding to the audio data;
[0017] The sound source separation submodule is further configured to count the total number of sound source signals contained in at least one sound source signal;
[0018] The sound source separation submodule is further configured to determine that the audio analysis result is a first result if the total number of signals is equal to the unit value, the first result indicating that the number of objects corresponding to the audio data is a single object;
[0019] The sound source separation submodule is further configured to determine that the audio analysis result is a second result if the total number of signals is greater than a unit value, and the second result indicates that the number of objects corresponding to the audio data is multiple.
[0020] In a possible implementation, the sound source separation submodule includes:
[0021] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;
[0022] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;
[0023] Sound source separation unit: used to perform sound source separation on the audio energy features based on the sound source separation model to obtain sound source separation data; the sound source separation model is obtained by constrained training of a preset first deep neural network for sound source separation with the sample audio energy features as input and the sample sound source signals corresponding to the sample audio energy features as the expected output. The sample audio energy features are obtained by energy feature extraction based on the audio data of a sample sound source or the mixed audio data corresponding to at least two sample sound sources.
[0024] In a possible implementation, the feature analysis module includes:
[0025] Noise determination submodule: used to determine the noise signal in the audio data;
[0026] Amplitude calculation submodule: used to calculate the noise field amplitude of the noise signal to obtain noise index data, which is used to indicate the ambient noise intensity corresponding to the audio data;
[0027] The judgment module is further configured to: if the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determine that the audio analysis result is a first result, where the first result is used to indicate that the number of objects corresponding to the audio data is a single object;
[0028] The judgment module is also used to: if the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, determine that the audio analysis result is a second result, and the second result is used to indicate that the number of objects corresponding to the audio data is multiple.
[0029] In a possible implementation, the noise determination submodule includes:
[0030] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;
[0031] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;
[0032] Noise estimation unit: used to perform noise estimation on audio energy features based on a noise estimation model to obtain a noise signal; the noise estimation model is obtained by constrained training of a preset second deep neural network for noise estimation with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output. The sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to at least two sample sound sources.
[0033] In a possible implementation, the terminal is further provided with a plurality of audio acquisition devices arranged in an array, and the data acquisition module includes:
[0034] Initial audio acquisition submodule: used to acquire multiple initial audio data collected by multiple audio collection devices;
[0035] Sound source localization submodule: used to localize the sound source of each initial audio data and obtain the sound source position information of each initial audio data;
[0036] Audio filtering submodule: used to determine the audio signal whose sound source is located in the first spatial range of the terminal from multiple initial audio data based on the sound source position information, and obtain audio data. The first spatial range is located in front of the terminal's display screen and within the field of view of the image acquisition device.
[0037] In a possible implementation manner, the device further includes:
[0038] A face detection module is configured to perform face detection on the image data based on a face detection model to obtain multiple facial images if the audio analysis result indicates that the number of objects corresponding to the audio data is multiple;
[0039] Face screening module: used to screen out a target face image of a target object from multiple face images;
[0040] The facial recognition module is further used to perform facial recognition on the target facial image based on the facial recognition model to obtain target facial data.
[0041] In a possible implementation manner, the device further includes:
[0042] The instruction sound source localization module is used to perform sound source localization on the instruction audio data carried by the facial payment voice instruction before responding to the facial recognition payment request of the target object, and obtain instruction sound source location information. The instruction sound source location information is used to indicate the relative position of the object issuing the instruction audio data relative to the image acquisition device;
[0043] Payment request generation module: used to generate a facial recognition payment request corresponding to the facial payment voice instruction if the instruction sound source location information indicates that the issuing object is located within a second spatial range, and the second spatial range is located in front of the terminal's display screen and the distance between it and the display screen is within a preset distance range.
[0044] In a possible implementation, the facial screening module includes:
[0045] Facial position acquisition submodule: used to obtain facial position information corresponding to each of the multiple facial images, each facial position information is used to represent the position of the object to which the corresponding facial image belongs relative to the image acquisition device;
[0046] Facial image matching submodule: used to determine the facial image whose facial position information in multiple facial images matches the command sound source position information as the target facial image.
[0047] On the other hand, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the above-mentioned payment processing method based on facial recognition.
[0048] On the other hand, a computer-readable storage medium is provided, in which a computer program is stored. The computer program includes program instructions. When the program instructions are executed by a processor, the facial recognition-based payment processing method as described above is executed.
[0049] On the other hand, a server is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned payment processing method based on facial recognition.
[0050] On the other hand, a terminal is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned payment processing method based on facial recognition.
[0051] Another aspect provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above-described facial recognition-based payment processing method.
[0052] The facial recognition-based payment processing method, apparatus, device, storage medium, server, terminal, and computer program product provided in this application have the following technical effects:
[0053] After receiving the facial recognition payment request from the target object, the image data captured by the image acquisition device and the audio data corresponding to the image data can be obtained in response to the facial recognition payment request from the target object. The audio data is then subjected to feature analysis for the number of objects to obtain an audio analysis result. If the audio analysis result indicates that the number of objects corresponding to the audio data is a single object, there is no need to load the facial detection module. The image data can be directly subjected to facial recognition based on the facial recognition model to obtain the target facial data of the target object, thereby executing the payment operation corresponding to the facial recognition payment request based on the target facial data. In this way, before loading the facial detection module, a feature analysis is first performed based on the audio data to obtain an audio analysis result that can indicate the number of objects in front of the current terminal. When the audio analysis result indicates that the current environment is a scene with a single object, facial recognition is directly performed without triggering the loading of related memory-intensive business modules such as facial detection of multiple objects. This optimizes the load time of the facial payment process, improves payment efficiency, and improves the product experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] FIG1 is a schematic diagram of an application environment provided by an embodiment of the present application;
[0055] FIG2 is a flow chart of a payment processing method based on facial recognition provided in an embodiment of the present application;
[0056] FIG3 is a flow chart of another payment processing method based on facial recognition provided in an embodiment of the present application;
[0057] FIG4 is a flow chart of another payment processing method based on facial recognition provided in an embodiment of the present application;
[0058] FIG5 is a flow chart of another payment processing method based on facial recognition provided in an embodiment of the present application;
[0059] FIG6 is a schematic diagram of a payment processing device based on facial recognition according to an embodiment of the present application;
[0060] FIG7 is a hardware structure block diagram of a computer device for executing a payment processing method based on facial recognition provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0062] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server comprising a series of steps or submodules is not necessarily limited to those steps or submodules clearly listed, but may include other steps or submodules that are not clearly listed or inherent to these processes, methods, products or devices.
[0063] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0064] 3D camera: Similar to traditional cameras, it adds living body-related hardware and software, including a depth camera and an infrared camera, to ensure information security.
[0065] SN: string serial number, which can uniquely identify a device ID.
[0066] SQLite: is a lightweight database and a relational database management system that complies with ACID.
[0067] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0068] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0069] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0070] Deep learning: The concept of deep learning originates from the study of artificial neural networks. A multilayer perceptron with multiple hidden layers is an example of a deep learning architecture. Deep learning discovers distributed feature representations of data by combining lower-level features to form more abstract higher-level representations of attribute categories or features.
[0071] Computer vision (CV) is the science of making machines "see." Specifically, it refers to the use of cameras and computers to replace the human eye in identifying, detecting, and measuring objects, and further image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0072] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0073] Please refer to Figure 1, which is a schematic diagram of an application environment provided by an embodiment of the present application. As shown in Figure 1, the application environment may include at least a terminal 01 and a server 02. In actual applications, the terminal 01 and the server 02 may be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions thereon.
[0074] The server 02 in the embodiment of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0075] Specifically, cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology can be applied in a variety of fields, such as healthcare cloud, cloud IoT, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calling, and cloud social networking. Based on the cloud computing business model, cloud technology distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides resources is called the "cloud." To users, the resources in the cloud appear infinitely scalable and can be accessed at any time, used on demand, and expanded at any time, with a pay-per-use policy. Providers of cloud computing infrastructure establish a cloud computing resource pool (referred to as a cloud platform, commonly referred to as IaaS (Infrastructure as a Service)) and deploy various types of virtual resources within the resource pool for external clients to choose from. The cloud computing resource pool primarily includes computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0076] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging apps. Generally speaking, SaaS and PaaS are upper layers relative to IaaS.
[0077] Specifically, the server 02 mentioned above may include a physical device, which may specifically include a network communication submodule, a processor, a memory, etc., and may also include software running in the physical device, which may specifically include an application program, etc.
[0078] Specifically, terminal 01 may include physical devices such as smart phones, desktop computers, tablet computers, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices, and may also include software running in physical devices, such as applications.
[0079] In an embodiment of the present application, terminal 01 can be used to generate a facial recognition payment request in response to a facial recognition payment-related trigger operation submitted by a target object, and then collect image data and audio data to perform feature analysis on the audio data based on the number of objects to obtain an audio analysis result. Wherein, the object here can refer to a user or an intelligent robot, and the number of objects can refer to the number of objects contained in the audio data. The feature analysis of the audio data based on the number of objects can be performed on the audio data to analyze the number of objects contained in the audio data. The obtained audio analysis result is also a result that can reflect the number of objects. Based on the number of objects corresponding to the audio data indicated by the audio analysis result, it can be determined whether to trigger and load a facial detection model (here, the facial detection model can refer to a model for detecting the faces of multiple objects, also known as a multi-person detection model) to perform facial detection and target facial image screening of the target object, and send a related service payment request to server 02 based on the target facial data finally obtained, so that server 02 performs corresponding payment processing / payment operation in response to the service payment request. Server 02 provides payment functions and data storage functions, such as using a SQLite database to store facial data for payment comparison and matching.
[0080] In addition, it can be understood that what is shown in FIG1 is merely an application environment of a payment processing method based on facial recognition. The application environment may include more or fewer nodes, and this application does not impose any limitation thereto.
[0081] The application environment involved in the embodiments of the present application, or the terminal 01 and server 02 in the application environment, etc., can be a distributed system formed by connecting a client and multiple nodes (any form of computing device connected to the network, such as a server, a terminal) through network communication. The distributed system can be a blockchain system, which can provide the above-mentioned payment processing function based on facial recognition, model training function and data storage function. In one embodiment, the application environment can run a dialogue system, which can store basic corpus information and include a role setting component, a question and answer component (QA component), an expression processing component and a reply generation component (see Figure 7).
[0082] The following introduces the technical solution of the present application based on the above-mentioned application environment and dialogue system. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. Please refer to Figure 2. Figure 2 is a flow chart of a payment processing method based on facial recognition provided by an embodiment of the present application. This specification provides method operation steps such as the embodiments or flow charts, but more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, a parallel processor or a multi-threaded processing environment). Specifically, as shown in Figure 2, the following steps S201-S207 may be included:
[0083] S201: In response to a facial recognition payment request from a target object, image data captured by an image capture device and audio data corresponding to the image data are acquired.
[0084] Specifically, terminal 01 can be provided with a display screen, an image acquisition device and an audio acquisition device. The display screen can be used to display a facial acquisition frame, etc., so as to display payment-related operation information and interfaces to the target object (the target object in this application can refer to any object). The field of view of the image acquisition device is located in front of the display screen, and is used to collect image data directly in front of the display screen for payment processing based on facial recognition. There can be multiple audio acquisition devices, such as multiple audio acquisition devices can be arranged in an array, for audio data acquisition and corresponding sound source positioning.
[0085] In some embodiments, to enhance the accuracy and security of facial recognition of an object, the image acquisition device may include a 3D camera, and the acquired image data may include, in addition to color images (such as RGB images), images including depth information such as depth maps, and may also include infrared images, etc.
[0086] In an embodiment of the present application, the terminal may run a facial recognition application, which may specifically include a facial payment component, an algorithm auxiliary component, and an algorithm scheduling component. The facial payment component is used for facial acquisition and image screening. The algorithm auxiliary component runs a facial detection component and an attention detection component, etc., to achieve multi-object facial detection. The algorithm scheduling component is used to perform facial recognition on the acquired facial image of the target object to obtain facial data for payment processing. Facial acquisition refers to calling an image acquisition device to perform image acquisition. It can be understood that the acquired image stream data is image stream data, such as RGB image stream, depth image stream, and infrared image stream. Image screening is used to filter out image data for subsequent facial recognition from the image stream data. Specifically, it can be based on a comprehensive evaluation of parameters such as facial area size, facial angle, image contrast, image brightness and clarity to determine the optimal image data containing facial images. Existing facial recognition payment solutions will introduce multi-object facial detection components and attention detection components before the image screening stage and the subsequent facial recognition stage. Through the multi-object facial detection algorithm and combined with the attention detection mechanism, it is determined whether there are facial image areas of multiple objects in the current image. In actual scenarios, there are payment scenarios with a single object, that is, there are no faces of multiple objects in front of the terminal. The use of the above-mentioned existing solutions will still load the multi-object facial detection component and attention detection component. The algorithm parameters of the multi-object facial detection component and the attention detection component are large, and the model occupies a high memory. Loading and detection are time-consuming, resulting in redundant resource occupation and affecting the facial payment experience. The facial payment component of the present application also runs a number detection component for the number of objects, which is used to detect the number of objects in front of the terminal.
[0087] It can be understood that there may be multiple objects in front of the terminal. The target object refers to the object that initiates facial payment for the current terminal. The facial recognition payment request refers to the request generated by the terminal based on the facial payment trigger operation of the target object. The request is used to request to execute payment by recognizing the face. In other words, facial recognition payment refers to the process of recognizing the face and determining whether to execute payment based on the recognition result. The triggering operation of facial payment may include but is not limited to voice command triggering, button submission, etc.
[0088] Specifically, the audio data can be mixed audio collected by the audio collection device, including sound signals in the environment surrounding the terminal. The audio data corresponding to the image data refers to the audio stream that matches the image collection time of the image data, that is, audio that contains the image collection time and the time difference between the audio collection time and the image collection time is within a preset time difference range.
[0089] In some embodiments, the terminal is provided with multiple audio collection devices arranged in an array, such as a microphone array, preferably a linear microphone array, in which multiple microphones are arranged in a straight line at a certain interval. Accordingly, referring to FIG3 , S201 may include S301-S305:
[0090] S301: Acquire multiple initial audio data collected by multiple audio collection devices;
[0091] S303: Perform sound source localization on each initial audio data to obtain the sound source position information of each initial audio data;
[0092] S305: Determine, from the plurality of initial audio data, based on the sound source position information, an audio signal whose sound source is located within a first spatial range of the terminal, and obtain audio data.
[0093] Specifically, each initial audio data is data obtained by pre-processing the original audio collected by the audio acquisition device. The pre-processing may include but is not limited to filtering, noise reduction, gain control, etc. Sound source localization refers to the technology of determining the location of the sound source by analyzing the propagation characteristics of sound waves and the received audio signal. It should be understood that multiple audio acquisition devices (such as microphones) can be distributed in different spatial positions, and the received audio signals have different time delays and amplitude differences. Based on the time delays and amplitude differences between each audio signal, its sound source position can be calculated. The sound source position information here is used to characterize the relative position between the sound source of the initial audio data and the audio acquisition device, and is used to determine the orientation of the sound source of the initial audio data relative to the terminal.
[0094] Specifically, a sound source localization algorithm is used to process each initial audio data obtained from the preprocessing to determine the location of its sound source and obtain sound source location information. The sound source localization algorithm used may include, but is not limited to, beamforming algorithms, covariance matching algorithms, and least squares methods, and is not specifically limited here. Furthermore, a sound source localization and tracking algorithm may be used to track the sound source to maintain the positional relationship between the audio acquisition device and the sound source. The sound source localization and tracking algorithm may include, but is not limited to, Kalman filtering, particle filtering, and extended Kalman filtering, and is not specifically limited here.
[0095] Specifically, the first spatial range is located in front of the terminal's display screen and within the field of view of the image acquisition device (the field of view here may refer to the range that the image acquisition device can capture). That is, the first spatial range is limited to using only the initial audio data generated by the sound source directly in front of the display screen (or image acquisition device). It can be understood that the image acquisition device is facing the front of the terminal display screen, and the above scheme accurately filters out the audio data originating from directly in front of the terminal to reduce interference from stray sound sources, reduce the difficulty of sound source separation, and improve the efficiency of subsequent object location.
[0096] In some embodiments, the sound source position information can be used to directly filter the audio data, and the sound source direction corresponding to the sound source position information and the camera direction can be matched to determine the angular difference between the sound source direction and the camera direction. The audio data can be filtered through the angle difference. For example, if the angular difference of a certain sound source position information is less than a certain angle threshold (the angle threshold can be specifically set based on actual business needs), then the sound source can be considered to be located directly in front of the image acquisition device, i.e., the display screen, and its corresponding audio signal can be determined to be within the first spatial range, and it can be determined as the audio data in this application; otherwise, it can be determined that it is not within the first spatial range and can be discarded. In other embodiments, sound source separation processing can be used in combination with the aforementioned sound source position information to perform sound source separation processing on multiple initial audio data to separate the sound signal directly in front and obtain audio data. The sound source separation algorithm used can be, but is not limited to, blind source separation, non-blind source separation, and separation based on a sound source separation model.
[0097] S203: Perform feature analysis on the audio data based on the number of objects to obtain an audio analysis result.
[0098] Specifically, audio analysis results are used to characterize the sound source index characteristics of the audio data, and may include, but are not limited to, at least one of sound source separation data and noise index data. The number of objects refers to the number of objects to which the sound sources that form the audio data belong, that is, the number of objects that emit the sound sources that form the audio data.
[0099] Specifically, feature analysis analyzes the audio energy features of audio data to obtain an audio analysis result. The audio analysis result refers to a result used to reflect the number of objects in the audio data, which may include a first result and a second result. The first result can be used to indicate that the number of objects in the audio data is single, and the second result can be used to characterize the number of objects in the audio data as multiple (multiple generally refers to two or more). The audio energy feature is used to characterize the energy of each frame of audio data in the time domain, that is, to describe the characteristics of audio energy, including but not limited to short-time energy, short-time average amplitude, short-time power spectrum, etc.
[0100] In some embodiments, S203 includes S401: performing sound source separation detection on the audio data to obtain sound source separation data.
[0101] Specifically, sound source separation detection refers to separating different sound sources in audio data to isolate different sound source signals from a piece of audio data. Accordingly, the sound source separation data includes at least one sound source signal corresponding to the audio data. It can be understood that if only a single sound source signal exists, it indicates that the current audio data originates from the same object, and the scene emitted by the audio data is a single-object scene. If there are more than one (at least two) sound source signals, it can be determined that the audio data originates from multiple objects, indicating that it hits a multi-source, multi-object scene.
[0102] In some cases, when sound source separation processing is used in the aforementioned S305 to separate the sound signal directly in front and obtain audio data, the sound source separation data in S401 can reuse the result obtained by the sound source separation processing in S305, thereby reducing the amount of data processing. In other cases, sound source separation processing is not used in the aforementioned S305, and the sound source separation algorithm used in the sound source separation detection in S401 may include blind source separation, non-blind source separation, and separation based on a sound source separation model. The above-mentioned sound separation data is obtained by performing sound source separation processing on the audio energy features, and the audio energy features are obtained by performing frame processing on the audio data to obtain multiple audio frames, and extracting energy features from the multiple audio frames.
[0103] Accordingly, in one embodiment, S401 includes: performing frame processing on the audio data to obtain multiple audio frames; performing energy feature extraction on the multiple audio frames to obtain audio energy features, that is, obtaining energy features of each audio frame, such as short-time energy, short-time average amplitude, short-time power spectrum, etc.; further, performing sound source separation based on the audio energy features of the sound source separation model to obtain sound source separation data. In this way, using the intelligent model to perform sound source separation can improve the accuracy of sound source separation.
[0104] Specifically, the sound source separation model mentioned above is obtained by constrained training of the preset first deep neural network for sound source separation with the sample audio energy feature as input and the sample sound source signal corresponding to the sample audio energy feature as the expected output. The sample audio energy feature is obtained by extracting energy features based on the audio data of a sample sound source or the mixed audio data corresponding to at least two sample sound sources. The audio data of the sample sound source can be human voice audio, and the mixed audio data can refer to audio including multiple human voices. The sample sound source signal serves as a training label, which refers to the sound signal emitted by a single sample sound source, and is used to optimize the training of the preset deep neural network (such as the first deep neural network). It can be understood that the sample audio energy feature here is similar to the method of obtaining the aforementioned audio energy feature, and will not be elaborated on.
[0105] In one embodiment, an audio data set containing multiple human voices is obtained, and the voices of different human voices are separated separately as audio data of sample sound sources, and then the audio data of the sample sound source is converted into a computer-processable feature vector, and the methods used include but are not limited to short-time Fourier transform (STFT), Mel-frequency cepstral coefficients (MFCC) and energy feature algorithms. An initial model is constructed using a preset deep neural network (such as a convolutional neural network (CNN), etc.), and then the initial model is trained. During the training process, optimizers such as Adam or SGD can be used for optimization, and the learning rate can be set based on the needs and the existing loss function can be used for loss calculation, such as cross entropy. The audio data to be separated is feature processed and input into the trained sound source separation model, and the sound source signals of different human voices predicted by the model are output. In the case of long audio data, long audio data can be processed by sliding window.
[0106] Accordingly, after obtaining at least one sound source signal corresponding to the audio data, an audio analysis result may be determined based on the at least one sound source signal. The specific implementation process may include at least the following steps S501 to S503:
[0107] S501: Count the total number of sound source signals included in at least one sound source signal.
[0108] S502: If the total number of signals is equal to the unit value, determine that the audio analysis result is a first result, and the first result indicates that the number of objects corresponding to the audio data is single.
[0109] Specifically, the unit value can be 1. If the total number of signals is equal to the unit value, it can indicate that the sound source separation data includes one sound source signal. At this time, the audio analysis result can be determined to be the first result, that is, the number of objects corresponding to the audio data is single.
[0110] S503: If the total number of signals is greater than the unit value, determine that the audio analysis result is a second result, and the second result indicates that the number of objects corresponding to the audio data is multiple.
[0111] Specifically, if the total number of signals is greater than a unit value, it may indicate that the sound source separation data includes at least two sound source signals. In this case, the audio analysis result may be determined to be the second result, that is, the number of objects corresponding to the audio data is multiple.
[0112] In this way, the energy characteristics are analyzed through the intelligent model to achieve modeling and separation of the sound characteristics of different objects, and the audio data is divided into different sound sources. The number of separated different sound source signals is then the different objects, and the number of objects can be obtained. If there are more than one, it means that there may be multiple objects in the surrounding area. In this case, relevant business components such as the multi-object facial detection module can be loaded to detect multiple facial images; if there is equal to one, there is no need to call the multi-object facial detection module and other related business components. The facial recognition model can be directly called to perform facial recognition and payment processing on the facial data of the target object to improve payment efficiency.
[0113] In some other embodiments, S203 includes S403-S405:
[0114] S403: Determine a noise signal in the audio data;
[0115] S405: Calculate the noise field amplitude of the noise signal to obtain noise index data, which is used to indicate the intensity of the ambient noise corresponding to the audio data, such as the noise decibel value.
[0116] Specifically, the noise signal refers to the noise audio separated from the audio data. Methods for obtaining the noise signal include, but are not limited to, framing the audio data to obtain multiple audio frames; after obtaining the multiple audio frames, extracting energy features from the multiple audio frames to obtain audio energy features; and further, performing noise estimation on the obtained audio energy features based on a noise estimation model to obtain the noise signal. It is understood that the noise signal acquisition process and the audio frames and audio energy features used in the sound source separation process can be reused.
[0117] Specifically, the noise estimation model is obtained by constrained training of a preset second deep neural network for noise estimation with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as the expected output. The sample audio energy features are obtained by energy feature extraction based on audio data of a sample sound source or mixed audio data corresponding to at least two sample sound sources.
[0118] Understandably, the sample noise signal is equivalent to a training label used to implement model training. The sample audio energy features, sample sound source audio data, and mixed audio data here can be shared with the training data of the sound source separation model, or they can be separate training data. The second deep neural network can be, for example, a convolutional neural network. In this way, noise signals are generated through intelligent models to assist in determining environmental noisiness, thereby improving the accuracy of multi-person scene judgment.
[0119] The extracted features are analyzed using a noise estimation model to obtain noise energy distribution characteristics, such as average energy and peak energy, to estimate the noise and generate a noise signal. Specifically, noise field amplitude calculation refers to calculating the noise field decibel value of the noise signal using a noise intensity algorithm.
[0120] Accordingly, after the noise index data is calculated, the audio analysis result of the audio data can be determined based on the ambient noise intensity corresponding to the audio data indicated by the noise index data. The specific implementation process can include at least the following steps S505-S507:
[0121] S505: If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining the audio analysis result as a first result, where the first result is used to indicate that the number of objects corresponding to the audio data is a single object;
[0122] S507: If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, the audio analysis result is determined to be a second result, where the second result is used to indicate that the number of objects corresponding to the audio data is multiple.
[0123] It is understandable that when the number of objects around the terminal is too large, high-decibel noise will lead to inaccurate sound separation. Therefore, this application introduces noise signal detection and noise intensity index calculation as separate judgment logic. If the decibel value exceeds the preset intensity threshold (the threshold can be set based on actual business needs), it can be judged that the decibel value of the noise field is unqualified, indicating that there may be multiple objects around. At this time, it is necessary to load the multi-object facial detection component. If it is lower than the preset intensity threshold, it can be determined as a single-object scene. At this time, there is no need to load the multi-object facial detection component. In this way, the judgment accuracy of multi-object scenes and single-object scenes can be improved through noise detection.
[0124] In some embodiments, sound source separation detection and noise field amplitude calculation of the noise signal can be performed in parallel. After obtaining the audio data, the audio data can be subjected to sound source separation detection and noise signal estimation and corresponding noise amplitude calculation, thereby obtaining the sound source separation result and noise index data respectively. In some cases, if any of the sound source separation results and noise index data indicates that the number of objects is multiple, execution of S209 can be triggered, and if both the sound source separation results and noise index data indicate that the number of objects is single, execution of S205 can be triggered. In other cases, if any of the sound source separation results and noise index data indicates that the number of objects is single, execution of S205 can be triggered, and if both the sound source separation results and noise index data indicate that the number of objects is multiple, execution of S209 can be triggered. In this way, through parallel detection, the two methods complement each other to achieve accurate discrimination between multi-person and single-person scenes.
[0125] In other embodiments, after obtaining the audio data, a sound source separation test may be performed first. If the sound source separation data includes more than one sound source signal, it may be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, and execution of S209 may be triggered. If the sound source separation data includes one sound source signal, noise signal estimation and noise field amplitude calculation may be performed on the audio data. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, it may be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is single, and execution of S205 may be triggered. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, it may be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, and execution of S209 may be triggered. In this way, by selectively triggering noise detection, detection accuracy can be ensured while reducing the resource usage of object number detection.
[0126] In other embodiments, after obtaining audio data, noise signal estimation and noise field amplitude calculation can be performed first. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, it can be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, and execution of S209 can be triggered. If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than the preset intensity threshold, it can be determined that the noise index result indicates a single-object scene. Further, sound source separation data of the audio data can be obtained. If the sound source separation data includes one sound source signal, it can be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is single, and execution of S205 can be triggered. If the sound source separation data includes more than one sound source signal, it can be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, and execution of S209 can be triggered. In this way, by selectively triggering sound source separation detection, detection accuracy can be ensured while reducing resource usage for object number detection.
[0127] It can be understood that audio energy features are used in both noise signal estimation and sound source separation detection. Therefore, in the above-mentioned various schemes, frame processing and energy feature extraction are only performed once, that is, frame processing is first performed to obtain multiple audio frames, and then energy features of the multiple audio frames are extracted to obtain audio energy features, and the audio energy features are used for noise signal estimation and sound source separation processing.
[0128] This solution is based on noise and sound source detection, pre-detection, and identification of the complexity of the objects contained in the current scene. If the scene does not contain excessive noise or related human voice conversations, it is likely a single-object payment scenario. In this case, the algorithm can directly perform facial recognition without loading the relevant capabilities of multi-object facial detection. If there is excessive noise or unnecessary human voice conversations, the multi-object facial detection component is loaded to improve payment efficiency and reduce resource loss.
[0129] S205: If the audio analysis result indicates that the number of objects corresponding to the audio data is single, facial recognition is performed on the image data based on a facial recognition model to obtain target facial data of the target object.
[0130] In some embodiments, if the result based on noise signal detection is that the number of objects in the audio data is a single object or the result based on sound source separation detection is a single object, it can be confirmed that the audio analysis result indicates that the number of objects corresponding to the audio data is a single object, and then the algorithm scheduling component can directly call the facial recognition model to perform facial recognition processing to obtain the target facial data without loading the multi-object facial detection algorithm and the corresponding attention detection component, etc. The target facial data can be a feature vector representing the facial image area of the target object and has object specificity. The facial recognition model can be constructed using existing recognition models and trained with facial recognition training data, and there is no limitation here. If the result based on noise signal detection is more than one person and the result based on sound source separation detection is more than one person, it can be confirmed that the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, and then the multi-object facial detection algorithm can be called accordingly to perform multi-object facial detection.
[0131] In other embodiments, if the results obtained based on noise signal detection and sound source separation detection both indicate that the number of objects is single, then it can be determined that the audio analysis result indicates that the number of objects corresponding to the audio data is single, that is, it is clear that the current payment scenario is a single-object payment scenario, and facial recognition can be performed directly without loading a multi-object facial detection algorithm and corresponding attention detection components, etc. If the result based on noise signal detection is more than one person or the result based on sound source separation detection is more than one person, it can be confirmed that the audio analysis result indicates that the number of objects corresponding to the audio data is more than one, and then the multi-object facial detection algorithm can be called accordingly to perform multi-object facial detection. In this way, the XOR judgment of noise and sound source separation can be achieved, the accuracy of single-object payment scenario recognition can be improved, and the missed detection rate of multi-object payment scenarios can be reduced.
[0132] S207: Execute the payment operation corresponding to the facial recognition payment request based on the target facial data.
[0133] It is understood that the facial recognition model can be run on the terminal, and the terminal can then send the obtained target facial data to the server for facial data matching operations to enable payment operations. Alternatively, the facial recognition model can be run on the server. After the terminal determines a single-object payment scenario, it can send the filtered image data to the server for input into the facial recognition model and output target facial data for facial data matching to enable payment operations.
[0134] Based on the above technical solution, before loading the multi-object facial detection module, this application can first perform feature analysis based on audio data to obtain an audio analysis result that can indicate the number of objects in front of the current terminal. When it indicates that the current environment is a single-object payment scenario, facial recognition can be performed directly without triggering the loading of multi-object facial detection and other related memory-intensive business components. This can optimize the load time of the facial payment process, improve payment efficiency, and improve product experience.
[0135] Specifically, after receiving image data uploaded from a terminal, the facial recognition component extracts features from the current image data to obtain target facial data. This data is then compared with features stored in a database. The feature data with the highest matching score is then compared with facial data in the backend database to determine whether the target party's payment request is compliant and the corresponding payment information. If compliant, payment is then executed. For example, after determining compliance, the downstream server returns the target party's final payment account information or payment code information for payment interaction verification.
[0136] Based on some or all of the above embodiments, in the embodiment of the present application, referring to FIG4 , the method further includes S209-S213:
[0137] S209: If the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, perform face detection on the image data based on the face detection model to obtain multiple facial images;
[0138] S211: Filtering a target facial image of a target object from a plurality of facial images;
[0139] S213: Perform facial recognition on the target facial image based on the facial recognition model to obtain target facial data.
[0140] It can be understood that the facial detection model may refer to a facial detection model for multiple objects. An existing detection model may be used. It may be constructed based on a deep neural network and trained using facial images of multiple objects for image detection. It may be run on a terminal to obtain multiple facial images directly at the terminal, or it may be run on a server. The terminal sends image data to the server and carries detection trigger information of multiple objects, so that the server calls the multi-object detection model to perform facial detection on the image data. After obtaining multiple facial images, the image of the target object needs to be screened. The screening process can be performed at the terminal or by the server. In this way, the detection of multi-object payment scenarios is achieved by pre-judging the number of objects. When it is determined that there are multiple objects, the multi-object facial detection model is loaded, and then the target facial image to be identified is obtained through facial screening, thereby improving the overall payment efficiency while ensuring payment security.
[0141] In some cases, facial images can be screened based on information such as the image area size and image area position corresponding to multiple facial images in the image data. It can be understood that the larger the image area size and the closer the image area position is to the image center, the higher the probability that it is the target facial data. The image area size, image area position and other information can be scored based on a voting scoring method, and then the scores can be added together to determine the facial image with the highest score as the target facial image.
[0142] Based on some or all of the above embodiments, in some embodiments, facial payment scenarios may be superimposed with voice commands, that is, facial recognition payment requests may be triggered by voice commands. Accordingly, referring to FIG. 5 , before S201 , the method further includes S101 - S103:
[0143] S101: In response to a facial payment voice command, performing sound source positioning on the instruction audio data carried by the facial payment voice command to obtain instruction sound source location information;
[0144] S103: If the instruction sound source location information indicates that the issuing object is located within the second spatial range, a facial recognition payment request corresponding to the facial payment voice instruction is generated.
[0145] Specifically, after the audio acquisition device of the terminal collects the original audio data and obtains the initial audio data, it can perform voice recognition with it. If the voice recognition result shows that the corresponding voice text hits the facial payment voice instruction, such as "face recognition payment", the position of the audio signal corresponding to the voice text in the initial audio data is located through the sound source localization algorithm to obtain the instruction sound source position information. The instruction sound source position indicates the relative position of the issuing object of the instruction audio data relative to the image acquisition device, and then the relative position of the issuing object relative to the display screen is determined based on the position information of the image acquisition device, the position information of the display screen and the instruction sound source position information. If the determined relative position is within the second spatial range, the instruction sound source position information indicates that the issuing object is within the second spatial range. Otherwise, it is determined that the position of the issuing object exceeds the second spatial range. The sound source localization algorithm here can be similar to the sound source localization processing in the previous article, and will not be repeated here.
[0146] If the instruction sound source location information indicates that the issuing object is located within the second spatial range, the object is determined to be the target object for initiating the facial payment operation of the current terminal, thereby generating a facial recognition payment request. The second spatial range is located in front of the terminal's display screen and the distance between it and the display screen is within a preset distance range. If the instruction sound source location information indicates that the issuing object is outside the second spatial range, it means that the issuing object is not the target object, and then no facial recognition payment request is generated, but the sound signals around the terminal are continuously collected. In this way, in the case of multiple payment terminals, crosstalk from voice commands to other terminals is avoided, and the accuracy of voice triggering is improved.
[0147] Accordingly, in the scenario where a voice command triggers a facial payment operation, if it is determined based on the aforementioned audio analysis result that there are multiple objects, then S211 filters out the target facial image of the target object from the multiple facial images, including:
[0148] S601: Obtain facial position information corresponding to each of a plurality of facial images;
[0149] S603: Determine a facial image, among the multiple facial images, whose facial position information matches the command sound source position information, as a target facial image.
[0150] Specifically, the facial position information is used to represent the position of the object to which the facial image belongs relative to the image acquisition device. It can be obtained by calculating the relative position of the object and the image acquisition device based on the image area position of the facial image in the image data and the camera parameters used to capture the image data. The image area position can be obtained based on the facial detection in S209.
[0151] Furthermore, the facial position information in each facial image is compared with the position of the command sound source, and the facial image with the closest position is determined as the target facial image. In this way, the target facial image is determined in combination with the positioning results of the voice command, improving the efficiency of target determination in multi-object payment scenarios.
[0152] In some embodiments, in order to avoid comparison errors caused by changes in the position of the target object in the time interval between the issuance of the voice command and the facial detection of multiple objects, after determining that a facial recognition payment request has been generated, the terminal generates a prompt sound for re-entering the voice signal, such as "Repeat facial payment command", to prompt the target object to re-enter the voice command, and then separate the sound source of the voice command to update the command sound source location information as comparison data for the facial position information, thereby improving the accuracy of the position comparison.
[0153] This application builds an audio acquisition device into the payment terminal device, which continuously collects the sound source directly in front of the device in silent mode, and performs timbre and noise judgment on the sound source. If multiple timbres or high noise are hit, the multi-person detection and attention judgment module will be additionally loaded when starting facial recognition payment. Otherwise, it will not be loaded, thereby optimizing the time consumption caused by the algorithm load and improving the business experience.
[0154] An embodiment of the present application also provides a facial recognition-based payment processing device 700, which is arranged in a terminal, as shown in Figure 6. Figure 6 shows a structural schematic diagram of a facial recognition-based payment processing device provided in an embodiment of the present application. The device may include the following modules.
[0155] Data acquisition module 10: for acquiring image data captured by an image acquisition device and audio data corresponding to the image data in response to a facial recognition payment request from a target object;
[0156] Feature analysis module 20: used to perform feature analysis on the audio data based on the number of objects to obtain audio analysis results;
[0157] Facial recognition module 30: configured to perform facial recognition on the image data based on a facial recognition model corresponding to a single-person scene to obtain target facial data of the target object if the audio analysis result indicates that the number of objects corresponding to the audio data is a single person;
[0158] Payment module 40: used to perform the payment operation corresponding to the facial recognition payment request based on the target facial data.
[0159] In some embodiments, the feature analysis module 20 includes a sound source separation submodule: configured to perform sound source separation detection on the audio data to obtain sound source separation data, the sound source separation data including at least one sound source signal corresponding to the audio data;
[0160] The sound source separation submodule is further configured to count the total number of sound source signals contained in at least one sound source signal;
[0161] The sound source separation submodule is further configured to determine that the audio analysis result is a first result if the total number of signals is equal to the unit value, the first result indicating that the number of objects corresponding to the audio data is a single object;
[0162] The sound source separation submodule is further configured to determine that the audio analysis result is a second result if the total number of signals is greater than a unit value, and the second result indicates that the number of objects corresponding to the audio data is multiple.
[0163] In some embodiments, the sound source separation submodule includes:
[0164] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;
[0165] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;
[0166] Sound source separation unit: used to perform sound source separation on the audio energy features based on the sound source separation model to obtain sound source separation data; the sound source separation model is obtained by constrained training of a preset first deep neural network for sound source separation with the sample audio energy features as input and the sample sound source signals corresponding to the sample audio energy features as the expected output. The sample audio energy features are obtained by energy feature extraction based on the audio data of a sample sound source or the mixed audio data corresponding to at least two sample sound sources.
[0167] In some embodiments, the feature analysis module 20 includes:
[0168] Noise determination submodule: used to determine the noise signal in the audio data;
[0169] Amplitude calculation submodule: used to calculate the noise field amplitude of the noise signal to obtain noise index data, which is used to indicate the ambient noise intensity corresponding to the audio data;
[0170] The judgment module is further configured to: if the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determine that the audio analysis result is a first result, where the first result is used to indicate that the number of objects corresponding to the audio data is a single object;
[0171] The judgment module is also used to: if the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, determine that the audio analysis result is a second result, and the second result is used to indicate that the number of objects corresponding to the audio data is multiple.
[0172] In some embodiments, the noise determination submodule includes:
[0173] Framing unit: used to perform framing processing on audio data to obtain multiple audio frames;
[0174] Feature extraction unit: used to extract energy features from multiple audio frames to obtain audio energy features;
[0175] Noise estimation unit: used to perform noise estimation on audio energy features based on a noise estimation model to obtain a noise signal; the noise estimation model is obtained by constrained training of a preset second deep neural network for noise estimation with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output. The sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to at least two sample sound sources.
[0176] In some embodiments, the terminal is further provided with a plurality of audio acquisition devices arranged in an array, and the data acquisition module 10 includes:
[0177] Initial audio acquisition submodule: used to acquire multiple initial audio data collected by multiple audio collection devices;
[0178] Sound source localization submodule: used to localize the sound source of each initial audio data and obtain the sound source position information of each initial audio data;
[0179] Audio filtering submodule: used to determine the audio signal whose sound source is located in the first spatial range of the terminal from multiple initial audio data based on the sound source position information, and obtain audio data. The first spatial range is located in front of the terminal's display screen and within the field of view of the image acquisition device.
[0180] In some embodiments, the apparatus further comprises:
[0181] A face detection module is configured to perform face detection on the image data based on a face detection model to obtain multiple facial images if the audio analysis result indicates that the number of objects corresponding to the audio data is multiple;
[0182] Face screening module: used to screen out a target face image of a target object from multiple face images;
[0183] The facial recognition module is further used to perform facial recognition on the target facial image based on the facial recognition model to obtain target facial data.
[0184] In some embodiments, the apparatus further comprises:
[0185] The instruction sound source localization module is used to perform sound source localization on the instruction audio data carried by the facial payment voice instruction before responding to the facial recognition payment request of the target object, and obtain instruction sound source location information. The instruction sound source location information is used to indicate the relative position of the object issuing the instruction audio data relative to the image acquisition device;
[0186] Payment request generation module: used to generate a facial recognition payment request corresponding to the facial payment voice instruction if the instruction sound source location information indicates that the issuing object is located within a second spatial range, and the second spatial range is located in front of the terminal's display screen and the distance between it and the display screen is within a preset distance range.
[0187] In some embodiments, the facial screening module includes:
[0188] Facial position acquisition submodule: used to obtain facial position information corresponding to each of the multiple facial images, each facial position information is used to represent the position of the object to which the corresponding facial image belongs relative to the image acquisition device;
[0189] Facial image matching submodule: used to determine the facial image whose facial position information in multiple facial images matches the command sound source position information as the target facial image.
[0190] It should be noted that the above device embodiments and method embodiments are based on the same implementation method.
[0191] An embodiment of the present application provides a device, which may be a terminal or a server, including a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement a facial recognition-based payment processing method as provided in the above-mentioned method embodiment.
[0192] The memory can be used to store software programs and modules. The processor executes various functional applications and anomaly detection by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for functions, etc.; the data storage area can store data created based on the use of the device, etc. In addition, the memory can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.
[0193] The method embodiments provided in the embodiments of the present application can be executed in a computer device such as a mobile terminal, a computer terminal, a server, or a similar computing device. Figure 7 is a hardware block diagram of a computer device for a facial recognition-based payment processing method provided in an embodiment of the present application. As shown in Figure 7, the computer device 900 may vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 910 (the processor 910 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 930 for storing data, and one or more storage media 920 (e.g., one or more mass storage devices) for storing application programs 923 or data 922. The memory 930 and storage medium 920 may be either ephemeral or persistent storage. The program stored in the storage medium 920 may include one or more modules, each of which may include a series of instruction operations on the computer device. Furthermore, the CPU 910 may be configured to communicate with the storage medium 920 to execute the series of instruction operations in the storage medium 920 on the computer device 900. The computer device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input and output interfaces 940, and / or one or more operating systems 921, such as Windows Server 2003 or Windows Server 2008. TM , Mac OS X TM , Unix TM , LinuxTM, FreeBSDTM, etc.
[0194] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communications provider of the computer device 900. In one embodiment, the input / output interface 940 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the input / output interface 940 can be a radio frequency (RF) module for wireless communication with the Internet.
[0195] Those skilled in the art will appreciate that the structure shown in FIG7 is merely illustrative and does not limit the structure of the electronic device. For example, the computer device 900 may include more or fewer components than those shown in FIG7 , or may have a configuration different from that shown in FIG7 .
[0196] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a computer device to store a computer program related to an anomaly detection method in a method embodiment. The computer program is loaded and executed by the processor to implement the anomaly detection method provided by the above method embodiment.
[0197] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0198] According to one aspect of the present application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the methods provided in the various optional implementations described above.
[0199] The facial recognition-based payment processing method, apparatus, device, storage medium, server, terminal, and program product provided by the present application have the following technical effects:
[0200] After receiving the facial recognition payment request from the target object, the image data captured by the image acquisition device and the audio data corresponding to the image data can be obtained in response to the facial recognition payment request from the target object. The audio data is then subjected to feature analysis for the number of objects to obtain an audio analysis result. If the audio analysis result indicates that the number of objects corresponding to the audio data is a single object, there is no need to load the facial detection module. The image data can be directly subjected to facial recognition based on the facial recognition model to obtain the target facial data of the target object, thereby executing the payment operation corresponding to the facial recognition payment request based on the target facial data. In this way, before loading the facial detection module, a feature analysis is first performed based on the audio data to obtain an audio analysis result that can indicate the number of objects in front of the current terminal. When the audio analysis result indicates that the current environment is a scene with a single object, facial recognition is directly performed without triggering the loading of related memory-intensive business modules such as facial detection of multiple objects. This optimizes the load time of the facial payment process, improves payment efficiency, and improves the product experience.
[0201] It should be noted that the order of the embodiments of the present application described above is for descriptive purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0202] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device, equipment, and storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0203] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0204] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A payment processing method based on facial recognition, characterized in that: The method is executed by a terminal, the terminal is provided with an image acquisition device, and the method includes: In response to a facial recognition payment request of a target object, acquiring image data captured by the image acquisition device and audio data corresponding to the image data; Performing feature analysis on the audio data with respect to the number of objects to obtain an audio analysis result; If the audio analysis result indicates that the number of objects corresponding to the audio data is single, performing facial recognition on the image data based on a facial recognition model to obtain target facial data of the target object; A payment operation corresponding to the facial recognition payment request is performed based on the target facial data.
2. The method according to claim 1, characterized in that The performing feature analysis on the audio data for the number of objects to obtain the audio analysis result comprises: Performing sound source separation detection on the audio data to obtain the sound source separation data, wherein the sound source separation data includes at least one sound source signal corresponding to the audio data; Counting the total number of sound source signals included in the at least one sound source signal; If the total number of signals is equal to a unit value, determining that the audio analysis result is a first result, the first result indicating that the number of objects corresponding to the audio data is a single object; If the total number of signals is greater than the unit value, the audio analysis result is determined to be a second result, and the second result indicates that the number of objects corresponding to the audio data is multiple.
3. The method according to any one of claims 1 to 2, characterized in that: The performing sound source separation detection on the audio data to obtain the sound source separation data comprises: Performing frame processing on the audio data to obtain multiple audio frames; Extracting energy features from the multiple audio frames to obtain audio energy features; The audio energy features are subjected to sound source separation based on a sound source separation model to obtain the sound source separation data; the sound source separation model is obtained by performing constraint training for sound source separation on a preset first deep neural network with sample audio energy features as input and sample sound source signals corresponding to the sample audio energy features as expected output, and the sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to at least two sample sound sources.
4. The method according to any one of claims 1 to 3, characterized in that: The performing feature analysis on the audio data with respect to the number of objects to obtain the audio analysis result comprises: determining a noise signal in the audio data; Calculating the noise field amplitude of the noise signal to obtain noise index data, where the noise index data is used to indicate the intensity of the ambient noise corresponding to the audio data; If the noise index data indicates that the ambient noise intensity corresponding to the audio data is lower than a preset intensity threshold, determining that the audio analysis result is a first result, where the first result is used to indicate that the number of objects corresponding to the audio data is a single object; If the noise index data indicates that the ambient noise intensity corresponding to the audio data is higher than a preset intensity threshold, the audio analysis result is determined to be a second result, and the second result is used to indicate that the number of objects corresponding to the audio data is multiple.
5. The method according to any one of claims 1 to 4, characterized in that: The determining of the noise signal in the audio data comprises: Performing frame processing on the audio data to obtain multiple audio frames; Extracting energy features from the multiple audio frames to obtain audio energy features; The audio energy feature is subjected to noise estimation based on a noise estimation model to obtain the noise signal; the noise estimation model is obtained by performing constrained training of noise estimation on a preset second deep neural network with sample audio energy features as input and sample noise signals corresponding to the sample audio energy features as expected output, and the sample audio energy features are obtained by extracting energy features based on audio data of a sample sound source or mixed audio data corresponding to at least two sample sound sources.
6. The method according to any one of claims 1 to 5, characterized in that: The terminal is also provided with a plurality of audio acquisition devices arranged in an array, and the method for acquiring the audio data includes: Acquire a plurality of initial audio data collected by a plurality of the audio collection devices; Performing sound source positioning on each of the initial audio data to obtain the sound source position information of each of the initial audio data; Based on the sound source position information, an audio signal whose sound source is located within a first spatial range of the terminal is determined from the multiple initial audio data to obtain the audio data. The first spatial range is located in front of the display screen of the terminal and within the field of view of the image acquisition device.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: If the audio analysis result indicates that the number of objects corresponding to the audio data is multiple, performing face detection on the image data based on a face detection model to obtain multiple face images; Filtering a target facial image of the target object from the multiple facial images; Performing facial recognition on the target facial image based on the facial recognition model to obtain the target facial data.
8. The method according to any one of claims 1 to 7, characterized in that Before responding to the facial recognition payment request of the target object, the method further includes: In response to a facial payment voice instruction, performing sound source positioning on the instruction audio data carried by the facial payment voice instruction to obtain the instruction sound source position information, wherein the instruction sound source position information is used to indicate the relative position of the issuing object of the instruction audio data relative to the image acquisition device; If the command sound source location information indicates that the issuing object is located within a second spatial range, a facial recognition payment request corresponding to the facial payment voice command is generated, and the second spatial range is located in front of the display screen of the terminal and the distance between the second spatial range and the display screen is within a preset distance range.
9. The method according to any one of claims 1 to 8, characterized in that: Screening out a target facial image of the target object from the multiple facial images comprises: Acquire facial position information corresponding to each of the plurality of facial images, each facial position information being used to represent a position of an object to which the corresponding facial image belongs relative to the image acquisition device; A facial image among the multiple facial images whose facial position information matches the instruction sound source position information is determined as the target facial image.
10. A payment processing device based on facial recognition, characterized in that: The device is applied to a terminal, the terminal is provided with an image acquisition device, and the device comprises: Data acquisition module: used for acquiring the image data acquired by the image acquisition device and the audio data corresponding to the image data in response to the facial recognition payment request of the target object; Feature analysis module: used to perform feature analysis on the audio data according to the number of objects to obtain audio analysis results; A facial recognition module: configured to perform facial recognition on the image data based on a facial recognition model to obtain target facial data of the target object if the audio analysis result indicates that the number of objects corresponding to the audio data is a single object; Payment module: used to execute the payment operation corresponding to the facial recognition payment request based on the target facial data.
11. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the facial recognition-based payment processing method as described in any one of claims 1 to 9.
12. A computer device, characterized in that: The device comprises a processor and a memory, wherein a computer program is stored in the memory, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method, device and system
CN109784157A
Face positioning method and device
CN110503045A
Voice matching method and related equipment
CN111091824A
Payment processing method and device, equipment, medium and program product
CN116258496A
Face recognition method and device, equipment and storage medium
CN116311413A