Task execution method and device and storage medium

By segmenting and rationally judging the collected images and sounds, ensuring that the background is reasonable and then biometric identification is carried out, the problem of thin feature information of traditional palm image authentication is solved, and the accuracy of identity authentication and the security of task execution is improved.

CN120258809APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410008510.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The traditional palm image-based identity authentication method has poor identity authentication results and poses security risks.

Method used

By obtaining the collected images of the target account and the sound during the acquisition process, image segmentation and background sound separation are carried out, background images and sound are judged reasonably, and biological images are recognized before performing tasks after the background is reasonable.

Benefits of technology

Improves the accuracy of identity authentication and the security of task execution, reducing the risk of consuming identity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258809A_ABST
    Figure CN120258809A_ABST
Patent Text Reader

Abstract

The invention discloses a task execution method and device and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, smart traffic and Internet of Vehicles, and the method comprises the following steps: responding to a target task execution request of a target account; obtaining a target collection image of the target account and a target sound in a collection process of the target collection image; performing image segmentation processing on the target acquisition image to obtain a target biological image and a target background image; performing a background sound separation operation on the target sound to obtain a target background sound; performing rationality judgment on the target background image and the target background sound to obtain a target judgment result; under the condition that the target judgment result represents that the target background image and the target background sound are reasonable, the target biological image is recognized, and a target recognition result is obtained; and under the condition that the target identification result is an identification passing result, executing the target task. According to the invention, the security of task execution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a task execution method, apparatus, and storage medium. Background Art

[0002] With the popularization of mobile payment, biometric payment methods such as face brushing payment and palm brushing payment are increasingly favored by people. However, these emerging payment methods also bring new security risks. Before payment, the user can be verified, and the order can be paid only after determining the true identity of the user to prevent the situation of using someone else's identity for payment. In this process, how to better manage the user's identity verification has gradually become the focus of attention of all parties. In the traditional method of implementing identity authentication based on the palm, identity authentication is often performed based on a palm image from a single angle, and the feature information of the palm image involved in the identity authentication process is thin, which will affect the effect of identity authentication. Summary of the Invention

[0003] This application provides a task execution method, apparatus, and storage medium, which can improve the security of task execution.

[0004] On the one hand, this application provides a task execution method, and the method includes:

[0005] In response to a target task execution request of a target account, obtain a target acquisition image of the target account and a target sound during the acquisition process of the target acquisition image; the target acquisition image includes the biometric features of the target account;

[0006] Perform image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image;

[0007] Perform background sound separation operation on the target sound to obtain a target background sound;

[0008] Perform a rationality judgment on the target background image and the target background sound to obtain a target judgment result;

[0009] In the case where the target judgment result indicates that both the target background image and the target background sound are reasonable, perform identification on the target biometric image to obtain a target identification result;

[0010] In the case where the target identification result is a passed identification result, execute the target task.

[0011] Optionally, the performing image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image includes:

[0012] Input the target acquisition image into the background image extraction model for background image extraction processing to obtain the target background image; the background image extraction model is obtained by training a third preset model based on sample acquisition images.

[0013] Segment the target background image from the target acquisition image to obtain the target biological image.

[0014] Among them, the training method of the background image extraction model includes:

[0015] Obtain sample acquisition images; the sample acquisition images are labeled with sample background image labels.

[0016] Input the sample acquisition image into the third preset model for background image extraction to obtain the sample background image result.

[0017] Train the third preset model according to the difference between the sample background image result and the sample background image label to obtain the background image extraction model.

[0018] Optionally, the operation of separating the background sound from the target sound to obtain the target background sound includes:

[0019] Input the target sound into the background sound extraction model for background sound separation operation to obtain the target background sound; the background sound extraction model is obtained by training a fourth preset model based on sample sounds.

[0020] Among them, the training method of the background sound extraction model includes:

[0021] Obtain sample sounds; the sample sounds are labeled with sample background sound labels.

[0022] Input the sample sound into the fourth preset model for background sound extraction to obtain the sample background sound result.

[0023] Train the fourth preset model according to the difference between the sample background sound result and the sample background sound label to obtain the background sound extraction model.

[0024] Optionally, when the first judgment result is a reasonable result, match the target background image with the target acquisition scene to obtain the first matching result, including:

[0025] When the first judgment result is a reasonable result, query the reference scene image matching the target acquisition scene based on the preset scene image library; the preset scene image library stores the mapping relationship between the preset acquisition scene and the preset reference scene image.

[0026] Match the target background image with the reference scene image to obtain a first matching result.

[0027] Optionally, when the second judgment result is a reasonable result, match the target background sound with the target acquisition scene to obtain a second matching result, including:

[0028] When the second judgment result is a reasonable result, query the reference sound that matches the target acquisition scene based on a preset scene sound library; the preset scene sound library stores the mapping relationship between the preset acquisition scene and the preset reference scene sound;

[0029] Match the target background sound with the reference sound to obtain a second matching result.

[0030] On the other hand, a task execution device is provided, and the device includes:

[0031] A target information acquisition module, configured to acquire a target acquisition image of the target account and a target sound during the acquisition process of the target acquisition image in response to a target task execution request of the target account; the target acquisition image includes the biometric features of the target account;

[0032] A target image determination module, configured to perform image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image;

[0033] A target sound determination module, configured to perform background sound separation operation on the target sound to obtain a target background sound;

[0034] A judgment module, configured to perform a rationality judgment on the target background image and the target background sound to obtain a target judgment result;

[0035] A target recognition module, configured to recognize the target biometric image to obtain a target recognition result when the target judgment result indicates that both the target background image and the target background sound are reasonable;

[0036] A task execution module, configured to execute the target task when the target recognition result is a recognition pass result.

[0037] Exemplarily, the judgment module includes:

[0038] A first judgment unit, configured to perform a rationality judgment on the target background image to obtain a first judgment result;

[0039] A second judgment unit, configured to perform a rationality judgment on the target background sound to obtain a second judgment result;

[0040] A third judgment unit, configured to determine a target judgment result based on the first judgment result and the second judgment result.

[0041] Exemplarily, the first judgment unit includes:

[0042] A first judgment subunit, configured to perform a rationality judgment on the target background image based on a first preset strategy to obtain a first judgment result; the first preset strategy includes determining that the preset background image is unreasonable if there is at least one of a virtual background, a fictional biometric background, a paper background, a synthetic background, and a posture-distorted background in the preset background image;

[0043] Exemplarily, the second judgment unit includes:

[0044] A second judgment subunit, configured to perform a rationality judgment on the target background sound based on a second preset strategy to obtain a second judgment result; the second preset strategy includes determining that the preset background sound is unreasonable if there is at least one of silence, multiple sounds, noise, and interrupted sound in the preset background sound.

[0045] Exemplarily, the first judgment subunit includes:

[0046] A first detection unit, configured to detect whether there is a virtual background in the target background image based on the first preset strategy;

[0047] A second detection unit, configured to detect whether there is a fictional biometric background in the target background image;

[0048] A third detection unit, configured to detect whether there is a paper background in the target background image;

[0049] A fourth detection unit, configured to detect whether there is a synthetic background in the target background image;

[0050] A fifth detection unit, configured to detect whether there is a posture-distorted background in the target background image;

[0051] A sixth detection unit, configured to determine that the first judgment result is an unreasonable result if there is at least one of a virtual background, a fictional biometric background, a paper background, a synthetic background, and a posture-distorted background in the target background image.

[0052] Exemplarily, the second judgment subunit includes:

[0053] A silence detection unit, configured to detect whether there is silence in the target background sound based on the second preset strategy;

[0054] A multiple-sound detection unit, configured to detect whether there are multiple sounds in the target background sound;

[0055] An interruption sound detection unit, configured to detect whether there is an interruption sound in the target background sound;

[0056] A result determination unit, configured to determine that the second judgment result is an unreasonable result if there is at least one of silence, multiple sounds, noise, and interruption sound in the target background sound.

[0057] Exemplarily, the third judgment unit includes:

[0058] A target scene determination subunit, configured to determine a target acquisition scene corresponding to the target acquisition image according to the target task;

[0059] A first matching subunit, configured to match the target background image with the target acquisition scene to obtain a first matching result when the first judgment result is a reasonable result;

[0060] A second matching subunit, configured to match the target background sound with the target acquisition scene to obtain a second matching result when the second judgment result is a reasonable result;

[0061] A result determination subunit, configured to determine a target judgment result according to the first matching result and the second matching result.

[0062] Exemplarily, the first judgment unit includes:

[0063] A first prediction subunit, configured to input the target background image into an image judgment model for reasonableness judgment to obtain a first judgment result; the image judgment model is obtained by training a first preset model for reasonableness judgment based on sample background images;

[0064] Exemplarily, the second judgment unit includes:

[0065] A second prediction subunit, configured to input the target background sound into a sound judgment model for reasonableness judgment to obtain a second judgment result; the sound judgment model is obtained by training a second preset model for reasonableness judgment based on sample background sounds.

[0066] Exemplarily, the device further includes:

[0067] A sample background image acquisition module, configured to acquire sample background images, and the sample background images are labeled with sample image judgment result labels;

[0068] A sample image prediction module, configured to input the sample background images into the first preset model for reasonableness judgment to obtain sample image prediction results;

[0069] A first parameter adjustment module, configured to adjust the model parameters of the first preset model based on the difference between the prediction result of the sample image and the label of the sample image judgment result until the training end condition is satisfied;

[0070] An image model determination module, configured to determine the first preset model at the end of training as the image judgment model.

[0071] Exemplarily, the apparatus further includes:

[0072] A sample background sound acquisition module, configured to acquire a sample background sound, where the sample background sound is labeled with a sample sound judgment result label;

[0073] A sample sound prediction module, configured to input the sample background sound into the second preset model for rationality judgment to obtain a sample sound prediction result;

[0074] A second parameter adjustment module, configured to adjust the model parameters of the first preset model based on the difference between the sample sound prediction result and the label of the sample sound judgment result until the training end condition is satisfied;

[0075] A sound model determination module, configured to determine the second preset model at the end of training as the sound judgment model.

[0076] Exemplarily, the target sound determination module includes:

[0077] A sound separation unit, configured to perform a background sound separation operation on the target sound to obtain a target device acquisition sound and a target account sound; the target device acquisition sound is the sound of an image acquisition device for acquiring the target biological acquisition image;

[0078] A sound determination unit, configured to remove the target device acquisition sound and the target account sound from the target sound to obtain a target background sound.

[0079] Exemplarily, the target image determination module includes:

[0080] An image processing unit, configured to perform image preprocessing operations on the target acquisition image to obtain a target processed image; the image preprocessing includes at least one of image denoising and image enhancement;

[0081] An image segmentation unit, configured to perform image segmentation processing on the target processed image.

[0082] Exemplarily, the target sound determination module includes:

[0083] A voice preprocessing unit for performing voice preprocessing operations on the target voice to obtain a target processed voice; the voice preprocessing includes at least one of voice denoising and filtering processing;

[0084] A background voice separation unit for performing background voice separation operations on the target processed voice.

[0085] Exemplarily, the target voice determination module includes:

[0086] A first segmentation unit for performing image segmentation processing on the target captured image to obtain a target palm-sweeping image and a target background image if the target captured image includes a palm-sweeping image;

[0087] A second segmentation unit for performing image segmentation processing on the target captured image to obtain a target face image and a target background image if the target captured image includes a face image.

[0088] On the other hand, a task execution device is provided, the device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the task execution method as described above.

[0089] On the other hand, a computer storage medium is provided, and at least one instruction or at least one program segment is stored in the computer storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the task execution method as described above.

[0090] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes to implement the task execution method as described above.

[0091] The task execution method, device and storage medium provided by this application have the following technical effects:

[0092] In response to a target task execution request of a target account, this application acquires a target acquisition image of the target account and target sound during the acquisition process of the target acquisition image; the target acquisition image includes the biometric features of the target account; performs image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image; performs background sound separation operation on the target sound to obtain a target background sound; performs a rationality judgment on the target background image and the target background sound to obtain a target judgment result; in the case where the target judgment result indicates that both the target background image and the target background sound are reasonable, it indicates that both the background image and the background sound corresponding to the target account are reasonable, that is, both the background image and the background sound corresponding to the target account conform to the scene corresponding to the target task. At this time, the target biometric image is further recognized to obtain a target recognition result; thereby improving the accuracy of the recognition result; in the case where the target recognition result is a recognition passed result, the target task is executed; further improving the security of task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions and advantages in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0094] Figure 1 is a schematic diagram of a task execution system provided by an embodiment of this specification;

[0095] Figure 2 is a schematic flowchart of a task execution method provided by an embodiment of this specification;

[0096] Figure 3 is a schematic flowchart of a method for performing a rationality judgment on the above-mentioned target background image and the above-mentioned target background sound to obtain a target judgment result;

[0097] Figure 4 is a schematic flowchart of a method for performing a rationality judgment on the above-mentioned target background image based on a first preset strategy to obtain a first judgment result;

[0098] Figure 5 is a schematic flowchart of a method for training an image judgment model provided by an embodiment of this specification;

[0099] Figure 6 is a schematic flowchart of a method for performing a rationality judgment on the above-mentioned target background sound based on a second preset strategy to obtain a second judgment result;

[0100] Figure 7 It is a schematic flowchart of a method for training a voice judgment model provided by an embodiment of this specification;

[0101] Figure 8 It is a schematic structural diagram of a task execution device provided by an embodiment of this specification;

[0102] Figure 9 It is a schematic structural diagram of a server provided by an embodiment of this specification. Detailed implementation manners

[0103] Next, in combination with the accompanying drawings in the embodiments of this specification, the technical solutions in the embodiments of this specification will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0104] First, some nouns or terms that appear in the process of describing the embodiments of this specification are explained as follows:

[0105] Multi-modal security verification technology: The palm brushing recognition background sound risk control module in this application adopts multi-modal security verification technology, that is, fuses and verifies multiple types of information (such as sound, image, etc.), improving the security and accuracy in scenarios such as palm brushing payment. This technology can avoid the deficiencies of single-modal verification and improve the security and robustness of biometric recognition payment.

[0106] Background image extraction technology: The palm brushing recognition background sound risk control module in this application also adopts background image extraction technology to extract the background image in the palm brushing scenario. This technology can avoid the interference of the background image on palm brushing recognition and improve the recognition accuracy and stability in scenarios such as palm brushing payment.

[0107] Background sound extraction technology: The palm brushing recognition background sound risk control module in this application adopts background sound extraction technology to extract the background sound from the sound in the palm brushing scenario. This technology can avoid the interference of other sounds on palm brushing recognition and improve the recognition accuracy and stability in scenarios such as palm brushing payment.

[0108] Risk control: The palm brushing recognition background sound risk control module in this application adopts a risk control strategy to judge whether the background sound in the palm brushing scenario is reasonable and to judge possible risks. Through this technology, the security and accuracy in scenarios such as palm brushing payment can be improved.

[0109] Palm Swipe Recognition: The background sound risk control module for palm swipe recognition in this application adopts palm swipe recognition technology, which is used to identify background sounds and palm swipe actions in palm swipe scenarios. This technology can achieve automatic recognition and verification in scenarios such as palm swipe payment, improving the convenience and efficiency of payment.

[0110] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0111] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0112] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0113] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0114] The solution provided by the embodiments of this application relates to technologies such as machine learning in artificial intelligence, and will be specifically described through the following embodiments.

[0115] It can be understood that in the specific implementation of this application, data related to the user's biological images, etc. is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0116] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0117] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0118] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a task execution system provided by the embodiments of this specification. As Figure 1 shown, the task execution system can at least include a server 01 and a client 02.

[0119] Specifically, in the embodiments of this specification, the above-mentioned server 01 may include an independently operating server, a distributed server, or a server cluster composed of multiple servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms. The server 01 may include a network communication unit, a processor, a memory, and so on. Specifically, the above-mentioned server 01 may be used to obtain the target acquisition image of the above-mentioned target account and the target sound during the acquisition process of the above-mentioned target acquisition image; perform image segmentation processing on the above-mentioned target acquisition image to obtain a target biological image and a target background image; perform background sound separation operation on the above-mentioned target sound to obtain a target background sound; perform a rationality judgment on the above-mentioned target background image and the above-mentioned target background sound to obtain a target judgment result; when the above-mentioned target judgment result indicates that both the above-mentioned target background image and the above-mentioned target background sound are reasonable, perform recognition on the above-mentioned target biological image to obtain a target recognition result; when the above-mentioned target recognition result is a recognition-passed result, execute the above-mentioned target task.

[0120] Specifically, in the embodiments of this specification, the above-mentioned client 02 may include entity devices of types such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, smart wearable devices, smart speakers, vehicle-mounted terminals, smart TVs, etc. It may also include software running on the entity devices, such as web pages provided to users by some service providers, or applications provided to users by these service providers. Specifically, the above-mentioned client 02 may be used to display the target recognition result.

[0121] The following introduces a task execution method of this application. Figure 2 It is a flowchart of a task execution method provided by the embodiments of this specification. This specification provides the method operation steps as described in the embodiments or the flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step order listed in the embodiments is only one of the many step execution orders and does not represent the only execution order. When the actual system or server product executes, it may be executed in the order of the embodiments or as shown in the drawings, or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing).

[0122] Specifically, as Figure 2 shown, the above method may include:

[0123] S201: In response to a target task execution request of a target account, obtain the target acquisition image of the target account and the target sound during the acquisition process of the target acquisition image; the target acquisition image includes the biometric features of the target account.

[0124] In the embodiments of this specification, the target account may be a user on the terminal side, and the target task may include, but is not limited to, payment tasks, opening doors, unlocking locks, etc. The method of this embodiment can be applied to scenarios such as palm brushing payment, face brushing payment, entering and exiting through palm brushing (face brushing), and unlocking through palm brushing (face brushing). The target acquisition image may be an image of the target account collected by an image acquisition device, and the image may include biometric features such as the face, palm, sole, eyes, etc. of the target account, and may also include the image acquisition background.

[0125] In the embodiments of this specification, the target sound can be obtained by recording the sound during the acquisition process of the target acquisition image, and the target sound may include the sound of the image acquisition device, the sound of the target account, and the background sound.

[0126] S203: Perform image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image.

[0127] In the embodiments of this specification, the performing image segmentation processing on the target acquisition image includes:

[0128] Perform image preprocessing operations on the target acquisition image to obtain a target processed image; the image preprocessing includes at least one of image denoising and image enhancement;

[0129] Perform image segmentation processing on the target processed image.

[0130] In the embodiments of this specification, at least one image preprocessing operation of image denoising and image enhancement can be first performed on the target acquisition image to obtain a target processed image; then, image segmentation is performed according to the target processed image.

[0131] In the embodiments of this specification, the performing image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image includes:

[0132] If the target acquisition image includes a palm brushing image, perform image segmentation processing on the target acquisition image to obtain a target palm brushing image and a target background image;

[0133] If the target acquisition image includes a face image, perform image segmentation processing on the target acquisition image to obtain a target face image and a target background image.

[0134] In the embodiments of this specification, the above-mentioned target acquisition image can be subjected to image segmentation processing according to the acquired biometric features to obtain a target biometric image and a target background image; the target biometric image can include, but is not limited to, a target palm image, a target face image, a target eye image, etc.; if the above-mentioned target acquisition image includes a palm image, the above-mentioned target acquisition image is subjected to image segmentation processing to obtain a target palm image and a target background image; if the above-mentioned target acquisition image includes a face image, the above-mentioned target acquisition image is subjected to image segmentation processing to obtain a target face image and a target background image, so as to separate the background image from the target acquisition image.

[0135] In the embodiments of this specification, image acquisition and separation can also be performed through a background image extraction module, and the background image extraction module is implemented based on digital image processing technology. Exemplarily, when the biometric feature is a palm, this module takes a live photo of the palm when the user performs palm brushing recognition, and uses image segmentation technology to separate the palm from the background and extract the background image. This module can also use deep learning algorithms to extract features and perform image recognition on the background, improving the accuracy and speed of background image extraction. The background image extraction module mainly includes sub-modules such as image acquisition, preprocessing, segmentation, feature extraction, and recognition (these are the steps of conventional image processing). Among them, the image acquisition sub-module is responsible for obtaining a live photo of the palm when the user performs palm brushing recognition; the preprocessing sub-module preprocesses the image, such as denoising, enhancement, etc.; the segmentation sub-module separates the palm from the background; the feature extraction sub-module extracts features from the background image; the recognition sub-module uses machine learning algorithms to recognize the background image.

[0136] When the target account (user) performs palm brushing recognition, the background image extraction module first obtains a live photo of the palm through the camera. Then, this module preprocesses the image, such as denoising, enhancement, etc. Next, image segmentation technology is used to separate the palm from the background and extract the background image. Finally, this module can use deep learning algorithms to extract features and perform image recognition on the background, further improving the accuracy and speed of background image extraction. By using the background image extraction module, the palm and the background can be accurately separated, and the background image can be extracted. This module can effectively improve the accuracy and security of multi-modal authentication and is applicable to the authentication requirements in various scenarios.

[0137] In an exemplary embodiment, performing image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image includes:

[0138] Inputting the target acquisition image into a background image extraction model for background image extraction processing to obtain a target background image; the background image extraction model is obtained by training a third preset model based on sample acquisition images.

[0139] Segment the target background image from the target captured image to obtain the target biological image.

[0140] Among them, the training method of the background image extraction model includes:

[0141] Obtain a sample captured image; the sample captured image is labeled with a sample background image label;

[0142] Input the sample captured image into a third preset model for background image extraction to obtain a sample background image result;

[0143] Train the third preset model according to the difference between the sample background image result and the sample background image label to obtain a background image extraction model.

[0144] In the embodiments of this specification, the difference between the sample background image result and the sample background image label can be used to construct third loss data; and the model parameters of the third preset model can be adjusted according to the third loss data until the training end condition is met; and the third preset model at the end of training is determined as the background image extraction model; thus facilitating the rapid and accurate extraction of the target background image through the background image extraction model.

[0145] S205: Perform a background sound separation operation on the above-mentioned target sound to obtain a target background sound.

[0146] In the embodiments of this specification, the above-mentioned background sound separation operation on the above-mentioned target sound includes:

[0147] Perform a sound preprocessing operation on the above-mentioned target sound to obtain a target processed sound; the above-mentioned sound preprocessing includes at least one of sound denoising and filtering;

[0148] Perform a background sound separation operation on the above-mentioned target processed sound.

[0149] In the embodiments of this specification, the target processed sound can be obtained by performing at least one background sound separation operation of sound denoising and filtering on the above-mentioned target sound; and then a background sound separation operation is performed on the above-mentioned target processed sound to obtain a target background sound.

[0150] In the embodiments of this specification, the above-mentioned background sound separation operation on the above-mentioned target sound to obtain a target background sound includes:

[0151] Perform a background sound separation operation on the above-mentioned target sound to obtain a target device captured sound and a target account sound; the above-mentioned target device captured sound is the sound of the image capture device capturing the above-mentioned target biological captured image;

[0152] Remove the sound collected by the target device and the sound of the target account from the above target sound to obtain the target background sound.

[0153] In the embodiments of this specification, during the sound separation operation, the sound collected by the target device and the sound of the target account can be extracted first; the sound of the target account is the voice or action sound emitted by the target account during image acquisition; then, by removing the sound collected by the target device and the sound of the target account from the above target sound, the target background sound can be obtained quickly and accurately.

[0154] In an exemplary embodiment, the target background sound in the target sound can be extracted by a background sound extraction module; the background sound extraction module is implemented based on digital signal processing technology. This module records the background sound when the user performs palm brushing recognition, and uses audio separation technology to separate the sounds of the device, the user, etc., which are mainly divided into the sounds of basic actions, the sounds of the palm-brushing user, and the background sound. Then, this module can use machine learning algorithms to analyze and identify the separated audio to achieve the verification and identification of the user's identity.

[0155] The background sound extraction module mainly includes sub-modules such as audio acquisition, preprocessing, separation, feature extraction, and recognition. Among them, the audio acquisition sub-module is responsible for recording the background sound when the user performs palm brushing recognition; the preprocessing sub-module preprocesses the audio, such as denoising, filtering, etc.; the separation sub-module separates the sounds of the device, the user, etc.; the feature extraction sub-module extracts features from the separated audio; the recognition sub-module uses machine learning algorithms to identify the separated audio.

[0156] When the user performs palm brushing recognition, the background sound extraction module first records the background sound when the user performs palm brushing recognition through a microphone. Then, this module preprocesses the audio, such as denoising, filtering, etc. Next, it uses audio separation technology to separate the sounds of the device, the user, etc., and extracts the background sound. Finally, this module can perform feature extraction and audio recognition on the background sound by using machine learning algorithms to further improve the accuracy and speed of background sound extraction.

[0157] By using the background sound extraction module, the sounds of the device, the user, etc. can be accurately separated, and the background sound can be extracted. This module can effectively improve the accuracy and security of multi-modal identity verification and is applicable to the identity verification requirements in various scenarios.

[0158] Exemplarily, the operation of separating the background sound from the target sound to obtain the target background sound includes:

[0159] Input the target sound into the background sound extraction model for background sound separation operation to obtain the target background sound; the background sound extraction model is obtained by training a fourth preset model based on sample sounds.

[0160] Among them, the sample sound is the sound during the image acquisition process of the sample acquisition image;

[0161] The training method of the background sound extraction model includes:

[0162] Obtain sample sounds; the sample sounds are labeled with sample background sound labels;

[0163] Input the sample sound into the fourth preset model for background sound extraction to obtain the sample background sound result;

[0164] According to the difference between the sample background sound result and the sample background sound label, train the fourth preset model to obtain the background sound extraction model.

[0165] In the embodiments of this specification, the fourth loss data can be constructed according to the difference between the sample background sound result and the sample background sound label; and the model parameters of the fourth preset model are adjusted according to the fourth loss data until the training end condition is met; and the fourth preset model at the end of training is determined as the background sound extraction model; thus facilitating the rapid and accurate extraction of the target background sound through the background sound extraction model.

[0166] S207: Perform a rationality judgment on the above-mentioned target background image and the above-mentioned target background sound to obtain a target judgment result.

[0167] In the embodiments of this specification, as Figure 3 shown, the above-mentioned performing a rationality judgment on the above-mentioned target background image and the above-mentioned target background sound to obtain a target judgment result includes:

[0168] S2071: Perform a rationality judgment on the above-mentioned target background image to obtain a first judgment result;

[0169] In the embodiments of this specification, the above-mentioned performing a rationality judgment on the above-mentioned target background image to obtain a first judgment result includes:

[0170] Perform a rationality judgment on the above-mentioned target background image based on a first preset strategy to obtain a first judgment result; the first preset strategy includes determining that the above-mentioned preset background image is unreasonable if there is at least one of a virtual background, a fictional biological feature background, a paper background, a synthetic background, and a posture-distorted background in the preset background image.

[0171] In some embodiments, as Figure 4As shown above, the rationality of the target background image is judged based on the first preset strategy to obtain a first judgment result, including:

[0172] S401: Detect whether there is a virtual background in the target background image based on the first preset strategy;

[0173] S403: Detect whether there is a fictional biometric background in the target background image;

[0174] S405: Detect whether there is a paper background in the target background image;

[0175] S407: Detect whether there is a synthetic background in the target background image;

[0176] S409: Detect whether there is a posture distortion background in the target background image;

[0177] S4011: If there is at least one of a virtual background, a fictional biometric background, a paper background, a synthetic background, and a posture distortion background in the target background image, determine that the first judgment result is an unreasonable result.

[0178] In the embodiments of this specification, the abnormal images in the target background image can be quickly identified through the constructed first preset strategy, so as to obtain the first judgment result; if there is no virtual background, fictional biometric background, paper background, synthetic background, and posture distortion background in the target background image, it is determined that the first judgment result is a reasonable result.

[0179] In the embodiments of this specification, the rationality of the target background image can be judged by the background image risk control module (i.e., the image judgment model) to obtain the first judgment result. The background image risk control module is implemented based on deep learning technology. This module is trained using a large number of on-site photos of users who normally perform palm brushing recognition and the labeled backgrounds, as well as a large number of risk scenarios, such as virtual backgrounds, fake hand backgrounds, paper backgrounds, synthetic backgrounds, and backgrounds with distorted standing postures of people. Through the learning and training of these data, this module can judge whether the on-site photos and the corresponding backgrounds in the actual scenario are reasonable and whether there are risks.

[0180] In the embodiments of this specification, the rationality of the target background image is judged to obtain the first judgment result, including:

[0181] Input the target background image into the image judgment model for rationality judgment to obtain the first judgment result; the image judgment model is obtained by training the first preset model for rationality judgment based on sample background images.

[0182] In the embodiments of this specification, such as Figure 5As shown, the training method of the above image judgment model includes:

[0183] S501: Obtain a sample background image, and the above sample background image is labeled with a sample image judgment result label;

[0184] S503: Input the above sample background image into the above first preset model for rationality judgment to obtain a sample image prediction result;

[0185] S505: Based on the difference between the above sample image prediction result and the above sample image judgment result label, adjust the model parameters of the above first preset model until the training end condition is met;

[0186] S507: Determine the first preset model at the end of training as the above image judgment model.

[0187] In the embodiments of this specification, the sample image judgment result label can be determined according to the first preset strategy, and the sample background image can be labeled. The first loss data can be constructed according to the difference between the sample image prediction result and the above sample image judgment result label; and the model parameters of the first preset model can be adjusted according to the first loss data until the training end condition is met; and the first preset model at the end of training can be determined as the image judgment model; thereby facilitating obtaining the first judgment result quickly and accurately through the image judgment model. In the embodiments of this specification, the first preset model can include, but is not limited to, deep learning models such as convolutional neural networks (CNNs). First, data collection of sample background images can be carried out: collect a large number of on-site photos of users who are normally performing palm brushing recognition and the labeled backgrounds, as well as a large number of training data for risky scenarios, such as virtual backgrounds, fake hand and fake background, paper backgrounds, synthetic backgrounds, backgrounds with distorted standing postures of people, etc. These data can be obtained through methods such as web crawling, manual collection, and simulation generation. Then, data preprocessing is performed: the collected data is preprocessed, including operations such as image format conversion, image enhancement, and noise removal, to improve the quality and accuracy of the data. Then, a deep learning model such as a convolutional neural network (CNN) is used for construction. CNN can effectively extract the feature information of images and classify and identify images. The preprocessed training data set is input into the model for training. The training process uses deep learning-related technologies such as the backpropagation algorithm and the gradient descent method to train and optimize the model. During the training process, methods such as cross-validation can be used to evaluate and optimize the model to improve the accuracy and generalization ability of the model.

[0188] In the embodiments of this specification, after the model training is completed, the validation dataset and the test dataset can be used to validate and test the model. The validation dataset is used to adjust the model parameters, and the test dataset is used to evaluate the overall performance of the model. Through validation and testing, indicators such as the accuracy, recall rate, and precision rate of the model, as well as the robustness and generalization ability of the model, can be evaluated.

[0189] During the application process, when the user performs palm brushing recognition, the background image risk control module first obtains the background image when the user brushes the palm. Then, this module inputs the background image into the neural network module for processing, extracts the feature information of the background image, and makes a risk judgment. Finally, this module outputs a judgment result to determine whether the background image is reasonable and whether there is a risk. Suppose a company uses palm brushing recognition technology to control the entry and exit of employees through the company gate. To ensure security, the company introduces a background image risk control module to identify whether the background image during palm brushing is reasonable.

[0190] Exemplarily, when the collected biometric feature of the target account is a palm image, a virtual background can be detected: the background image risk control module will detect whether there is a virtual background in the image, such as a green screen or a situation using virtual background software. Detect a fake hand and fake background: the background image risk control module will detect whether there is a fake hand and fake background in the image, such as using a photo or video to replace the real palm brushing process. Detect a paper background: the background image risk control module will detect whether there is a paper background in the image, such as using a paper to replace the real palm process. Detect a synthetic background: the background image risk control module will detect whether there is a synthetic background in the image, such as using synthetic software to combine multiple background images into one image. Detect a distorted background of the person's standing posture: the background image risk control module will detect whether there is a distorted background of the person's standing posture in the image, such as using special equipment or technology to distort the background; thus, risky background images can be effectively identified, improving the security and accuracy of palm brushing recognition.

[0191] In this embodiment, the image judgment model can make a risk judgment on the background image during palm brushing recognition, improving the accuracy and security of identity verification. This module can effectively identify risky background images such as virtual backgrounds, fake hand and fake backgrounds, paper backgrounds, synthetic backgrounds, and distorted backgrounds of the person's standing posture, thereby preventing identity verification from being attacked or deceived.

[0192] S2073: Make a rationality judgment on the above target background sound to obtain a second judgment result;

[0193] In some embodiments, the above making a rationality judgment on the above target background sound to obtain a second judgment result includes:

[0194] Judging the rationality of the above-mentioned target background sound based on a second preset strategy to obtain a second judgment result; the above-mentioned second preset strategy includes determining that the above-mentioned preset background sound is unreasonable if there is at least one of silence, multiple sounds, noise, and interrupted sounds in the preset background sound.

[0195] In some embodiments, the above-mentioned judging the rationality of the above-mentioned target background sound based on a second preset strategy to obtain a second judgment result includes:

[0196] Judging the rationality of the above-mentioned target background sound based on a second preset strategy to obtain an initial judgment result;

[0197] In the case where the above-mentioned initial judgment result indicates that the target background sound is unreasonable, based on the image acquisition device emitting infrasound waves, collecting the target sound waves reflected by the above-mentioned target biological image;

[0198] Judging the rationality of the above-mentioned target sound waves to obtain a second judgment result.

[0199] Exemplarily, judging the rationality of the above-mentioned target sound waves to obtain a second judgment result may include: obtaining acquisition scene information; determining a sound transmission medium according to the acquisition scene information; determining a standard sound wave range corresponding to the target biological image according to the sound transmission medium; based on the comparison result between the target sound waves and the standard sound wave range, judging the rationality of the above-mentioned target sound waves to obtain a second judgment result.

[0200] In the embodiments of this specification, in the case where the above-mentioned initial judgment result indicates that the target background sound is reasonable, the initial judgment result is determined as the second judgment result. In the case where the above-mentioned initial judgment result indicates that the target background sound is unreasonable, the sound transmission medium may be determined according to the current acquisition scene information; different sound transmission media correspond to different sound waves; the standard sound wave range corresponding to the target biological image may be determined according to the sound transmission medium; then it is judged whether the target sound waves are within the standard sound wave range; when the target sound waves are within the standard sound wave range, the second judgment result is determined as a reasonable result; when the target sound waves are not within the standard sound wave range, the second judgment result is determined as an unreasonable result; the image acquisition device may be a face-scanning device or a palm-scanning device. When the image acquisition device is a palm-scanning device, the sound waves emitted by the palm may be collected through the palm-scanning device and the rationality may be judged. In this embodiment, in the case where the initial judgment result indicates that the target background sound is unreasonable, further judgment can be made by collecting sound waves twice, thereby improving the accuracy of the judgment result.

[0201] In the embodiments of this specification, as Figure 6 shown, the above-mentioned judging the rationality of the above-mentioned target background sound based on a second preset strategy to obtain a second judgment result includes:

[0202] S601: Detect whether there is silence in the above-mentioned target background sound based on the above-mentioned second preset policy;

[0203] S603: Detect whether there are multiple sounds in the above-mentioned target background sound;

[0204] S605: Detect whether there is an interrupted sound in the above-mentioned target background sound;

[0205] S607: If there is at least one of silence, multiple sounds, background noise, and interrupted sound in the above-mentioned target background sound, determine that the above-mentioned second judgment result is an unreasonable result.

[0206] In the embodiments of this specification, when the user makes a palm payment in a store, the background sound during palm payment can be analyzed and judged to determine whether the background sound is reasonable. Its risk control strategy may include the following aspects:

[0207] Detecting silence: The background sound risk control module will detect whether there is silence in the background sound. If so, it is judged as unreasonable background sound.

[0208] Detecting chaotic sounds: The background sound risk control module will detect whether there is a chaotic sound situation in the background sound, such as multiple sounds appearing simultaneously, or there is background noise interference, etc. If so, it is judged as unreasonable background sound.

[0209] Detecting discontinuous sounds: The background sound risk control module will detect whether there is a discontinuous sound situation in the background sound, such as intermittent sounds or sound interruptions, etc. If so, it is judged as unreasonable background sound.

[0210] Through the above risk control strategy, risky background sounds can be effectively identified, thereby improving the security and accuracy of palm payment. If the background sound is determined to be unreasonable, this palm payment can be intercepted to prevent unsafe payment behaviors from occurring.

[0211] In the embodiments of this specification, the rationality of the above-mentioned target background sound can be judged through the background sound risk control module (sound judgment model) to obtain the second judgment result.

[0212] In the embodiments of this specification, the above-mentioned rationality judgment of the above-mentioned target background sound to obtain the second judgment result includes:

[0213] Input the above-mentioned target background sound into the sound judgment model for rationality judgment to obtain the second judgment result; the above-mentioned sound judgment model is obtained by training the second preset model for rationality judgment based on sample background sounds.

[0214] In the embodiments of this specification, as Figure 7 shown, the training method of the above-mentioned sound judgment model includes:

[0215] S701: Obtain the sample background sound, and the above sample background sound is labeled with a sample sound judgment result label;

[0216] S703: Input the above sample background sound into the above second preset model for rationality judgment to obtain a sample sound prediction result;

[0217] S705: Based on the difference between the above sample sound prediction result and the above sample sound judgment result label, adjust the model parameters of the above first preset model until the training end condition is met;

[0218] S707: Determine the second preset model at the end of training as the above sound judgment model.

[0219] In the embodiments of this specification, the sample sound judgment result label can be determined according to a second preset strategy, and the sample background sound can be labeled.

[0220] In the embodiments of this specification, the second preset model can be a neural network model. The neural network model is constructed using deep learning models such as convolutional neural networks (CNNs) or long short-term memory networks (LSTMs) for feature extraction and classification of background sounds. The training dataset includes background sounds, device sounds, user sounds, and pure background sounds recorded from a large number of examples of users who normally perform palm brushing recognition, as well as a large amount of sound training data for special scenarios, such as the sounds when users use electronic screens, paper, or prosthetic hands for palm brushing, silence, chaotic sounds, discontinuous sounds, etc. The training process uses deep learning-related technologies such as gradient descent to train and optimize the model. The second loss data can be constructed according to the difference between the sample sound prediction result and the above sample sound judgment result label; and the model parameters of the second preset model can be adjusted according to the second loss data until the training end condition is met; and the second preset model at the end of training is determined as the sound judgment model; thus facilitating the rapid and accurate obtaining of the second judgment result through the sound judgment model.

[0221] When a user performs palm brushing recognition, the background sound risk control module will analyze and judge the background sound to determine whether the background sound is reasonable. The background sound risk control module will extract features and classify the background sound, and then judge whether the background sound belongs to a normal palm brushing scenario or whether there are risks.

[0222] The following is an example of a background sound risk control module in a palm brushing payment scenario:

[0223] Suppose a company launches a palm - brushing payment function, and users can complete payments by brushing their palms in stores. To ensure security, the company introduces a background - sound risk - control module to identify whether the background sound during palm - brushing payment is reasonable. During the training process, this module uses the background sounds recorded from a large number of examples of users who are normally making palm - brushing payments, as well as the labeled device sounds, user sounds, and pure background sounds for training. At the same time, it also uses a large amount of sound training data for special scenarios, such as when users make palm - brushing payments in noisy environments, when the background sound is too noisy or interfering, etc.

[0224] S2075: Based on the above first judgment result and the above second judgment result, determine the target judgment result.

[0225] In the embodiments of this specification, the above - mentioned determining the target judgment result based on the above first judgment result and the above second judgment result includes:

[0226] According to the above - mentioned target task, determine the target acquisition scenario corresponding to the above - mentioned target acquisition image;

[0227] When the above first judgment result is a reasonable result, match the above - mentioned target background image with the above - mentioned target acquisition scenario to obtain a first matching result;

[0228] When the above second judgment result is a reasonable result, match the above - mentioned target background sound with the above - mentioned target acquisition scenario to obtain a second matching result;

[0229] Determine the target judgment result according to the above first matching result and the above second matching result.

[0230] In the embodiments of this specification, the target background image and the above - mentioned target background sound can be further matched according to the target acquisition scenario. When both the first matching result and the above - mentioned second matching result are successful matching results, determine that the target judgment result is a reasonable judgment result. When at least one of the first matching result and the above - mentioned second matching result is an unsuccessful matching result, determine that the target judgment result is an unreasonable judgment result.

[0231] In the embodiments of this specification, when the above first judgment result is a reasonable result, matching the above - mentioned target background image with the above - mentioned target acquisition scenario to obtain a first matching result includes:

[0232] When the above first judgment result is a reasonable result, query, based on a preset scenario image library, a reference scenario image that matches the target acquisition scenario; the preset scenario image library stores the mapping relationship between preset acquisition scenarios and preset reference scenario images;

[0233] Match the above - mentioned target background image with the above - mentioned reference scenario image to obtain a first matching result.

[0234] In the embodiments of the present specification, a preset scenario image library can be pre-constructed, which can store the mapping relationships between multiple preset acquisition scenarios and multiple preset reference scenario images. Among them, the preset acquisition scenarios and the preset reference scenario images can be in a one-to-one or one-to-many relationship. If the target background image matches the above reference scenario image, the first matching result is a reasonable result.

[0235] In the embodiments of the present specification, in the case where the above second judgment result is a reasonable result, the above target background sound is matched with the above target acquisition scenario to obtain a second matching result, including:

[0236] In the case where the above second judgment result is a reasonable result, based on the preset scenario sound library, a reference sound matching the target acquisition scenario is queried; the preset scenario sound library stores the mapping relationships between preset acquisition scenarios and preset reference scenario sounds;

[0237] The above target background sound is matched with the above reference sound to obtain a second matching result.

[0238] In the embodiments of the present specification, a preset scenario sound library can be pre-constructed, which can store the mapping relationships between multiple preset acquisition scenarios and multiple preset reference scenario sounds. Among them, the preset acquisition scenarios and the preset reference scenario sounds can be in a one-to-one or one-to-many relationship. If the target background sound matches the above target acquisition scenario, it is determined that the second matching result is a reasonable result. When both the first matching result and the second matching result are reasonable results, it is determined that the target judgment result is a reasonable result; if at least one of the first matching result and the second matching result is unreasonable, it is determined that the target judgment result is an unreasonable result.

[0239] In the embodiments of the present specification, the target background image and the target background sound can be further verified through the preset scenario image library and the preset scenario sound library, so as to improve the accuracy rate of the target judgment result.

[0240] S209: In the case where the above target judgment result indicates that both the above target background image and the above target background sound are reasonable, the above target biological image is recognized to obtain a target recognition result.

[0241] In the embodiments of the present specification, the above target biological image can be recognized to obtain a target recognition result in the case where the above target judgment result indicates that both the above target background image and the above target background sound are reasonable. The task can be stopped and a task execution failure message can be generated in the case where the above target judgment result indicates that at least one of the above target background image and the above target background sound is unreasonable.

[0242] S2011: When the above-mentioned target recognition result is a recognition pass result, execute the above-mentioned target task.

[0243] In the embodiments of this specification, when the above-mentioned target recognition result is a recognition pass result, the above-mentioned target task can be executed. Specifically, target tasks such as payment, entering and exiting the gate, and unlocking can be executed, thereby improving the security of task execution.

[0244] As can be seen from the technical solutions provided by the embodiments of this specification above, the embodiments of this specification respond to a target task execution request of a target account, obtain a target acquisition image of the target account and a target sound during the acquisition process of the target acquisition image; the target acquisition image includes the biometric features of the target account; perform image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image; perform background sound separation operation on the target sound to obtain a target background sound; perform a rationality judgment on the target background image and the target background sound to obtain a target judgment result; when the target judgment result indicates that both the target background image and the target background sound are reasonable, it means that both the background image and the background sound corresponding to the target account are reasonable, that is, both the background image and the background sound corresponding to the target account conform to the scene corresponding to the target task. At this time, the target biometric image is further recognized to obtain a target recognition result; thereby improving the accuracy of the recognition result; when the target recognition result is a recognition pass result, execute the target task; further improving the security of task execution.

[0245] The embodiments of this specification also provide a task execution device, as Figure 8 shown, the above-mentioned device includes:

[0246] A target information acquisition module 810, configured to respond to a target task execution request of a target account, and acquire the target acquisition image of the target account and a target sound during the acquisition process of the target acquisition image; the target acquisition image includes the biometric features of the target account;

[0247] A target image determination module 820, configured to perform image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image;

[0248] A target sound determination module 830, configured to perform background sound separation operation on the target sound to obtain a target background sound;

[0249] A judgment module 840, configured to perform a rationality judgment on the target background image and the target background sound to obtain a target judgment result;

[0250] A target recognition module 850, configured to recognize the target biological image when the above-mentioned target judgment result indicates that both the above-mentioned target background image and the above-mentioned target background sound are reasonable, so as to obtain a target recognition result;

[0251] A task execution module 860, configured to execute the above-mentioned target task when the above-mentioned target recognition result is a recognition pass result.

[0252] In some embodiments, the above-mentioned judgment module includes:

[0253] A first judgment unit, configured to perform a rationality judgment on the above-mentioned target background image to obtain a first judgment result;

[0254] A second judgment unit, configured to perform a rationality judgment on the above-mentioned target background sound to obtain a second judgment result;

[0255] A third judgment unit, configured to determine a target judgment result based on the above-mentioned first judgment result and the above-mentioned second judgment result.

[0256] In some embodiments, the above-mentioned first judgment unit includes:

[0257] A first judgment subunit, configured to perform a rationality judgment on the above-mentioned target background image based on a first preset strategy to obtain a first judgment result; the above-mentioned first preset strategy includes determining that the above-mentioned preset background image is unreasonable if at least one of a virtual background, a fictional biological feature background, a paper background, a synthetic background, and a posture-distorted background exists in the preset background image;

[0258] In some embodiments, the above-mentioned second judgment unit includes:

[0259] A second judgment subunit, configured to perform a rationality judgment on the above-mentioned target background sound based on a second preset strategy to obtain a second judgment result; the above-mentioned second preset strategy includes determining that the above-mentioned preset background sound is unreasonable if at least one of silence, multiple sounds, noise, and interrupted sound exists in the preset background sound.

[0260] In some embodiments, the above-mentioned first judgment subunit includes:

[0261] A first detection unit, configured to detect whether a virtual background exists in the above-mentioned target background image based on the above-mentioned first preset strategy;

[0262] A second detection unit, configured to detect whether a fictional biological feature background exists in the above-mentioned target background image;

[0263] A third detection unit, configured to detect whether a paper background exists in the above-mentioned target background image;

[0264] The fourth detection unit is used to detect whether there is a synthetic background in the above-mentioned target background image;

[0265] The fifth detection unit is used to detect whether there is a posture-distorted background in the above-mentioned target background image;

[0266] The sixth detection unit is used to determine that the above first judgment result is an unreasonable result if there is at least one of a virtual background, a fictional biometric background, a paper background, a synthetic background, and a posture-distorted background in the above-mentioned target background image.

[0267] In some embodiments, the above second judgment subunit includes:

[0268] The mute detection unit is used to detect whether there is a mute in the above-mentioned target background sound based on the above second preset strategy;

[0269] The multi-sound detection unit is used to detect whether there are multiple sounds in the above-mentioned target background sound;

[0270] The interrupted sound detection unit is used to detect whether there is an interrupted sound in the above-mentioned target background sound;

[0271] The result determination unit is used to determine that the above second judgment result is an unreasonable result if there is at least one of a mute, multiple sounds, noise, and an interrupted sound in the above-mentioned target background sound.

[0272] In some embodiments, the above third judgment unit includes:

[0273] The target scene determination subunit is used to determine the target acquisition scene corresponding to the above-mentioned target acquisition image according to the above-mentioned target task;

[0274] The first matching subunit is used to match the above-mentioned target background image with the above-mentioned target acquisition scene to obtain a first matching result when the above first judgment result is a reasonable result;

[0275] The second matching subunit is used to match the above-mentioned target background sound with the above-mentioned target acquisition scene to obtain a second matching result when the above second judgment result is a reasonable result;

[0276] The result determination subunit is used to determine a target judgment result according to the above first matching result and the above second matching result.

[0277] In some embodiments, the above first judgment unit includes:

[0278] The first prediction subunit is used to input the above-mentioned target background image into an image judgment model for reasonableness judgment to obtain a first judgment result; the above image judgment model is obtained by training a first preset model for reasonableness judgment based on sample background images;

[0279] In some embodiments, the second determination unit includes:

[0280] A second prediction subunit, configured to input the target background sound into a sound determination model for rationality determination to obtain a second determination result; the sound determination model is obtained by training a second preset model for rationality determination based on sample background sounds.

[0281] In some embodiments, the apparatus further includes:

[0282] A sample background image acquisition module, configured to acquire a sample background image, where the sample background image is labeled with a sample image determination result label;

[0283] A sample image prediction module, configured to input the sample background image into the first preset model for rationality determination to obtain a sample image prediction result;

[0284] A first parameter adjustment module, configured to adjust the model parameters of the first preset model based on the difference between the sample image prediction result and the sample image determination result label until the training end condition is satisfied;

[0285] An image model determination module, configured to determine the first preset model at the end of training as the image determination model.

[0286] In some embodiments, the apparatus further includes:

[0287] A sample background sound acquisition module, configured to acquire a sample background sound, where the sample background sound is labeled with a sample sound determination result label;

[0288] A sample sound prediction module, configured to input the sample background sound into the second preset model for rationality determination to obtain a sample sound prediction result;

[0289] A second parameter adjustment module, configured to adjust the model parameters of the first preset model based on the difference between the sample sound prediction result and the sample sound determination result label until the training end condition is satisfied;

[0290] A sound model determination module, configured to determine the second preset model at the end of training as the sound determination model.

[0291] In some embodiments, the target sound determination module includes:

[0292] A sound separation unit, configured to perform a background sound separation operation on the target sound to obtain a target device acquisition sound and a target account sound; the target device acquisition sound is the sound of an image acquisition device acquiring the target biological acquisition image;

[0293] A sound determination unit, configured to remove the sound collected by the target device and the sound of the target account from the above-mentioned target sound, so as to obtain a target background sound.

[0294] In some embodiments, the above-mentioned target image determination module includes:

[0295] An image processing unit, configured to perform image preprocessing operations on the above-mentioned target collected image to obtain a target processed image; the above-mentioned image preprocessing includes at least one of image denoising and image enhancement;

[0296] An image segmentation unit, configured to perform image segmentation processing on the above-mentioned target processed image.

[0297] In some embodiments, the above-mentioned target sound determination module includes:

[0298] A sound preprocessing unit, configured to perform sound preprocessing operations on the above-mentioned target sound to obtain a target processed sound; the above-mentioned sound preprocessing includes at least one of sound denoising and filtering processing;

[0299] A background sound separation unit, configured to perform background sound separation operations on the above-mentioned target processed sound.

[0300] In some embodiments, the above-mentioned target sound determination module includes:

[0301] A first segmentation unit, configured to perform image segmentation processing on the above-mentioned target collected image if the above-mentioned target collected image includes a palm brushing image, so as to obtain a target palm brushing image and a target background image;

[0302] A second segmentation unit, configured to perform image segmentation processing on the above-mentioned target collected image if the above-mentioned target collected image includes a face image, so as to obtain a target face image and a target background image.

[0303] The device in the above-mentioned device embodiment and the method embodiment are based on the same inventive concept.

[0304] An embodiment of this specification provides a task execution device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement the task execution method provided in the above-mentioned method embodiment.

[0305] An embodiment of this application also provides a computer storage medium. The above-mentioned storage medium can be set in a terminal to store at least one instruction or at least one program segment related to implementing a task execution method in the method embodiment. The at least one instruction or at least one program segment is loaded and executed by the processor to implement the task execution method provided in the above-mentioned method embodiment.

[0306] Embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the task execution method provided in the foregoing method embodiments.

[0307] Optionally, in the embodiments of the present specification, the storage medium may be located in at least one network server among multiple network servers of a computer network. Optionally, in this embodiment, the foregoing storage medium may include, but is not limited to, various media capable of storing program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc.

[0308] The memory in the embodiments of the present specification may be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for functions, etc.; the data storage area may store data created according to the use of the foregoing device, etc. In addition, the memory may include a high-speed random access memory, and may further include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory may further include a memory controller to provide the processor with access to the memory.

[0309] The task execution method embodiments provided in the embodiments of the present specification may be executed on a mobile terminal, a computer terminal, a server, or a similar computing device. Taking running on a server as an example, Figure 9 is a hardware structure block diagram of a server for a task execution method provided in the embodiments of the present specification. As Figure 9As shown, the server 900 can vary significantly due to differences in configuration or performance, and may include one or more central processing units (CPUs) 910 (the central processing unit 910 may include, but is not limited to, processing devices such as microprocessor MCUs or programmable logic devices FPGAs), a memory 930 for storing data, and one or more storage media 920 for storing application programs 923 or data 922 (such as one or more mass storage devices). Among them, the memory 930 and the storage media 920 can be transient storage or persistent storage. The program stored in the storage media 920 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processing unit 910 can be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the server 900. The server 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0310] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the server 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 940 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0311] Those of ordinary skill in the art can understand that Figure 9 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the server 900 may also include more or fewer components than Figure 9 shown in the figure, or have a different configuration from Figure 9 that shown in the figure.

[0312] As can be seen from the embodiments of the task execution method, apparatus, device, or storage medium provided by the present application above, in response to a target task execution request of a target account, the present application acquires a target capture image of the target account and a target sound during the capture process of the target capture image; the target capture image includes the biometric features of the target account; performs image segmentation processing on the target capture image to obtain a target biometric image and a target background image; performs background sound separation operation on the target sound to obtain a target background sound; performs a rationality judgment on the target background image and the target background sound to obtain a target judgment result; in the case where the target judgment result indicates that both the target background image and the target background sound are reasonable, it indicates that both the background image and the background sound corresponding to the target account are reasonable, that is, both the background image and the background sound corresponding to the target account conform to the scenario corresponding to the target task. At this time, the target biometric image is further recognized to obtain a target recognition result; thereby improving the accuracy of the recognition result; in the case where the target recognition result is a recognition pass result, the target task is executed; further improving the security of task execution.

[0313] It should be noted that: the above sequence of the embodiments of this specification is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0314] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, device, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0315] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer storage medium. The above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.

[0316] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A task execution method, characterized in that, The method includes: In response to a target task execution request of a target account, obtaining a target acquisition image of the target account and a target sound during the acquisition process of the target acquisition image; the target acquisition image includes biometric features of the target account; Performing image segmentation processing on the target acquisition image to obtain a target biometric image and a target background image; Performing background sound separation operation on the target sound to obtain a target background sound; Performing a rationality judgment on the target background image and the target background sound to obtain a target judgment result; When the target judgment result indicates that both the target background image and the target background sound are reasonable, performing identification on the target biometric image to obtain a target identification result; When the target identification result is a passed identification result, executing the target task.

2. The method according to claim 1, characterized in that, The performing a rationality judgment on the target background image and the target background sound to obtain a target judgment result includes: Performing a rationality judgment on the target background image to obtain a first judgment result; Performing a rationality judgment on the target background sound to obtain a second judgment result; Determining the target judgment result based on the first judgment result and the second judgment result.

3. The method according to claim 2, wherein The performing a rationality judgment on the target background image to obtain a first judgment result includes: Performing a rationality judgment on the target background image based on a first preset policy to obtain a first judgment result; the first preset policy includes determining that the preset background image is unreasonable if there is at least one of a virtual background, a fictional biometric feature background, a paper background, a synthetic background, and a posture distortion background in the preset background image; The performing a rationality judgment on the target background sound to obtain a second judgment result includes: Performing a rationality judgment on the target background sound based on a second preset policy to obtain a second judgment result; the second preset policy includes determining that the preset background sound is unreasonable if there is at least one of silence, multiple sounds, noise, and interrupted sound in the preset background sound.

4. The method according to claim 3, wherein The performing a rationality judgment on the target background image based on the first preset policy to obtain a first judgment result includes: Detecting whether there is a virtual background in the target background image based on the first preset policy; Detecting whether there is a fictional biometric feature background in the target background image; Detecting whether there is a paper background in the target background image; Detecting whether there is a synthetic background in the target background image; Detecting whether there is a posture distortion background in the target background image; If there is at least one of a virtual background, a fictional biometric feature background, a paper background, a synthetic background, and a posture distortion background in the target background image, determining that the first judgment result is an unreasonable result.

5. The method according to claim 3, characterized in that, The performing a rationality judgment on the target background sound based on the second preset policy to obtain a second judgment result includes: Detecting whether there is silence in the target background sound based on the second preset policy; Detecting whether there are multiple sounds in the target background sound; Detecting whether there is an interrupted sound in the target background sound; If at least one of silence, multiple sounds, noise, and interrupted sound exists in the target background sound, determine that the second judgment result is an unreasonable result.

6. The method according to claim 4 or 5, characterized in that, Determining a target judgment result based on the first judgment result and the second judgment result includes: Determine a target acquisition scene corresponding to the target acquisition image according to the target task; When the first judgment result is a reasonable result, match the target background image with the target acquisition scene to obtain a first matching result; When the second judgment result is a reasonable result, match the target background sound with the target acquisition scene to obtain a second matching result; Determine a target judgment result according to the first matching result and the second matching result.

7. The method according to claim 2, characterized in that, Performing a rationality judgment on the target background image to obtain a first judgment result includes: Input the target background image into an image judgment model for rationality judgment to obtain a first judgment result; the image judgment model is obtained by training a first preset model for rationality judgment based on sample background images; Performing a rationality judgment on the target background sound to obtain a second judgment result includes: Input the target background sound into a sound judgment model for rationality judgment to obtain a second judgment result; the sound judgment model is obtained by training a second preset model for rationality judgment based on sample background sounds.

8. The method according to claim 7, wherein The training method of the image judgment model includes: Obtain a sample background image, and the sample background image is labeled with a sample image judgment result label; Input the sample background image into the first preset model for rationality judgment to obtain a sample image prediction result; Based on the difference between the sample image prediction result and the sample image judgment result label, adjust the model parameters of the first preset model until the training end condition is met; Determine the first preset model at the end of training as the image judgment model.

9. The method according to claim 7, characterized in that, The training method of the sound judgment model includes: Obtain a sample background sound, and the sample background sound is labeled with a sample sound judgment result label; Input the sample background sound into the second preset model for rationality judgment to obtain a sample sound prediction result; Based on the difference between the sample sound prediction result and the sample sound judgment result label, adjust the model parameters of the first preset model until the training end condition is met; Determine the second preset model at the end of training as the sound judgment model.

10. The method according to claim 1, characterized in that, Performing a background sound separation operation on the target sound to obtain a target background sound includes: Perform a background sound separation operation on the target sound to obtain a target device acquisition sound and a target account sound; the target device acquisition sound is the sound of the image acquisition device acquiring the target biological acquisition image; Eliminate the target device acquisition sound and the target account sound in the target sound to obtain a target background sound.

11. The method according to claim 1, characterized in that, Performing an image segmentation process on the target acquisition image includes: Perform image preprocessing operations on the target captured image to obtain a target processed image; the image preprocessing includes at least one of image denoising and image enhancement; Perform image segmentation processing on the target processed image.

12. The method according to claim 1, characterized in that, The operation of separating background sound from the target sound includes: Perform sound preprocessing operations on the target sound to obtain a target processed sound; the sound preprocessing includes at least one of sound denoising and filtering processing; Perform background sound separation operations on the target processed sound.

13. According to the method described in claim 1, the operation of performing image segmentation processing on the target captured image to obtain a target biological image and a target background image includes: If the target captured image includes a palm brushing image, perform image segmentation processing on the target captured image to obtain a target palm brushing image and a target background image; If the target captured image includes a face image, perform image segmentation processing on the target captured image to obtain a target face image and a target background image.

14. A task execution device, characterized in that, The device includes: A target information acquisition module, configured to acquire the target captured image of the target account and the target sound during the acquisition process of the target captured image in response to a target task execution request of the target account; the target captured image includes the biological characteristics of the target account; A target image determination module, configured to perform image segmentation processing on the target captured image to obtain a target biological image and a target background image; A target sound determination module, configured to perform background sound separation operations on the target sound to obtain a target background sound; A judgment module, configured to perform a rationality judgment on the target background image and the target background sound to obtain a target judgment result; A target recognition module, configured to recognize the target biological image to obtain a target recognition result when the target judgment result indicates that both the target background image and the target background sound are reasonable; A task execution module, configured to execute the target task when the target recognition result is a recognition pass result.

15. A computer storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the task execution method described in any one of claims 1-13.