A live body attack prevention method, device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202310154331.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2026-08-07
- Estimated Expiration
- 2043-02-16
AI Technical Summary
例如刷脸支付,面部门禁,面部考勤以及面部进站等技术都需要依赖生物识别,但是,随着生物识别技术越来越广泛的应用,生物识别场景下的活体检测需求也越来越凸出,例如面部考勤、刷脸进站、刷脸支付等生物识别场景得到了广泛应用,在生物识别为人们提供方便的同时,也带来了新的风险挑战
[0016] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
Smart Images

Figure CN116246322B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method, device, storage medium, and electronic device for preventing live attacks. Background Technology
[0002] With the rapid development of computer technology, biometric technology has been widely applied to people's production and daily lives. For example, facial recognition payment, facial access control, facial attendance, and facial recognition station entry all rely on biometrics. However, as biometric technology becomes more widely used, the need for liveness detection in biometric scenarios is becoming increasingly prominent. While facial attendance, facial recognition station entry, and facial recognition payment have been widely adopted, they have also brought new risks and challenges. The most common means of threatening the security of biometric systems is liveness attacks, which involve attempting to bypass image biometric verification through means such as device screens or printed photos. To detect liveness attacks, liveness prevention technology has become an essential component of biometric scenarios. Summary of the Invention
[0003] This specification provides a method, device, storage medium, and electronic device for preventing liveness attacks. The technical solution is as follows:
[0004] Firstly, this specification provides a method for preventing attacks on living individuals, the method comprising:
[0005] Collect the first object image and the first environmental audio of the target object in the surrounding environment;
[0006] Based on the first object image, the first environmental speech is subjected to speech enhancement processing to obtain the second environmental speech;
[0007] Based on the first object image, the target object is subjected to liveness detection processing to obtain the liveness detection result.
[0008] Based on the liveness detection results, the second environmental voice and the first object image are used to determine the voice alarm level for the environment, and the voice alarm level is used to perform voice anti-attack alarm.
[0009] Secondly, this specification provides a live anti-attack device, the device comprising:
[0010] The data acquisition module is used to acquire the first object image and the first environmental voice of the target object in the environment.
[0011] The data processing module is used to perform speech enhancement processing on the first environmental speech based on the first object image to obtain the second environmental speech;
[0012] The liveness detection module is used to perform liveness detection processing on the target object based on the first object image to obtain a liveness detection result. Based on the liveness detection result, the module uses the second environmental voice and the first object image to determine the voice alarm level for the environment and uses the voice alarm level to perform a voice anti-attack alarm.
[0013] Thirdly, this specification provides a computer storage medium storing at least one instruction adapted for loading by a processor and executing method steps of one or more embodiments of this specification.
[0014] Fourthly, this specification provides a computer program product storing at least one instruction adapted to be loaded by a processor and to execute the method steps of one or more embodiments of this specification.
[0015] Fifthly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the method steps of one or more embodiments of this specification.
[0016] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0017] In one or more embodiments of this specification, an electronic device can acquire a first object image and a first ambient voice of a target object in its environment. Based on the first object image, it performs voice enhancement processing on the first ambient voice to obtain a second ambient voice. Based on the first object image, it performs liveness detection processing on the target object to obtain a liveness detection result. Based on the liveness detection result, it can use the second ambient voice and the first object image to determine the voice alarm level for voice attack prevention warning. This achieves a closed loop from liveness attack detection to system attack prevention, optimizing the situation in related technologies where only relevant objects are reminded to retry while ignoring system risks, thus reducing system attack risks and providing an attack prevention warning effect. Furthermore, based on the enhanced ambient voice and image, the voice alarm level that matches the current environment can be accurately determined, resulting in a better system attack prevention warning effect. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of a live anti-attack system provided in this manual;
[0020] Figure 2 This is a flowchart illustrating a live target anti-attack method provided in this manual;
[0021] Figure 3 This is a schematic diagram of model training for a cross-modal data augmentation model provided in this manual;
[0022] Figure 4 This is a schematic diagram of the model training for a liveness defense model provided in this manual;
[0023] Figure 5 This is a schematic diagram of the model training for an initial voice warning level model provided in this manual;
[0024] Figure 6 This is a schematic diagram of another live anti-attack device provided in this manual;
[0025] Figure 7 This is a schematic diagram of the structure of an electronic device provided in this specification;
[0026] Figure 8 This is a schematic diagram of the operating system and user space provided in this manual;
[0027] Figure 9 yes Figure 8 Architecture diagram of the Android operating system in China;
[0028] Figure 10 yes Figure 8 Architecture diagram of the iOS operating system. Detailed Implementation
[0029] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0030] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0031] In related technologies, liveness detection is performed by acquiring images of objects (such as users) in their natural state to determine if an attack is imminent. If an attack is identified, the liveness detection fails, and subsequent actions typically only prompt the relevant object (such as the user) to retry. This gives attackers more opportunities to attempt attacks, increasing system risk.
[0032] The present specification will now be described in detail with reference to specific embodiments.
[0033] Please see Figure 1 This is a schematic diagram of a live anti-attack system provided in this specification. Figure 1 As shown, the liveness detection system may include at least a client cluster and a service platform 100.
[0034] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0035] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0036] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0037] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster. Based on this communication connection, data interaction is completed during the liveness detection process, such as online transaction data interaction. For example, the client can collect a first object image and a first environmental voice of the target object in its environment and send them to the service platform 100. The service platform 100 then executes the liveness detection method corresponding to one or more embodiments of this specification to perform voice anti-attack warnings. Alternatively, the service platform 100 can instruct the client to execute the liveness detection method corresponding to one or more embodiments of this specification to perform voice anti-attack warnings.
[0038] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0039] The liveness detection system embodiments provided in this specification and the liveness detection methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the liveness detection method in one or more embodiments of this specification can be an electronic device, such as the aforementioned service platform 100; the execution entity electronic device corresponding to the liveness detection method in one or more embodiments of this specification can also be a client, depending on the actual application environment. The implementation process of the liveness detection system embodiments can be detailed in the following method embodiments, and will not be repeated here.
[0040] based on Figure 1 The following is a detailed description of the liveness protection method provided by one or more embodiments of this specification, illustrated in the scenario diagram.
[0041] Please see Figure 2 This document provides a flowchart illustrating a method for preventing liveness detection attacks, tailored to one or more embodiments. This method can be implemented using a computer program and can run on a liveness detection attack device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The liveness detection attack device can be a service platform.
[0042] Specifically, this live-target defense method includes:
[0043] Understandably, liveness detection is an application based on liveness detection, a method used in identity verification scenarios to determine the true physiological characteristics of an object. In liveness detection applications, at least the acquired target liveness detection image must be verified to confirm whether the image represents a real, living object. Image liveness detection needs to effectively resist common liveness attack methods such as photos, face swaps, masks, occlusion, and screen captures. Therefore, liveness detection processing is performed based on the image liveness detection results to help users identify fraudulent activities and protect their rights.
[0044] S102: Acquire the first object image and the first environmental speech of the target object in the environment;
[0045] The target object can be a person, an animal, or other similar object;
[0046] The first object image can be understood as object image data collected for the target object (such as a user, animal, etc.) in a liveness detection scenario;
[0047] In one or more embodiments of this specification, the image modality type corresponding to the first object image can be a fit of one or more of the following image modality types: video modality type, color image modality (rgb) type, short video modality type, animation modality type, depth image modality type, infrared image modality type, near-infrared modality (NIR) type, etc.
[0048] It should be noted that the image modality type of the first object image is not limited in this specification, and is determined based on the actual liveness detection application environment.
[0049] The first environmental speech is the speech information obtained after collecting the speech data of the environment when collecting the first object image of the target object in the environment.
[0050] Optionally, the voice acquisition time of the first environmental speech and the image acquisition time of the first object image can be the same or different.
[0051] Understandably, when the target object of an electronic device triggers the authentication function, the electronic device can collect the first object image and the first environmental voice of the target object in the environment through corresponding information acquisition devices (such as image acquisition devices and audio acquisition devices).
[0052] S104: Perform speech enhancement processing on the first environmental speech based on the first object image to obtain the second environmental speech;
[0053] In one or more embodiments of this specification, not only are first object images (image type) and first environmental speech (environmental speech type) acquired, but the first environmental speech is also enhanced in terms of speech quality based on the first object image. The second environmental speech with enhanced speech quality is then combined with the first object image to help determine the speech alarm level that matches the environment of the target object, so as to output a liveness detection and anti-attack reminder using a speech reminder method with the corresponding speech alarm level.
[0054] In a schematic manner, by extracting image features from the first object image, identifying the image environment information of the environment based on the image features, and adaptively adjusting the environmental speech components of the first environmental speech according to the image environment information, a second environmental speech with adjusted quality can be obtained. The second environmental speech with adjusted quality can reduce noise error components and increase the proportion of environmental information components.
[0055] Optionally, the aforementioned image environment information can be understood as the environment type of the current environment, such as a public environment, a private environment, an office environment, etc.
[0056] Optionally, adaptive quality adjustment of the environmental speech components of the first environmental speech can be understood as adjusting the speech information of the first environmental speech, including but not limited to pitch, intensity, duration, and timbre. By combining image environmental information to adjust the speech information of the first environmental speech, the quality of its environmental speech components can be enhanced. This can reduce objective error interference (such as noise and equipment influence) during the speech acquisition process, thereby enhancing the quality of the speech components in the first environmental speech.
[0057] If the aforementioned image environment information indicates that it is a private environment, then the voice information of the first environment voice should be adjusted to better match the audio characteristics of the private environment.
[0058] If the aforementioned image environment information indicates that it is an office environment, then the voice information of the first environment voice should be adjusted to better match the audio characteristics of the office environment.
[0059] Optionally, a voice information mode adjustment method can be adopted, which pre-constructs a mapping relationship between image environment information (such as which type of environment) and voice information adjustment mode (such as a certain type of environment can correspond to a certain voice information adjustment mode). The target voice information adjustment mode corresponding to the image environment information is determined through the adjustment mode mapping relationship, and the voice information of the first environment voice is adjusted by using the target voice information adjustment mode.
[0060] Optionally, a cross-modal data augmentation model based on a machine learning model is trained for speech quality enhancement. The cross-modal data augmentation model takes the first object image and the first environmental speech as input, and combines the image features of the first object image with the cross-modal data augmentation model to enhance the speech quality of the first environmental speech, thereby outputting the second environmental speech.
[0061] S106: Perform liveness detection processing on the target object based on the first object image to obtain the liveness detection result;
[0062] The liveness detection result may consist of at least a liveness attack result that reflects the liveness attack capability and a system risk result that reflects the damage of the liveness attack to the system;
[0063] In related technologies, liveness detection only provides feedback on the probability of a liveness attack, which is a type of liveness attack. It cannot directly measure the degree of damage that the liveness attack causes to the system or the risk of the system's defense against liveness attacks.
[0064] In one feasible implementation, a liveness detection model can be pre-built and trained based on a machine learning model. The liveness detection model is then used to perform liveness detection and system risk detection processing based on a first object image to obtain the liveness detection result.
[0065] It should be noted that the machine learning models involved in one or more embodiments of this specification include, but are not limited to, fitting of one or more of the following machine learning models: Convolutional Neural Network (CNN) model, Deep Neural Network (DNN) model, Recurrent Neural Networks (RNN) model, embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, Residual Network (ResNet) model, etc.
[0066] S108: Based on the liveness detection result, the second environmental voice and the first object image are used to determine the voice alarm level for the environment, and the voice alarm level is used to perform voice anti-attack alarm.
[0067] The voice alarm level is used to indicate what type of voice reminder method the device uses to provide a reminder.
[0068] Indicatively, the voice alarm level can correspond to a corresponding voice prompt method. By determining the voice alarm level and the corresponding voice prompt method, voice anti-attack alarms can be issued according to the voice prompt method. For example, the voice alarm level can be used to indicate the use of no voice prompt method, low decibel voice prompt method, medium decibel voice prompt method, high decibel voice warning, etc.
[0069] Optionally, based on the liveness detection result, a voice anti-attack warning is determined. This involves extracting the first image features of the first object image, performing environmental resolution processing based on the extracted first image features and the enhanced second environmental voice, determining the current environment type, and identifying the appropriate voice warning level for the current environment type. In other words, the voice warning level corresponding to the current environment type is determined, and a reminder is given based on the voice reminder method corresponding to the voice warning level.
[0070] In one or more embodiments of this specification, data is collected and preprocessed. In addition to collecting a first object image of the target object, a second environmental voice is also collected. The image information is used for liveness detection and risk classification to obtain liveness anti-attack detection results. Based on the liveness anti-attack detection results, an assessment of the current system attack risk is introduced in addition to the liveness attack risk assessment of the target object, thus optimizing the liveness attack detection effect and achieving a better liveness anti-attack level system alarm. Furthermore, the environmental components in the first environmental voice are enhanced based on the first object image, which can assist in the subsequent accurate determination of the voice alarm level and the implementation of voice alarm.
[0071] Optionally, the step of performing speech enhancement processing on the first environmental speech based on the first object image to obtain the second environmental speech can be:
[0072] The first object image and the first environmental speech are input into a cross-modal data augmentation model for cross-modal speech augmentation processing, and the second environmental speech is output.
[0073] Understandably, after acquiring the first object image and the first environmental speech, a pre-trained cross-modal data augmentation model can be used to input the first object image and the first environmental speech into the cross-modal data augmentation model to perform cross-modal speech augmentation processing on the first environmental speech and enhance the environmental components in the speech information of the first environmental speech.
[0074] For illustrative purposes, please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating the model training of a cross-modal data augmentation model proposed in one or more embodiments of this specification. Specifically:
[0075] S202: Create an initial cross-modal data augmentation model;
[0076] In one or more embodiments of this specification, in response to a cross-modal data augmentation task, an initial cross-modal data augmentation model can be created based on a machine learning model, and the model can be trained to obtain a cross-modal data augmentation model. The cross-modal data augmentation model can then be used to enhance the data quality of environmental speech.
[0077] S204: Obtain the first sample object image and the first sample environment speech, and perform noise addition processing on the first sample environment speech to obtain the second sample environment speech;
[0078] The first sample object image is the sample image data of the initial cross-modal data augmentation model during the model training phase. It can usually be obtained by using an image acquisition device (such as a camera) to acquire the sample object image of the sample object in the sample environment.
[0079] The first sample environment speech is the sample environment speech data of the initial cross-modal data augmentation model during the model training phase. It can usually be obtained by using a speech acquisition device (such as a microphone) to collect the speech data of the sample environment in which the sample object is located.
[0080] As an illustration, in order to better assist the training of the initial cross-modal data augmentation model, the first sample environmental speech is not used directly. Instead, noise components are added to it to generate the second sample environmental speech. Based on the noise-added second sample environmental speech, the model convergence can be accelerated, the model's adaptability to complex scenes can be improved, and the robustness of the model can be enhanced.
[0081] Furthermore, noise-adding techniques from related technologies can be used to add noise to the first sample environmental speech to obtain the second sample environmental speech.
[0082] S206: Input the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model for training until the initial cross-modal data augmentation model completes training, and obtain the trained cross-modal data augmentation model.
[0083] To illustrate, by acquiring multiple sets of first sample object images and second sample environmental speech for model training, a well-trained cross-modal data augmentation model can be obtained after the initial cross-modal data augmentation model has completed its training.
[0084] In one or more embodiments of this specification, the model termination condition may include, for example, the value of the loss function being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. The specific model termination condition can be determined based on actual circumstances and is not specifically limited here.
[0085] Schematic: The step of inputting the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model for training can be:
[0086] A2: Input the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model, determine the first sample image features and the first sample speech features based on the feature encoder of the initial cross-modal data augmentation model, determine the first sample fusion features corresponding to the first sample image features and the first sample speech features through the feature fusion module of the initial cross-modal data, and perform quality enhancement and denoising processing on the first sample environmental speech based on the first sample fusion features through the quality enhancement module of the initial cross-modal data to obtain sample environmental enhanced speech;
[0087] The internal model structure of an initial cross-modal data augmentation model can be constructed based on a machine learning model / network, for example, the model structure of the initial cross-modal data augmentation model can include at least three parts: the first part is a feature encoder, the second part is a feature fusion module, and the third part is a quality enhancement module.
[0088] The feature encoder processes the first sample object image and the second sample environment speech. The feature encoder can extract the first sample image features corresponding to the first sample object image and the first sample speech features corresponding to the first sample environment speech. It can be understood that the processing result of the feature encoder is the first sample image features and the first sample speech features.
[0089] The processing object of the feature fusion module is the processing result of the feature encoder, namely the first sample image features and the first sample speech features. By fusing the first sample image features and the first sample speech features through the feature fusion module, the first sample fused features can be obtained. It can be understood that the processing result of the feature fusion module is the first sample fused features.
[0090] The quality enhancement module processes the result of the feature fusion module, namely the first sample fusion feature. The quality enhancement module identifies the image environment information of the environment based on the first sample fusion feature, and performs adaptive quality adjustment and noise removal on the environmental speech components based on the image environment information, so as to generate a quality-enhanced and noise-removed sample environment-enhanced speech based on the first sample fusion feature.
[0091] A4: Calculate the first model loss based on the sample environment enhanced speech and the first sample environment speech, and adjust the model parameters of the initial cross-modal data augmentation model based on the first model loss.
[0092] Schematic, the calculation of the first model loss based on the enhanced speech of the sample environment and the first sample environment speech can be:
[0093] The speech denoising loss of the sample environment enhanced speech and the first sample environment speech is calculated using the first loss calculation formula, and the speech denoising loss is used as the first model loss.
[0094] The first loss calculation formula satisfies the following formula:
[0095] La = L2(x, x_recovery)
[0096] Wherein, La is the speech denoising loss, x is the first sample environment speech, x_recovery is the sample environment enhanced speech, and L2 represents the Euclidean distance operator.
[0097] For example, the Euclidean distance operation in the first loss calculation formula can also be expressed in the following form:
[0098]
[0099] Understandably, during the initial training of the cross-modal data augmentation model in each round, the electronic device inputs the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model to obtain the sample environmental augmented speech for each round. Then, the model loss is calculated based on the first sample environmental speech and the sample environmental augmented speech. Directly using the first sample environmental speech for quality enhancement model training makes it difficult to measure which quality enhancement target speech is better, and the model loss is difficult to represent well. This step transforms this by using noisy second sample environmental speech for model quality enhancement and denoising. The denoising effect of the processing result is used to measure the model's instruction augmentation effect, which can better represent the model loss. The model parameters of the initial cross-modal data augmentation model are adjusted by backpropagation based on the first model loss of each round until the model training termination condition is met, resulting in a well-trained cross-modal data augmentation model. This optimizes the model training process, reduces the resource consumption of model training, and accelerates the model convergence effect.
[0100] Optionally, the liveness detection result can be the liveness attack probability and the system risk level. The step of performing liveness detection processing on the target object based on the first object image to obtain the liveness detection result can be:
[0101] The electronic device inputs the first object image into the liveness detection model to perform liveness detection processing, and obtains the liveness attack probability P and the system risk level L.
[0102] Optionally, the liveness detection result can be the liveness attack probability, system risk level, and object image features. The step of performing liveness detection processing on the target object based on the first object image to obtain the liveness detection result can be:
[0103] The electronic device inputs the first object image into the liveness detection model to perform liveness detection processing, and obtains the liveness attack probability P, the system risk level L, and the object image features;
[0104] Furthermore, after obtaining the liveness detection result, the process of determining the voice alarm level for the environment based on the liveness detection result and using the second environmental voice and the first object image can be as follows:
[0105] 1) If the probability of a live attack is greater than the target probability threshold, then the object image features corresponding to the first object image are obtained, and the second environmental voice and the object image features are used to determine the voice warning level for the environment.
[0106] 2) If the system risk level is greater than the target level threshold, then obtain the object image features corresponding to the first object image, and use the second environmental voice and the object image features to determine the voice warning level for the environment.
[0107] 3) If the probability of a live attack is greater than the target probability threshold and the system risk level is greater than the target level threshold, then obtain the object image features corresponding to the first object image, and use the second environmental voice and the object image features to determine the voice warning level for the environment.
[0108] Wherein, the target probability threshold is a threshold or critical value set for the probability P of a live attack, and the target level threshold is a threshold or critical value set for the system risk level L.
[0109] As an illustration, electronic devices can directly obtain object image features based on the liveness detection model; or they can obtain the corresponding object image features by extracting features from the first object image.
[0110] For illustrative purposes, please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating the training of a liveness detection and anti-attack model described in this manual. Specifically:
[0111] S302: Create an initial liveness detection model;
[0112] In related technologies, traditional liveness detection models only determine the probability of a liveness attack, but cannot determine the probability of the attack breaching the system, i.e., the system risk level.
[0113] Indicatively, an initial liveness detection model is created in advance based on a machine learning model, and the internal model structure of the initial liveness detection model is constructed. For example, the ResNet18 residual network model can be used to construct the internal model structure of the initial liveness detection model.
[0114] S304: Obtain the second sample object image and the corresponding liveness classification label, attack classification label and risk level classification label of the second sample object image; input the second sample object image into the initial liveness anti-attack model; and determine the sample liveness classification result, attack classification result, risk level classification result and sample image features through the initial liveness anti-attack model.
[0115] In a schematic way, this step breaks down the traditional liveness attack classification result into two dimensions. The first is the liveness classification result of the liveness classification dimension, which is the liveness classification result predicted by the initial liveness anti-attack model from the liveness classification dimension. The liveness classification result is the liveness classification probability obtained from the liveness classification dimension. The second is the attack classification result of the attack classification dimension, which is the attack classification result predicted by the initial liveness anti-attack model from the attack classification dimension. The attack classification result is the attack classification probability obtained from the liveness classification dimension.
[0116] It should be noted that this step controls the initial liveness detection model to refine the classification characteristics from both liveness and attack classification dimensions, so that the model can have better separability and processing effect in liveness and attack classification. This optimizes the limitations of related technologies in the training phase of the liveness detection model, such as overfitting and long training time, which may be caused by using a single liveness attack probability.
[0117] Based on this, in the data labeling stage before model training, the input sample data, i.e., the second sample object image, is labeled with a liveness classification label from the liveness classification dimension and an attack classification label from the attack classification dimension.
[0118] Furthermore, in the data annotation stage, the corresponding second sample object image is labeled with a risk level label from the dimension of system attack risk. The risk level label is also the label of the specific system risk level.
[0119] Optionally, the sample image features are the image features of the second sample object image. Here, the sample image features are output separately in each round of model training to provide input data to assist in the subsequent processing stage of determining the alarm and the voice alarm level, which can avoid the secondary acquisition of image features in the deployment and application stage.
[0120] S306: The initial liveness anti-attack model is trained based on the sample liveness classification result, the liveness classification label, the attack class classification result, the attack classification label, the risk level classification result, and the risk level classification label until the initial liveness anti-attack model is trained, thus obtaining the trained liveness anti-attack model.
[0121] Schematic illustration: Training the initial liveness anti-attack model based on the sample liveness classification result, the liveness classification label, the attack class classification result, the attack classification label, the risk level classification result, and the risk level classification label can be as follows:
[0122] B2: Determine the liveness classification loss based on the sample liveness classification result and the liveness classification label; determine the attack classification loss based on the attack class classification result and the attack classification label; and determine the risk level prediction loss based on the risk level classification result and the risk level classification label.
[0123] Indicatively, the sample liveness classification result and the liveness classification label are input into the second loss calculation formula to obtain the liveness classification loss, the attack class classification result and the attack classification label are input into the second loss calculation formula to determine the attack classification loss, and the risk level classification result and the risk level classification label are input into the second loss calculation formula to determine the risk level prediction loss.
[0124] The second loss calculation formula satisfies the following formula:
[0125] Lb = CrossEntropy(f,y)
[0126] Wherein, Lb is the classification loss symbol corresponding to the second loss calculation formula, f is the classification result, y is the classification label, and CrossEntropy is the cross-entropy operator.
[0127] B4: Determine the second model loss based on the liveness classification loss, the attack classification loss, and the risk level prediction loss, and adjust the model parameters of the initial liveness anti-attack model based on the first model loss.
[0128] Furthermore, the liveness classification result of the sample, f, and the liveness classification label, y, are substituted into the second loss calculation formula above to calculate the liveness classification loss, Loss. live ;
[0129] Furthermore, the attack classification result of the sample, f, and the attack classification label, y, are substituted into the second loss calculation formula above to calculate the attack classification loss, Loss. attack ;
[0130] Furthermore, the risk level classification result (f) and the risk level classification label (y) are substituted into the second loss calculation formula to calculate the risk level predicted loss (Loss). risk ;
[0131] Specifically, the second model loss of the initial liveness defense model can be characterized by the following formula:
[0132] Loss total =Loss live +Loss attack +Loss risk
[0133] Wherein, the Loss total This represents the model loss of the initial liveness defense model.
[0134] The initial liveness prevention model is trained multiple times based on the above model structure and the second model loss to adjust the model parameters until the model ends training. After that, the trained liveness prevention model can be obtained.
[0135] In one or more embodiments of this specification, the traditional liveness attack classification is broken down and refined during the model training phase of the liveness attack prevention model. This refines the classification characteristics from both liveness and attack classification dimensions, enabling the model to achieve better separability and processing performance in both liveness and attack classification. This optimizes the limitations of related technologies that rely on a single liveness attack probability during the training phase of the liveness attack model, such as overfitting and long training times. Furthermore, it fully utilizes the image characteristics of different dimensions, enabling image liveness detection to possess excellent feature representation effects. This assists the liveness detection model in subsequent accurate liveness detection classification, achieving better liveness attack detection results. Additionally, the system risk level characteristic is introduced, allowing for the assessment of the probability of a liveness attack breaching the system, thus optimizing the shortcomings of related liveness detection processes in system risk prediction.
[0136] Optionally, determining the voice alarm level for the environment using the second environmental voice and the first object image can be:
[0137] The electronic device acquires the object image features corresponding to the first object image, inputs the second environmental voice and the object image features into the voice alarm level model, and outputs the voice alarm level for the environment.
[0138] Understandably, based on the liveness detection results, the electronic device determines to issue a voice anti-attack warning. It acquires the first image features of the first object image, and based on the acquired first image features and the enhanced second environmental voice input voice warning level model, it performs environmental calculation processing through the voice warning level model to determine the current environment type and determine what level of voice warning is required for the current environment type. That is, it determines the voice warning level corresponding to the current environment type and outputs the voice warning level. The electronic device then issues a voice reminder warning based on the voice reminder method corresponding to the voice warning level.
[0139] For illustrative purposes, please refer to Figure 5 , Figure 5 This is a schematic diagram illustrating the model training of an initial voice warning level model as described in this manual. Specifically:
[0140] S402: Create an initial voice warning level model;
[0141] In one or more embodiments of this specification, an initial voice alarm level model can be created in advance based on a machine learning model in response to a voice alarm task. Subsequently, the initial voice alarm level model is trained. Once the initial voice alarm level model meets the model end training conditions, a trained voice alarm level model can be obtained.
[0142] In one or more embodiments of this specification, the model termination condition may include, for example, the value of the loss function being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. The specific model termination condition can be determined based on actual circumstances and is not specifically limited here.
[0143] S404: Obtain the sample environment enhanced speech, sample image features and sample speech warning level labels, and input the sample environment enhanced speech and the sample image features into the initial speech warning level model to obtain the predicted speech warning level;
[0144] An illustrative internal model structure for building an initial voice warning level model based on a machine learning model / network: The model structure can include at least three parts: the first part is a voice feature encoder, the second part is a feature fusion module (transformer), and the third part is a warning level prediction module.
[0145] In a schematic manner, the sample environment-enhanced speech and the sample image features are input into the initial speech warning level model, so that the speech feature encoder based on the initial speech warning level model determines the second sample speech features, the feature fusion module based on the initial speech warning level model determines the second sample fusion feature corresponding to the second sample speech features and the sample image features, and the warning level prediction module based on the initial speech warning level model determines the predicted speech warning level corresponding to the second sample fusion feature;
[0146] To illustrate, the sample environment-enhanced speech is the output of the aforementioned cross-modal data augmentation model during the model training phase. Before training the initial speech warning level model, the sample environment-enhanced speech is labeled with the corresponding speech warning level label.
[0147] Understandably, the speech feature encoder processes the enhanced environment speech of the sample, and extracts the speech features of the enhanced environment speech through the speech feature encoder. The processing result of the speech feature encoder is the second sample speech features.
[0148] Understandably, the processing objects of the feature fusion module are the second sample speech features and the previously obtained sample image features. The feature fusion module performs cross-model feature fusion on the second sample speech features and the previously obtained sample image features. The processing result of the feature fusion module is the second sample fused features.
[0149] Understandably, the alarm level prediction module processes the second sample fusion features. It performs environmental calculations to determine the current environment type and the required predicted voice alarm level. The module then predicts and outputs the predicted voice alarm level corresponding to the current environment type.
[0150] S406: Train the initial voice warning level model based on the predicted voice warning level and the voice warning level label until the initial voice warning level model is trained to obtain the trained voice warning level model.
[0151] Indicatively, a third model loss is determined based on the predicted voice warning level and the voice warning level label, and the model parameters of the initial voice warning level model are adjusted using the third model loss.
[0152] Furthermore, determining the third model loss based on the predicted voice warning level and the voice warning level label can be:
[0153] The predicted voice warning level and the voice warning level label are input into the third loss calculation formula to obtain the predicted classification loss, and the predicted classification loss is used as the third model loss.
[0154] The third loss calculation formula satisfies the following formula:
[0155] Lc = CrossEntropy(p,q)
[0156] Wherein, Lc is the predicted classification loss, p is the predicted voice warning level, q is the voice warning level label, and CrossEntropy is the cross-entropy operator.
[0157] Indicatively, during the model training process of the initial voice warning level model in each round, the third model loss of the initial voice warning level model is calculated based on the above method. The model parameters of the initial voice warning level model are then adjusted by backpropagation in combination with the third model loss until the model training end condition is met, and a well-trained voice warning level model can be obtained.
[0158] In one or more embodiments of this specification, by training a voice warning level model, the current environment is analyzed based on enhanced voice dimension information and image dimension to determine the voice warning method. This enables the accurate determination of which voice warning method to switch the current device environment when there is a high system risk of liveness attack. This realizes a liveness anti-attack method based on risk classification intelligent voice reminder, which can improve the effectiveness of liveness anti-attack.
[0159] The following will combine Figure 6 This manual provides a detailed description of the live attack protection device provided. It should be noted that... Figure 6 The live anti-attack device shown is used to perform the functions described in this manual. Figures 1-5 The methods of the embodiments shown are illustrated only in connection with this specification for ease of explanation. For specific technical details not disclosed, please refer to this specification. Figures 1-5 The example shown.
[0160] Please see Figure 6 This diagram illustrates the structure of the liveness detection device described in this specification. The liveness detection device 1 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the liveness detection device 1 includes a data acquisition module 11, a data processing module 12, and a liveness detection module 13, specifically used for:
[0161] Data acquisition module 11 is used to acquire the first object image and the first environmental voice of the target object in the environment.
[0162] Data processing module 12 is used to perform speech enhancement processing on the first environmental speech based on the first object image to obtain the second environmental speech;
[0163] The liveness detection module 13 is used to perform liveness detection processing on the target object based on the first object image to obtain a liveness detection result. Based on the liveness detection result, the second environmental voice and the first object image are used to determine the voice alarm level for the environment, and the voice alarm level is used to perform voice anti-attack alarm.
[0164] Optionally, the data processing module 12 is used for:
[0165] The first object image and the first environmental speech are input into a cross-modal data augmentation model for cross-modal speech augmentation processing, and the second environmental speech is output.
[0166] Optionally, the data processing module 12 is used for:
[0167] Create an initial cross-modal data augmentation model;
[0168] Acquire a first sample object image and a first sample environment speech, and add noise to the first sample environment speech to obtain a second sample environment speech;
[0169] The first sample object image and the second sample environmental speech are input into the initial cross-modal data augmentation model for training until the initial cross-modal data augmentation model completes training, resulting in the trained cross-modal data augmentation model.
[0170] Optionally, the data processing module 12 is used for:
[0171] The first sample object image and the second sample environmental speech are input into an initial cross-modal data augmentation model. The feature encoder of the initial cross-modal data augmentation model determines the first sample image features and the first sample speech features. The feature fusion module of the initial cross-modal data determines the first sample fusion features corresponding to the first sample image features and the first sample speech features. The quality enhancement module of the initial cross-modal data performs quality enhancement and noise reduction processing on the first sample environmental speech based on the first sample fusion features to obtain sample environmental enhanced speech.
[0172] The first model loss is calculated based on the enhanced speech in the sample environment and the first sample environment speech. The model parameters of the initial cross-modal data augmentation model are then adjusted based on the first model loss.
[0173] Optionally, the data processing module 12 is used for:
[0174] The speech denoising loss of the sample environment enhanced speech and the first sample environment speech is calculated using the first loss calculation formula, and the speech denoising loss is used as the first model loss.
[0175] The first loss calculation formula satisfies the following formula:
[0176] La = L2(x, x_recovery)
[0177] Wherein, La is the speech denoising loss, x is the first sample environment speech, x_recovery is the sample environment enhanced speech, and L2 represents the Euclidean distance operator.
[0178] Optionally, the liveness detection module is used for:
[0179] The first object image is input into the liveness detection model to obtain the liveness detection probability and system risk level; or, the first object image is input into the liveness detection model to obtain the liveness detection probability, system risk level and object image features.
[0180] The step of determining the voice alert level for the environment based on the liveness detection result and using the second environmental voice and the first object image includes:
[0181] If the probability of a live attack is greater than the target probability threshold and / or the system risk level is greater than the target level threshold, then the object image features corresponding to the first object image are obtained, and the voice warning level for the environment is determined by using the second environmental voice and the object image features.
[0182] Optionally, the liveness detection module is used for:
[0183] Create an initial liveness detection model;
[0184] Obtain the second sample object image and the corresponding liveness classification label, attack classification label, and risk level classification label. Input the second sample object image into the initial liveness anti-attack model. Determine the sample liveness classification result, attack classification result, risk level classification result, and sample image features through the initial liveness anti-attack model.
[0185] The initial liveness prevention model is trained based on the sample liveness classification results, the liveness classification labels, the attack class classification results, the attack classification labels, the risk level classification results, and the risk level classification labels until the initial liveness prevention model is trained, resulting in the trained liveness prevention model.
[0186] Optionally, the liveness detection module is used for:
[0187] Based on the sample liveness classification results and the liveness classification labels, the liveness classification loss is determined; based on the attack class classification results and the attack classification labels, the attack classification loss is determined; based on the risk level classification results and the risk level classification labels, the risk level prediction loss is determined.
[0188] The second model loss is determined based on the liveness classification loss, the attack classification loss, and the risk level prediction loss. The model parameters of the initial liveness anti-attack model are adjusted based on the first model loss.
[0189] Optionally, the liveness detection module is used for:
[0190] The liveness classification result and the liveness classification label are input into the second loss calculation formula to obtain the liveness classification loss. The attack classification result and the attack classification label are input into the second loss calculation formula to determine the attack classification loss. The risk level classification result and the risk level classification label are input into the second loss calculation formula to determine the risk level prediction loss.
[0191] The second loss calculation formula satisfies the following formula:
[0192] Lb = CrossEntropy(f,y)
[0193] Wherein, Lb is the classification loss symbol corresponding to the second loss calculation formula, f is the classification result, y is the classification label, and CrossEntropy is the cross-entropy operator.
[0194] Optionally, the liveness detection module is used for:
[0195] Obtain the object image features corresponding to the first object image, input the second environmental voice and the object image features into the voice alarm level model, and output the voice alarm level for the environment.
[0196] Optionally, the liveness detection module is used for:
[0197] Create an initial voice warning level model;
[0198] Obtain sample environment-enhanced speech, sample image features, and sample speech warning level labels; input the sample environment-enhanced speech and the sample image features into the initial speech warning level model to obtain the predicted speech warning level.
[0199] The initial voice warning level model is trained based on the predicted voice warning level and the voice warning level label until the initial voice warning level model is trained, resulting in a trained voice warning level model.
[0200] Optionally, the liveness detection module is used for:
[0201] The sample environment enhanced speech and the sample image features are input into the initial speech warning level model. The speech feature encoder based on the initial speech warning level model determines the second sample speech features. The feature fusion module based on the initial speech warning level model determines the second sample fusion feature corresponding to the second sample speech features and the sample image features. The warning level prediction module based on the initial speech warning level model determines the predicted speech warning level corresponding to the second sample fusion feature.
[0202] The step of training the initial voice warning level model based on the predicted voice warning level and the voice warning level label includes:
[0203] A third model loss is determined based on the predicted voice warning level and the voice warning level label, and the model parameters of the initial voice warning level model are adjusted using the third model loss.
[0204] Optionally, the liveness detection module is used for:
[0205] The predicted voice warning level and the voice warning level label are input into the third loss calculation formula to obtain the predicted classification loss, and the predicted classification loss is used as the third model loss.
[0206] The third loss calculation formula satisfies the following formula:
[0207] Lc = CrossEntropy(p,q)
[0208] Wherein, Lc is the predicted classification loss, p is the predicted voice warning level, q is the voice warning level label, and CrossEntropy is the cross-entropy operator.
[0209] It should be noted that the above embodiments of the liveness detection device, when executing the liveness detection method, are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the liveness detection device and the liveness detection method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0210] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0211] In one or more embodiments of this specification, an electronic device can acquire a first object image and a first ambient voice of a target object in its environment. Based on the first object image, it performs voice enhancement processing on the first ambient voice to obtain a second ambient voice. Based on the first object image, it performs liveness detection processing on the target object to obtain a liveness detection result. Based on the liveness detection result, it can use the second ambient voice and the first object image to determine the voice alarm level for voice attack prevention warning. This achieves a closed loop from liveness attack detection to system attack prevention, optimizing the situation in related technologies where only relevant objects are reminded to retry while ignoring system risks, thus reducing system attack risks and providing an attack prevention warning effect. Furthermore, based on the enhanced ambient voice and image, the voice alarm level that matches the current environment can be accurately determined, resulting in a better system attack prevention warning effect.
[0212] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-5 The liveness detection method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-5 The specific details of the illustrated embodiments will not be elaborated here.
[0213] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-5 The liveness detection method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-5 The specific details of the illustrated embodiments will not be elaborated here.
[0214] Please refer to Figure 7 This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this specification. The electronic device in this specification may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.
[0215] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device via various interfaces and lines, and performs various functions and processes data of electronic device 100 by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of the following: central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately through a communication chip.
[0216] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems. The data storage area may also store data created by the electronic device during use, such as phonebook data, audio and video data, chat log data, etc.
[0217] See Figure 8As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in the user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly to the specific application scenario of the third-party application.
[0218] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0219] Taking the Android operating system as an example, the programs and data stored in memory 120 are as follows: Figure 8As shown, the memory 120 can store the Linux kernel layer 320, the system runtime library layer 340, the application framework layer 360, and the application layer 380. The Linux kernel layer 320, system runtime library layer 340, and application framework layer 360 belong to the operating system space, while the application layer 380 belongs to the user space. The Linux kernel layer 320 provides low-level drivers for various hardware components of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, and power management. The system runtime library layer 340 provides support for key features of the Android system through several C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D graphics support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library, which mainly provides core libraries that allow developers to write Android applications using the Java language. The Application Framework Layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application runs in the Application Layer 380. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera apps; or third-party applications developed by third-party developers, such as games, instant messaging, and photo editing apps.
[0220] Taking the operating system as an example (iOS), the programs and data stored in memory 120 are as follows: Figure 10As shown, the iOS system includes: Core OS layer 420, Core Services layer 440, Media layer 460, and Cocoa Touch layer 480. Core OS layer 420 includes the operating system kernel, drivers, and low-level program frameworks. These low-level program frameworks provide hardware-level functionality for use by the program frameworks located in Core Services layer 440. Core Services layer 440 provides system services and / or program frameworks required by applications, such as Foundation framework, account framework, advertising framework, data storage framework, network connectivity framework, geolocation framework, motion framework, etc. Media layer 460 provides applications with audiovisual interfaces, such as interfaces related to graphics and images, audio technology, video technology, and AirPlay (wireless playback of audio and video transmission technologies). Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development and is responsible for user touch interaction on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit frameworks, map frameworks, and so on.
[0221] exist Figure 10 The framework shown includes, but is not limited to, the base framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The base framework provides many basic object classes and data types, offering the most basic system services to all applications, and is independent of the UI. The UIKit framework, on the other hand, provides a basic UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UI, thus providing the application's infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
[0222] The methods and principles for implementing data communication between third-party applications and the operating system in the iOS system can be found in the Android system, and will not be repeated here.
[0223] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined into a touch screen, which is used to receive touch operations from the user using a finger, stylus, or any suitable object on or near it, and to display the user interface of various applications. The touch screen is usually located on the front panel of the electronic device. The touch screen can be designed as a full-screen, curved screen, or irregularly shaped screen. The touch screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; this specification does not limit this.
[0224] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0225] In this specification, the entity executing each step can be the electronic device described above. Optionally, the entity executing each step can be the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.
[0226] The electronic device described in this manual may also be equipped with a display device. This display device can be any device capable of displaying information, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an e-ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). Users can use the display device on electronic device 101 to view displayed text, images, videos, and other information. The electronic device may be a smartphone, tablet computer, gaming device, AR (Augmented Reality) device, automobile, data storage device, audio playback device, video playback device, laptop, desktop computing device, or wearable device such as an electronic watch, electronic glasses, electronic helmet, electronic bracelet, electronic necklace, or electronic clothing.
[0227] exist Figure 7 In the illustrated electronic device, the processor 110 can be used to call the application program stored in the memory 120 and specifically perform the following operations:
[0228] Collect the first object image and the first environmental audio of the target object in the surrounding environment;
[0229] Based on the first object image, the first environmental speech is subjected to speech enhancement processing to obtain the second environmental speech;
[0230] Based on the first object image, the target object is subjected to liveness detection processing to obtain the liveness detection result.
[0231] Based on the liveness detection results, the second environmental voice and the first object image are used to determine the voice alarm level for the environment, and the voice alarm level is used to perform voice anti-attack alarm.
[0232] In one embodiment, the processor 110 performs the following steps when executing the speech enhancement processing on the first environmental speech based on the first object image to obtain the second environmental speech:
[0233] The first object image and the first environmental speech are input into a cross-modal data augmentation model for cross-modal speech augmentation processing, and the second environmental speech is output.
[0234] In one embodiment, the processor 110 further performs the following steps when executing the liveness detection method:
[0235] Create an initial cross-modal data augmentation model;
[0236] Acquire a first sample object image and a first sample environment speech, and add noise to the first sample environment speech to obtain a second sample environment speech;
[0237] The first sample object image and the second sample environmental speech are input into the initial cross-modal data augmentation model for training until the initial cross-modal data augmentation model completes training, resulting in the trained cross-modal data augmentation model.
[0238] In one embodiment, the processor 110 performs the following steps when training an initial cross-modal data augmentation model by inputting the first sample object image and the second sample environmental speech:
[0239] The first sample object image and the second sample environmental speech are input into an initial cross-modal data augmentation model. The feature encoder of the initial cross-modal data augmentation model determines the first sample image features and the first sample speech features. The feature fusion module of the initial cross-modal data determines the first sample fusion features corresponding to the first sample image features and the first sample speech features. The quality enhancement module of the initial cross-modal data performs quality enhancement and noise reduction processing on the first sample environmental speech based on the first sample fusion features to obtain sample environmental enhanced speech.
[0240] The first model loss is calculated based on the enhanced speech in the sample environment and the first sample environment speech. The model parameters of the initial cross-modal data augmentation model are then adjusted based on the first model loss.
[0241] In one embodiment, the processor 110 performs the following steps when calculating the first model loss based on the sample environment enhanced speech and the first sample environment speech:
[0242] The speech denoising loss of the sample environment enhanced speech and the first sample environment speech is calculated using the first loss calculation formula, and the speech denoising loss is used as the first model loss.
[0243] The first loss calculation formula satisfies the following formula:
[0244] La = L2(x, x_recovery)
[0245] Wherein, La is the speech denoising loss, x is the first sample environment speech, x_recovery is the sample environment enhanced speech, and L2 represents the Euclidean distance operator.
[0246] In one embodiment, the processor 110 performs the following steps after executing the liveness detection processing on the target object based on the first object image to obtain the liveness detection result:
[0247] The first object image is input into the liveness detection model to obtain the liveness detection probability and system risk level; or, the first object image is input into the liveness detection model to obtain the liveness detection probability, system risk level and object image features.
[0248] The step of determining the voice alert level for the environment based on the liveness detection result and using the second environmental voice and the first object image includes:
[0249] If the probability of a live attack is greater than the target probability threshold and / or the system risk level is greater than the target level threshold, then the object image features corresponding to the first object image are obtained, and the voice warning level for the environment is determined by using the second environmental voice and the object image features.
[0250] In one embodiment, the processor 110 further performs the following steps when executing the liveness detection method:
[0251] Create an initial liveness detection model;
[0252] Obtain the second sample object image and the corresponding liveness classification label, attack classification label, and risk level classification label. Input the second sample object image into the initial liveness anti-attack model. Determine the sample liveness classification result, attack classification result, risk level classification result, and sample image features through the initial liveness anti-attack model.
[0253] The initial liveness prevention model is trained based on the sample liveness classification results, the liveness classification labels, the attack class classification results, the attack classification labels, the risk level classification results, and the risk level classification labels until the initial liveness prevention model is trained, resulting in the trained liveness prevention model.
[0254] In one embodiment, the processor 110, when training the initial liveness anti-attack model based on the sample liveness classification result, the liveness classification label, the attack class classification result, the attack classification label, the risk level classification result, and the risk level classification label, performs the following steps:
[0255] Based on the sample liveness classification results and the liveness classification labels, the liveness classification loss is determined; based on the attack class classification results and the attack classification labels, the attack classification loss is determined; based on the risk level classification results and the risk level classification labels, the risk level prediction loss is determined.
[0256] The second model loss is determined based on the liveness classification loss, the attack classification loss, and the risk level prediction loss. The model parameters of the initial liveness anti-attack model are adjusted based on the first model loss.
[0257] In one embodiment, the processor 110 performs the following steps when executing the process of determining a liveness classification loss based on the sample liveness classification result and the liveness classification label, determining an attack classification loss based on the attack class classification result and the attack classification label, and determining a risk level prediction loss based on the risk level classification result and the risk level classification label:
[0258] The liveness classification result and the liveness classification label are input into the second loss calculation formula to obtain the liveness classification loss. The attack classification result and the attack classification label are input into the second loss calculation formula to determine the attack classification loss. The risk level classification result and the risk level classification label are input into the second loss calculation formula to determine the risk level prediction loss.
[0259] The second loss calculation formula satisfies the following formula:
[0260] Lb = CrossEntropy(f,y)
[0261] Wherein, Lb is the classification loss symbol corresponding to the second loss calculation formula, f is the classification result, y is the classification label, and CrossEntropy is the cross-entropy operator.
[0262] In one embodiment, the processor 110, when performing the step of determining the voice alarm level for the environment using the second environmental voice and the first object image, executes the following steps:
[0263] Obtain the object image features corresponding to the first object image, input the second environmental voice and the object image features into the voice alarm level model, and output the voice alarm level for the environment.
[0264] In one embodiment, the processor 110 further includes the following when executing the liveness detection method:
[0265] Create an initial voice warning level model;
[0266] Obtain sample environment-enhanced speech, sample image features, and sample speech warning level labels; input the sample environment-enhanced speech and the sample image features into the initial speech warning level model to obtain the predicted speech warning level.
[0267] The initial voice warning level model is trained based on the predicted voice warning level and the voice warning level label until the initial voice warning level model is trained, resulting in a trained voice warning level model.
[0268] In one embodiment, the processor 110 performs the following steps when it inputs the sample environment-enhanced speech and the sample image features into the initial speech warning level model to obtain a predicted speech warning level:
[0269] The sample environment enhanced speech and the sample image features are input into the initial speech warning level model. The speech feature encoder based on the initial speech warning level model determines the second sample speech features. The feature fusion module based on the initial speech warning level model determines the second sample fusion feature corresponding to the second sample speech features and the sample image features. The warning level prediction module based on the initial speech warning level model determines the predicted speech warning level corresponding to the second sample fusion feature.
[0270] The step of training the initial voice warning level model based on the predicted voice warning level and the voice warning level label includes:
[0271] A third model loss is determined based on the predicted voice warning level and the voice warning level label, and the model parameters of the initial voice warning level model are adjusted using the third model loss.
[0272] In one embodiment, the processor 110 performs the following steps when determining the third model loss based on the predicted voice alert level and the voice alert level label:
[0273] The predicted voice warning level and the voice warning level label are input into the third loss calculation formula to obtain the predicted classification loss, and the predicted classification loss is used as the third model loss.
[0274] The third loss calculation formula satisfies the following formula:
[0275] Lc = CrossEntropy(p,q)
[0276] Wherein, Lc is the predicted classification loss, p is the predicted voice warning level, q is the voice warning level label, and CrossEntropy is the cross-entropy operator.
[0277] Electronic devices can acquire a first object image and a first ambient voice from the surrounding environment. Based on the first object image, the first ambient voice is enhanced to obtain a second ambient voice. Based on the first object image, the target object undergoes liveness detection processing to obtain a liveness detection result. Based on the liveness detection result, the second ambient voice and the first object image can be used to determine the voice alarm level for voice attack prevention warning. This achieves a closed loop from liveness attack detection to system attack prevention, optimizing related technologies that only remind the target object to retry while ignoring system risks, thus reducing system attack risks and providing an attack prevention warning effect. Furthermore, based on the enhanced ambient voice and image, the voice alarm level appropriate to the current environment can be accurately determined, resulting in a better system attack prevention warning effect.
[0278] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0279] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the first object image and the first environmental voice involved in this specification were obtained with full authorization.
[0280] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A method for preventing attacks on live targets, the method comprising: Collect the first object image and the first environmental audio of the target object in the surrounding environment; The first object image and the first environmental speech are input into a cross-modal data augmentation model for speech enhancement processing to obtain the second environmental speech; Based on the first object image, the target object is subjected to liveness detection processing to obtain the liveness detection result. Based on the liveness detection results, the second environmental voice and the first object image are used to determine the voice alarm level for the environment, and the voice alarm level is used to perform voice anti-attack alarm. The cross-modal data augmentation model was trained using the following method: An initial cross-modal data augmentation model is created by acquiring a first sample object image and a first sample environmental speech, adding noise to the first sample environmental speech to obtain a second sample environmental speech, and inputting the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model for training until the initial cross-modal data augmentation model completes training, thereby obtaining the trained cross-modal data augmentation model. During training, the initial cross-modal data augmentation model determines first sample image features and first sample speech features based on its feature encoder. The feature fusion module of the initial cross-modal data determines first sample fusion features corresponding to the first sample image features and the first sample speech features. The quality enhancement module of the initial cross-modal data performs quality enhancement and denoising processing on the first sample environmental speech based on the first sample fusion features to obtain enhanced environmental speech. A first model loss is calculated based on the enhanced environmental speech and the first sample environmental speech. Model parameters are adjusted based on the first model loss.
2. The method according to claim 1, wherein calculating the first model loss based on the enhanced speech of the sample environment and the first sample environment speech comprises: The speech denoising loss of the sample environment enhanced speech and the first sample environment speech is calculated using the first loss calculation formula, and the speech denoising loss is used as the first model loss. The first loss calculation formula satisfies the following formula: La = L2(x, x_recovery) Wherein, La is the speech denoising loss, x is the first sample environment speech, x_recovery is the sample environment enhanced speech, and L2 represents the Euclidean distance operator.
3. The method according to claim 1, wherein performing liveness detection processing on the target object based on the first object image to obtain a liveness detection result includes: The first object image is input into the liveness detection model to obtain the liveness detection probability and system risk level. Alternatively, the first object image can be input into a liveness detection model to obtain the liveness attack probability, system risk level, and object image features. The step of determining the voice alert level for the environment based on the liveness detection result and using the second environmental voice and the first object image includes: If the probability of a live attack is greater than the target probability threshold and / or the system risk level is greater than the target level threshold, then the object image features corresponding to the first object image are obtained, and the voice warning level for the environment is determined by using the second environmental voice and the object image features.
4. The method according to claim 3, further comprising: Create an initial liveness detection model; Obtain the second sample object image and the corresponding liveness classification label, attack classification label, and risk level classification label. Input the second sample object image into the initial liveness anti-attack model. Determine the sample liveness classification result, attack classification result, risk level classification result, and sample image features through the initial liveness anti-attack model. The initial liveness prevention model is trained based on the sample liveness classification results, the liveness classification labels, the attack class classification results, the attack classification labels, the risk level classification results, and the risk level classification labels until the initial liveness prevention model is trained, resulting in the trained liveness prevention model.
5. The method according to claim 4, wherein training the initial liveness anti-attack model based on the sample liveness classification result, the liveness classification label, the attack class classification result, the attack classification label, the risk level classification result, and the risk level classification label comprises: Based on the sample liveness classification results and the liveness classification labels, the liveness classification loss is determined; based on the attack class classification results and the attack classification labels, the attack classification loss is determined; based on the risk level classification results and the risk level classification labels, the risk level prediction loss is determined. The second model loss is determined based on the liveness classification loss, the attack classification loss, and the risk level prediction loss. The model parameters of the initial liveness anti-attack model are adjusted based on the first model loss.
6. The method according to claim 5, wherein determining the liveness classification loss based on the sample liveness classification result and the liveness classification label, determining the attack classification loss based on the attack class classification result and the attack classification label, and determining the risk level prediction loss based on the risk level classification result and the risk level classification label, comprises: The liveness classification result and the liveness classification label are input into the second loss calculation formula to obtain the liveness classification loss. The attack classification result and the attack classification label are input into the second loss calculation formula to determine the attack classification loss. The risk level classification result and the risk level classification label are input into the second loss calculation formula to determine the risk level prediction loss. The second loss calculation formula satisfies the following formula: Lb = CrossEntropy(f, y) Wherein, Lb is the classification loss symbol corresponding to the second loss calculation formula, f is the classification result, y is the classification label, and CrossEntropy is the cross-entropy operator.
7. The method according to claim 1, wherein determining the voice warning level for the environment using the second environmental voice and the first object image includes: Obtain the object image features corresponding to the first object image, input the second environmental voice and the object image features into the voice alarm level model, and output the voice alarm level for the environment.
8. The method according to claim 7, further comprising: Create an initial voice warning level model; Obtain sample environment-enhanced speech, sample image features, and sample speech warning level labels; input the sample environment-enhanced speech and the sample image features into the initial speech warning level model to obtain the predicted speech warning level. The initial voice warning level model is trained based on the predicted voice warning level and the voice warning level label until the initial voice warning level model is trained, resulting in a trained voice warning level model.
9. The method according to claim 8, wherein inputting the sample environment-enhanced speech and the sample image features into the initial speech warning level model to obtain the predicted speech warning level comprises: The sample environment enhanced speech and the sample image features are input into the initial speech warning level model. The speech feature encoder based on the initial speech warning level model determines the second sample speech features. The feature fusion module based on the initial speech warning level model determines the second sample fusion feature corresponding to the second sample speech features and the sample image features. The warning level prediction module based on the initial speech warning level model determines the predicted speech warning level corresponding to the second sample fusion feature. The step of training the initial voice warning level model based on the predicted voice warning level and the voice warning level label includes: A third model loss is determined based on the predicted voice warning level and the voice warning level label, and the model parameters of the initial voice warning level model are adjusted using the third model loss.
10. The method according to claim 9, wherein determining the third model loss based on the predicted voice warning level and the voice warning level label comprises: The predicted voice warning level and the voice warning level label are input into the third loss calculation formula to obtain the predicted classification loss, and the predicted classification loss is used as the third model loss. The third loss calculation formula satisfies the following formula: Lc = CrossEntropy(p, q) Wherein, Lc is the predicted classification loss, p is the predicted voice warning level, q is the voice warning level label, and CrossEntropy is the cross-entropy operator.
11. A live anti-attack device, the device comprising: The data acquisition module is used to acquire the first object image and the first environmental voice of the target object in the environment. The data processing module is used to input the first object image and the first environmental speech into a cross-modal data enhancement model for speech enhancement processing to obtain the second environmental speech; The liveness detection module is used to perform liveness detection processing on the target object based on the first object image to obtain a liveness detection result. Based on the liveness detection result, the second environmental voice and the first object image are used to determine the voice alarm level for the environment. The voice alarm level is used to perform voice anti-attack alarm. The cross-modal data augmentation model was trained using the following method: An initial cross-modal data augmentation model is created by acquiring a first sample object image and a first sample environmental speech, adding noise to the first sample environmental speech to obtain a second sample environmental speech, and inputting the first sample object image and the second sample environmental speech into the initial cross-modal data augmentation model for training until the initial cross-modal data augmentation model completes training, thereby obtaining the trained cross-modal data augmentation model. During training, the initial cross-modal data augmentation model determines first sample image features and first sample speech features based on its feature encoder. The feature fusion module of the initial cross-modal data determines first sample fusion features corresponding to the first sample image features and the first sample speech features. The quality enhancement module of the initial cross-modal data performs quality enhancement and denoising processing on the first sample environmental speech based on the first sample fusion features to obtain enhanced environmental speech. A first model loss is calculated based on the enhanced environmental speech and the first sample environmental speech. Model parameters are adjusted based on the first model loss.
12. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 10.
13. A computer program product storing at least one instruction, said at least one instruction being loaded by a processor and executing the method steps of any one of claims 1 to 10.
14. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 10.
Citation Information
Patent Citations
Human face living body detection method, related device, equipment and storage medium
CN111310575A
Multi-modal living body detection method and device, computer equipment and storage medium
CN112633201A
Living body identification method and system
CN116503962A