Voiceprint verification method and device, electronic equipment and storage medium
By extracting noisy audio from the audio to be verified and mixing it with the registered audio to generate mixed audio, the matching degree problem of voiceprint recognition when environmental noise changes is solved, thus improving the pass rate and accuracy of voiceprint verification.
Patent Information
- Application Number
- CN202111553029.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing voiceprint recognition technologies suffer from decreased matching between verified speech and voiceprint template libraries when environmental noise changes, affecting verification pass rate and accuracy.
Noise audio is extracted from the audio to be verified and mixed with pre-registered audio to generate mixed audio. The verification result is determined using the mixed audio and the audio to be verified.
It improves the pass rate and accuracy of voiceprint verification, adapts to different noise environments, and enhances the user experience.
Smart Images

Figure CN116343800B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of voiceprint recognition, and particularly relates to a voiceprint verification method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the continuous development of voice technology, voiceprint recognition as an important biometric authentication technology is widely used in various terminal devices. At present, in the process of voiceprint recognition, the user is guided to perform voiceprint registration in a relatively quiet environment in advance to generate a high-quality voiceprint template library, and then the verification voice of the user and the voiceprint template library are compared during voiceprint verification to determine whether the voiceprint verification is passed.
[0003] However, in this way, the environmental noise of voiceprint registration cannot match the situation during actual voiceprint verification, especially when the environmental noise of the user changes greatly, which will cause the matching degree of the verification voice and the voiceprint template library to drop sharply, which will affect the pass rate of voiceprint verification and reduce the accuracy of the verification result. SUMMARY
[0004] To overcome the problems in the related art, the present disclosure provides a voiceprint verification method, device, electronic device and storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, a voiceprint verification method is provided, the method comprising:
[0006] extracting noise audio from the to-be-verified audio;
[0007] mixing the noise audio and the pre-registered registration audio to obtain mixed audio;
[0008] determining a verification result of the to-be-verified audio according to the mixed audio and the to-be-verified audio.
[0009] Optionally, the extracting noise audio from the to-be-verified audio comprises:
[0010] extracting an audio segment containing noise from the to-be-verified audio as the noise audio by using a preset noise detection algorithm.
[0011] Optionally, the registration audio is multiple, and each registration audio corresponds to a registration voiceprint feature vector; the mixing the noise audio and the pre-registered registration audio to obtain mixed audio comprises:
[0012] obtaining a verification voiceprint feature vector corresponding to the to-be-verified audio;
[0013] determine a target registered audio from the plurality of registered audios according to the verification voiceprint feature vector and the registered voiceprint feature vector;
[0014] noise the target sample audio according to the noise audio to perform audio mixing on the noise audio and the target sample audio to obtain the mixed audio.
[0015] Optionally, the determining a target registered audio from the plurality of registered audios according to the verification voiceprint feature vector and the registered voiceprint feature vector comprises:
[0016] inputting the registered voiceprint feature vector corresponding to each registered audio and the verification voiceprint feature vector into a pre-trained voiceprint matching model to obtain a matching degree of each registered voiceprint feature vector and the verification voiceprint feature vector;
[0017] taking the registered audio corresponding to the registered voiceprint feature vector with the highest matching degree as the target registered audio.
[0018] Optionally, the determining a verification result of the to-be-verified audio according to the mixed audio and the to-be-verified audio comprises:
[0019] obtaining a target voiceprint feature vector corresponding to the mixed audio;
[0020] determining the verification result of the to-be-verified audio according to the verification voiceprint feature vector and the target voiceprint feature vector.
[0021] Optionally, the determining the verification result of the to-be-verified audio according to the verification voiceprint feature vector and the target voiceprint feature vector comprises:
[0022] inputting the verification voiceprint feature vector and the target voiceprint feature vector into a pre-trained voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector;
[0023] determining that the to-be-verified audio passes voiceprint verification in a case where the verification confidence is greater than or equal to a preset threshold.
[0024] According to a second aspect of the embodiments of the present disclosure, a voiceprint verification device is provided, and the device comprises:
[0025] a extraction module configured to extract a noise audio from a to-be-verified audio;
[0026] a mixing module configured to perform audio mixing on the noise audio and a pre-registered registered audio to obtain a mixed audio;
[0027] A determining module is configured to determine a verification result of the audio to be verified according to the mixed audio and the audio to be verified.
[0028] Optionally, the extracting module is configured to extract an audio segment containing noise from the audio to be verified as the noise audio by using a preset noise detection algorithm.
[0029] Optionally, the registration audios are multiple, and each registration audio corresponds to a registration voiceprint feature vector; the mixing module comprises:
[0030] A first obtaining sub-module is configured to obtain a verification voiceprint feature vector corresponding to the audio to be verified;
[0031] A first determining sub-module is configured to determine a target registration audio from the multiple registration audios according to the verification voiceprint feature vector and the registration voiceprint feature vector;
[0032] A mixing sub-module is configured to add noise to the target sample audio according to the noise audio to perform audio mixing on the noise audio and the target sample audio, so as to obtain the mixed audio.
[0033] Optionally, the first determining sub-module is configured to:
[0034] input the registration voiceprint feature vector corresponding to each registration audio and the verification voiceprint feature vector into a pre-trained voiceprint matching model to obtain a matching degree of each registration voiceprint feature vector and the verification voiceprint feature vector;
[0035] determine the registration audio corresponding to the registration voiceprint feature vector with the highest matching degree as the target registration audio.
[0036] Optionally, the determining module comprises:
[0037] A second obtaining sub-module is configured to obtain a target voiceprint feature vector corresponding to the mixed audio;
[0038] A second determining sub-module is configured to determine a verification result of the audio to be verified according to the verification voiceprint feature vector and the target voiceprint feature vector.
[0039] Optionally, the second determining sub-module is configured to:
[0040] input the verification voiceprint feature vector and the target voiceprint feature vector into a pre-trained voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector;
[0041] In a case where the verification confidence is greater than or equal to a preset threshold, it is determined that the to-be-verified audio passes the voiceprint verification.
[0042] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising:
[0043] a processor;
[0044] a memory for storing processor-executable instructions;
[0045] The processor is configured to perform the steps of the voiceprint verification method provided in the first aspect of the present disclosure.
[0046] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer program instructions, and the program instructions are executed by a processor to implement the steps of the voiceprint verification method provided in the first aspect of the present disclosure.
[0047] The technical solutions provided by the embodiments of the present disclosure can include the following beneficial effects:
[0048] The present disclosure first extracts noise audio from the to-be-verified audio, then performs audio mixing on the noise audio and the pre-registered registered audio to obtain mixed audio, and finally determines the verification result of the to-be-verified audio according to the mixed audio and the to-be-verified audio. The present disclosure can effectively utilize the to-be-verified audio, extract noise audio from the to-be-verified audio, and perform audio mixing with the registered audio to obtain mixed audio. The noise corresponding to the mixed audio matches the noise in the audio input environment where the to-be-verified audio is located. In this way, the verification result determined by using the mixed audio and the to-be-verified audio can effectively cope with various noises in the audio input environment where the to-be-verified audio is located, thereby improving the pass rate of voiceprint verification and ensuring the accuracy of the verification result.
[0049] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0050] The accompanying drawings, which are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.
[0051] Figure 1 is a flowchart of a voiceprint verification method according to an exemplary embodiment.
[0052] Figure 2 is a flowchart of step 102 according to the embodiment shown in Figure 1
[0053] Figure 3 is a flowchart of step 102 according to the embodiment shown in Figure 1 A flowchart of one step 103 is shown in the illustrated embodiment.
[0054] Figure 4 A block diagram of a voiceprint verification device is shown in the illustrated embodiment according to an example embodiment.
[0055] Figure 5 A block diagram of a mixing module is shown in the illustrated embodiment according to an example embodiment. Figure 4
[0056] Figure 6 A block diagram of a determination module is shown in the illustrated embodiment according to an example embodiment. Figure 4
[0057] Figure 7 A block diagram of an electronic device is shown in the illustrated embodiment according to an example embodiment.
[0058] Figure 8 A block diagram of another electronic device is shown in the illustrated embodiment according to an example embodiment. DETAILED DESCRIPTION
[0059] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals represent like elements or similar elements, unless otherwise indicated. The following description of example embodiments is not representative of all embodiments consistent with the present disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0060] Before introducing the voiceprint verification method, device, electronic device and storage medium provided by the present disclosure, first introduce the application scenarios involved in each embodiment of the present disclosure. The application scenarios can be human-computer interaction, information security, monitoring and multimedia entertainment, etc. which need to verify the voiceprint of the user input audio. For example, in the human-computer interaction scenario, when the user wakes up the terminal device through voice input, the voice audio input by the user can be verified by voiceprint, so that the user who has registered the voiceprint on the terminal device in advance can pass the voiceprint verification and unlock the terminal device for interactive use. The terminal device may, for example, be a mobile terminal such as a smart phone, a tablet computer, a smart watch, a smart bracelet, a PDA (English: Personal Digital Assistant, Chinese: Personal Digital Assistant) and the like, or a fixed terminal such as a desktop computer. For another example, in the information security scenario, the voiceprint verification can be applied to the call center of the telecommunication and financial systems, and the voice audio of the user can be used as the verification code of the account to ensure the security of the user account. For another example, in the monitoring scenario, the voiceprint verification can be applied to monitor whether a specified user appears in the telephone conversation process and monitor and track the content of the conversation.
[0061] Figure 1 is a flowchart of a voiceprint verification method according to an example embodiment. As shown in Figure 1 the method can include the following steps:
[0062] In step 101, noise audio is extracted from audio to be verified.
[0063] For example, in the voiceprint recognition process, the user needs to input multiple registration audios in advance for voiceprint registration to generate a voiceprint template library. In actual situations, the audio registration environment and the audio registration state in which the user performs the voiceprint registration process are generally fixed, which limits the diversity of the voiceprint template library. Meanwhile, when the user performs voiceprint verification, it can be performed in various voiceprint verification environments (for example, open outdoor environment, closed space environment, quiet environment, and noisy environment, etc.), so that the environmental noise in the audio registration environment and the environmental noise in the voiceprint verification environment are quite different (i.e. not matched), which will affect the pass rate of voiceprint verification and reduce the accuracy of the verification result.
[0064] To avoid the above problems, the environmental noise in the voiceprint verification environment can be used to update the voiceprint template library, so that the updated voiceprint template library can effectively cope with various complex environmental noises that occur during voiceprint verification, so as to reduce the influence of the difference between the environmental noise of voiceprint registration and voiceprint verification on the pass rate of voiceprint verification, thereby improving the pass rate of voiceprint verification. Specifically, first, the user input audio to be verified for voiceprint verification can be obtained. The audio to be verified can be user speech collected by an audio collection device (for example, a microphone), or a pre-stored preset audio directly obtained from a storage device. Then, the audio to be verified can be detected to detect whether there is a noise segment in the audio to be verified, and if there is a noise segment in the audio to be verified, the noise segment is taken as noise audio.
[0065] In step 102, the noise audio and the pre-registered registration audio are mixed to obtain mixed audio.
[0066] For example, before obtaining the user input audio to be verified, multiple registration audios input by the user in advance for voiceprint registration need to be obtained. Then, according to each registration audio, a voiceprint feature vector (i.e. the voiceprint template of the user) that can represent the voiceprint feature of the user can be obtained through a pre-trained voiceprint extraction model, and a voiceprint template library can be generated according to the voiceprint feature vector. Next, the noise audio and the registration audio can be mixed to obtain mixed audio. For example, based on the noise audio, each registration audio can be voice-noised to obtain mixed audio corresponding to each registration audio.
[0067] In step 103, a verification result of the to-be-verified audio is determined according to the mixed audio and the to-be-verified audio.
[0068] In this step, the voiceprint feature vector corresponding to the mixed audio can be obtained through the pre-trained voiceprint extraction model according to the obtained mixed audio, and the voiceprint template library is updated (that is, the diversity of the registered voiceprint template library is improved) according to the voiceprint feature vector corresponding to the mixed audio, so as to replace the voiceprint feature vector corresponding to the registered audio in the original voiceprint template library with the voiceprint feature vector corresponding to the mixed audio. Further, if there is no noise segment in the detection result of the to-be-verified audio, the original voiceprint template library remains unchanged. In this way, the original voiceprint template library generated by the registered audio can be fully utilized to improve the robustness of the voiceprint template library, and the voiceprint feature vector corresponding to the mixed audio is more consistent with the voiceprint feature in the voiceprint verification environment when the user inputs the to-be-verified audio, so as to more effectively cope with complex noise environments. Even when the environmental noise in the audio registration environment is quite different from the environmental noise in the voiceprint verification environment, the false rejection rate of voiceprint verification can be effectively reduced, and the application demand of the user in different complex noise environments can be adapted to improve the user experience.
[0069] Then, the voiceprint feature vector corresponding to the to-be-verified audio can be obtained through the pre-trained voiceprint extraction model according to the to-be-verified audio. The voiceprint feature vector corresponding to the to-be-verified audio and the voiceprint feature vector corresponding to the mixed audio in the updated voiceprint template library are compared, and the verification result of the to-be-verified audio is determined according to the comparison result.
[0070] In summary, the voiceprint feature vector corresponding to the to-be-verified audio can be obtained through the pre-trained voiceprint extraction model according to the to-be-verified audio. The voiceprint feature vector corresponding to the to-be-verified audio and the voiceprint feature vector corresponding to the mixed audio in the updated voiceprint template library are compared, and the verification result of the to-be-verified audio is determined according to the comparison result.
[0071] Optionally, step 101 can be implemented in the following way:
[0072] The audio segment containing noise is extracted from the to-be-verified audio as the noise audio by using a preset noise detection algorithm.
[0073] For example, after obtaining the audio input from the user, a preset noise detection algorithm can be used to detect the audio to be verified, extracting audio segments containing noise and using these segments as the noise audio. For instance, VAD (Voice Activity Detection) endpoint detection can be performed on the audio to be verified to extract valid noise and speech segments, and the extracted noise segments can be used as the noise audio.
[0074] Figure 2 It is based on Figure 1 The illustrated embodiment shows a flowchart of step 102. Multiple audio files are registered, each corresponding to a registered voiceprint feature vector, such as... Figure 2 As shown, step 102 may include the following steps:
[0075] In step 1021, the verification voiceprint feature vector corresponding to the audio to be verified is obtained.
[0076] In one scenario, multiple registered audio files may be obtained from voiceprint registration by different users (for example, in human-computer interaction scenarios, multiple users may be using the same terminal device simultaneously, requiring pre-registration of voiceprints for each user). In this case, the registered audio file matching the audio to be verified can be determined from the multiple registered audio files, and then voiceprint verification can be performed using the registered audio file matching the audio to be verified. Specifically, after obtaining the audio to be verified, it can be input into a pre-trained voiceprint extraction model to obtain the verification voiceprint feature vector corresponding to the audio to be verified.
[0077] In step 1022, the target registered audio is determined from multiple registered audios based on the verified voiceprint feature vector and the registered voiceprint feature vector.
[0078] In step 1023, noise is added to the target sample audio based on the noise audio to mix the noise audio and the target sample audio, resulting in mixed audio.
[0079] Furthermore, the registered voiceprint feature vector and the verification voiceprint feature vector corresponding to each registered audio can be input into a pre-trained voiceprint matching model to obtain the matching degree between each registered voiceprint feature vector and the verification voiceprint feature vector. The matching degree characterizes the similarity between the registered audio corresponding to the registered voiceprint feature vector and the voiceprint of the audio to be verified; a higher matching degree indicates a greater similarity. Then, the registered audio corresponding to the registered voiceprint feature vector with the highest matching degree to the verification voiceprint feature vector (i.e., the registered audio closest to the voiceprint of the audio to be verified) can be selected as the target registered audio. Finally, noisy audio can be used as a noise source to add noise to the target sample audio, resulting in mixed audio.
[0080] Figure 3 It is based on Figure 1 The illustrated embodiment shows a flowchart of step 103. For example... Figure 3 As shown, step 103 may include the following steps:
[0081] In step 1031, the target voiceprint feature vector corresponding to the mixed audio is obtained.
[0082] For example, after obtaining the mixed audio, the mixed audio can be input into a pre-trained voiceprint extraction model to obtain the target voiceprint feature vector corresponding to the mixed audio. The voiceprint template library is then updated based on the target voiceprint feature vector to replace the voiceprint feature vector corresponding to the registered audio in the original voiceprint template library.
[0083] It should be noted that after each voiceprint verification and the acquisition of a new audio to be verified, the current voiceprint template library is updated online. This makes the voiceprint feature vector in the voiceprint template library (i.e., the voiceprint features of the registered audio) closer to the voiceprint features of the audio to be verified, thereby effectively improving the problem of the decrease in the pass rate of voiceprint verification caused by environmental noise.
[0084] In step 1032, the verification result of the audio to be verified is determined based on the verification voiceprint feature vector and the target voiceprint feature vector.
[0085] In this step, the verification result of the audio to be verified can be determined based on the verification voiceprint feature vector and the target voiceprint feature vector. For example, the voiceprint feature vector corresponding to the mixed audio and the voiceprint feature vector corresponding to the audio to be verified can be matched (or compared in likelihood) to obtain the verification score of the audio to be verified. If the verification score is greater than a set threshold, the voiceprint verification of the audio to be verified is considered to have passed; otherwise, the voiceprint verification of the audio to be verified is considered to have failed.
[0086] In one scenario, step 1032 can be implemented in the following way:
[0087] input the verification voiceprint feature vector and the target voiceprint feature vector into the pre-trained voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector.
[0088] In a case where the verification confidence is greater than or equal to a preset threshold, it is determined that the audio to be verified passes the voiceprint verification.
[0089] For example, after obtaining the target voiceprint feature vector and the verification voiceprint feature vector, the verification voiceprint feature vector and the target voiceprint feature vector can be input into the voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector. The verification confidence can be understood as a verification score of the audio to be verified, and is used to represent a matching degree of the verification voiceprint feature vector and the target voiceprint feature vector. The higher the verification confidence is, the higher the matching degree of the verification voiceprint feature vector and the target voiceprint feature vector is, that is, the voiceprint of the audio to be verified is closer to the voiceprint of the registered audio. Then, in a case where the verification confidence is greater than or equal to a preset threshold (the preset threshold is a threshold of a confidence that can be considered as the audio to be verified being an audio input by a legal user), it is determined that the audio to be verified passes the voiceprint verification.
[0090] In summary, the voiceprint verification device according to the present disclosure first extracts the noise audio from the audio to be verified, then mixes the noise audio and the pre-registered registered audio to obtain the mixed audio, and finally determines the verification result of the audio to be verified according to the mixed audio and the audio to be verified. The voiceprint verification device according to the present disclosure can effectively utilize the audio to be verified, extract the noise audio from the audio to be verified, and mix the noise audio with the registered audio to obtain the mixed audio. The noise corresponding to the mixed audio matches the noise in the audio input environment where the audio to be verified is located. Thus, the verification result determined by using the mixed audio and the audio to be verified can effectively cope with various noises in the audio input environment where the audio to be verified is located, thereby improving the pass rate of the voiceprint verification and ensuring the accuracy of the verification result.
[0091] Figure 4 is a block diagram of a voiceprint verification device according to an example embodiment. As shown in Figure 4 The device 200 includes an extraction module 201, a mixing module 202, and a determination module 203.
[0092] The extraction module 201 is configured to extract the noise audio from the audio to be verified.
[0093] The mixing module 202 is configured to mix the noise audio and the pre-registered registered audio to obtain the mixed audio.
[0094] The determination module 203 is configured to determine the verification result of the audio to be verified according to the mixed audio and the audio to be verified.
[0095] Optionally, the extraction module 201 is configured to extract an audio segment containing noise from the audio to be verified as the noise audio by using a preset noise detection algorithm.
[0096] Figure 5 is a block diagram of a mixing module according to an embodiment shown in Figure 4 As shown in Figure 5 The registration audio is multiple, and each registration audio corresponds to a registration voiceprint feature vector. The mixing module 202 includes:
[0097] The first acquisition submodule 2021 is configured to acquire a verification voiceprint feature vector corresponding to the audio to be verified.
[0098] The first determination submodule 2022 is configured to determine a target registration audio from the multiple registration audios according to the verification voiceprint feature vector and the registration voiceprint feature vector.
[0099] The mixing submodule 2023 is configured to add noise to the target sample audio according to the noise audio to perform audio mixing on the noise audio and the target sample audio, to obtain a mixed audio.
[0100] Optionally, the first determination submodule is configured to:
[0101] input the registration voiceprint feature vector corresponding to each registration audio and the verification voiceprint feature vector into a pre-trained voiceprint matching model to obtain a matching degree of each registration voiceprint feature vector and the verification voiceprint feature vector.
[0102] The registration audio corresponding to the registration voiceprint feature vector with the highest matching degree with the verification voiceprint feature vector is taken as the target registration audio.
[0103] Figure 6 is a block diagram of a determination module according to an embodiment shown in Figure 4 As shown in Figure 6 The determination module 203 includes:
[0104] The second acquisition submodule 2031 is configured to acquire a target voiceprint feature vector corresponding to the mixed audio.
[0105] The second determination submodule 2032 is configured to determine a verification result of the audio to be verified according to the verification voiceprint feature vector and the target voiceprint feature vector.
[0106] Optionally, the second determination submodule 2032 is configured to:
[0107] input the verification voiceprint feature vector and the target voiceprint feature vector into a pre-trained voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector.
[0108] In a case where the verification confidence is greater than or equal to the preset threshold, it is determined that the to-be-verified audio passes the voiceprint verification.
[0109] As to the apparatus in the above-described embodiments, the specific manners in which the respective modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.
[0110] To sum up, the voiceprint verification method provided by the present disclosure first extracts noise audio from the to-be-verified audio, then performs audio mixing on the noise audio and the pre-registered registered audio to obtain mixed audio, and finally determines the verification result of the to-be-verified audio according to the mixed audio and the to-be-verified audio. The present disclosure can effectively utilize the to-be-verified audio, extract noise audio from the to-be-verified audio, and perform audio mixing on the noise audio and the registered audio to obtain mixed audio. The noise corresponding to the mixed audio matches the noise in the audio input environment where the to-be-verified audio is located. Thus, the verification result determined by using the mixed audio and the to-be-verified audio can effectively cope with various noises in the audio input environment where the to-be-verified audio is located, thereby improving the pass rate of voiceprint verification and ensuring the accuracy of the verification result.
[0111] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of the voiceprint verification method provided by the present disclosure.
[0112] Figure 7 is a block diagram of an electronic device according to an example embodiment. The electronic device 800 can be, for example, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0113] Referring to Figure 8 , the electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0114] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of the steps of the voiceprint verification method described above. In addition, the processing component 802 can include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0115] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or nonvolatile memory, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0116] The power component 806 provides power to various components of the electronic device 800. The power component 806 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0117] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. When the electronic device 800 is in an operation mode, such as a photographing mode or a video mode, the front camera and / or the back camera can receive an external multimedia data. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0118] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.
[0119] The I / O interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, etc. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0120] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature of the electronic device 800. The sensor component 814 can include an optical sensor for detecting ambient light, a proximity sensor for detecting nearby objects without any physical touch, a CMOS or CCD image sensor for use in imaging applications, or an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor in some embodiments.
[0121] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a corresponding communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technology.
[0122] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors or other electronic elements, to perform the voiceprint verification method described above.
[0123] In an example embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to perform the voiceprint verification method described above. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0124] In another exemplary embodiment, there is also provided a computer program product comprising a computer program being executable by a programmable apparatus, the computer program having code sections for performing the voiceprint verification method described above when executed by the programmable apparatus.
[0125] Figure 8 is a block diagram of another electronic device according to an exemplary embodiment. For example, the electronic device 1900 can be provided as a server. Referring to Figure 8 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application program, executable by the processing component 1922. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the voiceprint verification method described above.
[0126] The electronic device 1900 can also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0127] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure. This application is intended to cover any variations, uses or adaptations of the disclosure that are deemed to fall within the general principles of the disclosure and include the practice of the disclosure in its broadest form without departing from the scope of the disclosure. The specification and examples are to be considered exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0128] It should be understood that the present disclosure is not limited to the precise structures described and shown in the drawings, and that various modifications and changes can be made to the embodiments without departing from the scope of the present disclosure. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. A voiceprint verification method, characterized in that, The method comprises: extracting noise audio from the audio to be verified; mixing the noise audio and the pre-registered registered audio to obtain mixed audio; determining a verification result of the audio to be verified according to the mixed audio and the audio to be verified; The registered audio is multiple, and each registered audio corresponds to a registered voiceprint feature vector; the mixing of the noise audio and the pre-registered registered audio to obtain mixed audio comprises: obtaining a verification voiceprint feature vector corresponding to the audio to be verified; determining a target registered audio from the multiple registered audios according to the verification voiceprint feature vector and the registered voiceprint feature vector; adding noise to the target registered audio according to the noise audio to mix the noise audio and the target registered audio to obtain the mixed audio.
2. The method of claim 1, wherein, The extraction of the noise audio from the audio to be verified comprises: extracting an audio segment containing noise from the audio to be verified as the noise audio by using a preset noise detection algorithm.
3. The method of claim 1, wherein, The determination of the target registered audio from the multiple registered audios according to the verification voiceprint feature vector and the registered voiceprint feature vector comprises: inputting the registered voiceprint feature vector corresponding to each registered audio and the verification voiceprint feature vector into a pre-trained voiceprint matching model to obtain a matching degree of each registered voiceprint feature vector and the verification voiceprint feature vector; the registered audio corresponding to the registered voiceprint feature vector with the highest matching degree of the verification voiceprint feature vector is taken as the target registered audio.
4. The method of claim 1, wherein, The determination of the verification result of the audio to be verified according to the mixed audio and the audio to be verified comprises: obtaining a target voiceprint feature vector corresponding to the mixed audio; determining the verification result of the audio to be verified according to the verification voiceprint feature vector and the target voiceprint feature vector.
5. The method of claim 4, wherein, The determination of the verification result of the audio to be verified according to the verification voiceprint feature vector and the target voiceprint feature vector comprises: inputting the verification voiceprint feature vector and the target voiceprint feature vector into a pre-trained voiceprint verification model to obtain a verification confidence corresponding to the verification voiceprint feature vector; in the case that the verification confidence is greater than or equal to a preset threshold, it is determined that the audio to be verified passes the voiceprint verification.
6. A voiceprint verification apparatus, characterized in that, The device comprises: an extraction module configured to extract noise audio from the audio to be verified; a mixing module configured to mix the noise audio and the pre-registered registered audio to obtain mixed audio; a determination module configured to determine a verification result of the audio to be verified according to the mixed audio and the audio to be verified; The registered audio is multiple, and each registered audio corresponds to a registered voiceprint feature vector; the mixing module comprises: a first acquisition submodule configured to obtain a verification voiceprint feature vector corresponding to the audio to be verified; a first determination submodule configured to determine a target registered audio from the multiple registered audios according to the verification voiceprint feature vector and the registered voiceprint feature vector; The mixing submodule is configured to add noise to the target registration audio according to the noise audio to perform audio mixing on the noise audio and the target registration audio, so as to obtain the mixed audio.
7. An electronic device, comprising: Comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the steps of the method of any one of claims 1-5.
8. A computer-readable storage medium having stored thereon computer program instructions, wherein, The program instructions, when executed by the processor, implement the steps of the method of any one of claims 1-5.
Citation Information
Patent Citations
Identity recognition method and device
CN110880325A