Speech enhancement method, device, system, and storage medium

By superimposing noise components onto the registered speech to match the speech environment to be verified, the problems of low voiceprint recognition rate and poor user experience under the influence of noise are solved, achieving more accurate recognition and a better user experience.

CN113921013BActive Publication Date: 2026-01-30HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010650893.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-08
Publication Date
2026-01-30
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

In noisy environments, the recognition rate of voiceprint recognition technology is affected, and existing technologies are unable to effectively improve recognition accuracy and provide a poor user experience.

Method used

By recording the registered speech in a quiet environment and superimposing noise components corresponding to the speech to be verified onto the registered speech, the registered speech is enhanced to match the noisy environment of the speech to be verified. A voiceprint recognition algorithm is then used for comparison to improve recognition accuracy.

Benefits of technology

It improves the accuracy and robustness of voiceprint recognition. Users only need to record their registration voice in a quiet environment, avoiding repeated recording in multiple scenarios and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113921013B_ABST
    Figure CN113921013B_ABST
Patent Text Reader

Abstract

This application provides a speech enhancement method, terminal device, speech enhancement system, and computer-readable storage medium based on Artificial Intelligence (AI). An electronic device acquires speech to be verified, determines at least one of environmental noise and environmental feature parameters contained in the speech, and then enhances the registered speech based on the environmental noise and / or environmental feature parameters. Finally, the electronic device compares the speech to be verified with the enhanced registered speech to determine whether the speech to be verified and the registered speech come from the same user. In this embodiment, the registered speech is enhanced according to the noise components in the speech to be verified, so that the enhanced registered speech and the speech to be verified have similar noise components, thereby obtaining more accurate recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biometric technology, and in particular to a voice enhancement method, device, system, and computer-readable storage medium. Background Technology

[0002] Currently, biometric authentication technology based on biometric identification is gradually being promoted and applied in fields such as home life and public safety. Biometric features that can be used for biometric authentication include fingerprints, facial features, irises, DNA, and voiceprints. Among them, voiceprint recognition technology (also known as speaker recognition technology) uses voiceprints as the identification feature to collect voice samples in a non-contact manner. The collection method is more discreet and therefore more easily accepted by users.

[0003] In existing technologies, the voiceprint recognition rate is affected when there is noise in the environment where the sound sample is collected. Summary of the Invention

[0004] Some embodiments of this application provide a speech enhancement method, a terminal device, a speech enhancement system, and a computer-readable storage medium. The following describes this application from multiple aspects, and the embodiments and beneficial effects of the following aspects can be referred to each other.

[0005] In a first aspect, embodiments of this application provide a voice enhancement method applied to an electronic device, comprising: acquiring voice to be verified; determining environmental noise and / or environmental feature parameters contained in the voice to be verified; enhancing the registered voice based on the environmental noise and / or environmental feature parameters; comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user.

[0006] According to the embodiments of this application, the registered speech is enhanced based on the noise components in the speech to be verified, so that the enhanced registered speech and the speech to be verified have similar noise components. Thus, the main difference between the speech to be verified and the enhanced registered speech lies in the difference between their effective speech components. By comparing the two using a voiceprint recognition algorithm, a more accurate recognition result can be obtained. Furthermore, in this embodiment, the user only needs to record the registered speech in a quiet environment, eliminating the need to record the registered speech in multiple scenarios, thus providing a better user experience.

[0007] In some implementations, the registered speech is the speech from the registered speaker, collected in a quiet environment. This way, the registered speech contains no significant noise components, which can improve recognition accuracy.

[0008] In some implementations, enhancing the registered speech based on environmental noise includes superimposing environmental noise onto the registered speech. The method implemented in this application obtains enhanced registered speech by superimposing environmental noise onto the registered speech; the algorithm is simple.

[0009] In some embodiments, the ambient noise is the sound picked up by a secondary microphone of the electronic device. Embodiments of this application can conveniently determine the noise contained in the speech to be verified.

[0010] In some implementations, the duration of the voice message to be verified is shorter than the duration of the voice message to be registered. This allows users to record shorter voice messages to be verified, which improves the user experience.

[0011] In some implementations, the environmental feature parameters include the scene type corresponding to the speech to be verified; enhancing the registered speech based on the environmental feature parameters includes: determining the template noise corresponding to the scene type based on the scene type corresponding to the speech to be verified, and superimposing the template noise on the registered speech.

[0012] According to the embodiments of this application, the registered speech is enhanced by superimposing template noise on the registered speech, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0013] In some implementations, the scene type corresponding to the speech to be verified is determined by recognizing the speech using a scene recognition algorithm. In some implementations, the scene recognition algorithm is any one of the following: GMM algorithm; DNN algorithm.

[0014] In some implementations, the scenario type for the voice to be verified is any of the following: home scenario; in-vehicle scenario; noisy outdoor scenario; meeting room scenario; cinema scenario. The scenario types in the implementations of this application cover places where users engage in daily activities, which is beneficial to improving user experience.

[0015] In some implementations, the environmental parameter features of the voice to be verified include the distance between the user generating the voice and the electronic device; enhancing the registered voice based on the environmental feature parameters includes: performing far-field simulation on the registered voice according to the distance between the user generating the voice and the electronic device. Specifically, performing far-field simulation on the registered voice is used to simulate the acquisition distance of the registered voice (the distance between the voice acquisition device for the registered voice and the user generating the registered voice) to the acquisition distance of the voice to be verified (the distance between the voice acquisition device for the voice to be verified and the user generating the voice).

[0016] According to the embodiments of this application, by performing far-field simulation on the registered speech, the attenuation component of the speech to be verified during the propagation process can be taken into account, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0017] In some implementations, far-field simulation of the registered speech is performed based on the distance between the user who generated the speech to be verified and the electronic device, including: establishing an impulse response function of the acquisition location of the speech to be verified based on the mirror source model method according to the distance between the user who generated the speech to be verified and the electronic device; and convolving the impulse response function with the audio signal of the registered speech to perform far-field simulation of the registered speech.

[0018] In some implementations, the speech to be verified and the enhanced registered speech are both processed by the same front-end processing algorithm. Front-end processing can remove interference factors in the speech, which helps to improve the accuracy of voiceprint recognition.

[0019] In some implementations, the front-end processing algorithm includes at least one of the following processing algorithms: echo cancellation; dereverberation; active noise reduction; dynamic gain; directional sound pickup.

[0020] In some implementations, there are multiple registered voice messages; and, based on environmental noise and / or environmental characteristic parameters, each of the multiple registered voice messages is enhanced to obtain multiple enhanced registered voice messages.

[0021] According to the implementation method of this application, multiple enhanced registered voices are obtained. The voice to be verified can be matched with the multiple enhanced registered voices separately to obtain multiple similarity matching results. Then, the similarity between the voice of the speaker to be verified and the voice of the registered speaker can be judged comprehensively based on the multiple similarity matching results. This allows the error of a single matching result to be averaged, which is beneficial to improving the accuracy of voiceprint recognition and the robustness of the voiceprint recognition algorithm.

[0022] In some implementations, comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user includes: extracting feature parameters of the voice to be verified and feature parameters of the enhanced registered voice using a feature parameter extraction algorithm; performing parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice using a parameter recognition model to obtain the voice template of the speaker to be verified and the voice template of the registered speaker, respectively; and matching the voice template of the speaker to be verified and the voice template of the registered speaker using a template matching algorithm, and determining that the voice to be verified and the registered voice come from the same user based on the matching result.

[0023] In some implementations, the feature parameter extraction algorithm is the MFCC algorithm, the log mel algorithm, or the LPCC algorithm; and / or, the parameter recognition model is the identity vector model, the time-delay neural network model, or the ResNet model; and / or, the template matching algorithm is the cosine distance method, the linear discriminant method, or the probabilistic linear discriminant analysis method.

[0024] Secondly, embodiments of this application provide a voice enhancement method, comprising: a terminal device collecting voice to be verified and sending the voice to be verified to a server communicatively connected to the terminal device; the server determining environmental noise and / or environmental feature parameters contained in the voice to be verified; the server enhancing the registered voice based on the environmental noise and / or environmental feature parameters; the server comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user; and the server sending the determination result that the voice to be verified and the registered voice come from the same user to the terminal device.

[0025] According to the embodiments of this application, the registered speech is enhanced based on the noise components in the speech to be verified, so that the enhanced registered speech and the speech to be verified have similar noise components. Thus, the main difference between the speech to be verified and the enhanced registered speech lies in the difference between their effective speech components. By comparing the two using a speaker recognition algorithm, a more accurate recognition result can be obtained. Furthermore, in this embodiment, the user only needs to record the registered speech in a quiet environment, eliminating the need to record the registered speech in multiple scenarios, resulting in a better user experience. In this embodiment, the speaker recognition algorithm is implemented on a server, saving local computing resources on the terminal device.

[0026] In some implementations, the registered speech is the speech from the registered speaker, collected in a quiet environment. This way, the registered speech contains no significant noise components, which can improve recognition accuracy.

[0027] In some implementations, enhancing the registered speech based on environmental noise includes superimposing environmental noise onto the registered speech. The method implemented in this application obtains enhanced registered speech by superimposing environmental noise onto the registered speech; the algorithm is simple.

[0028] In some embodiments, the ambient noise is the sound picked up by the secondary microphone of the terminal device. Embodiments of this application can conveniently determine the noise contained in the speech to be verified.

[0029] In some implementations, the duration of the voice message to be verified is shorter than the duration of the voice message to be registered. This allows users to record shorter voice messages to be verified, which improves the user experience.

[0030] In some implementations, the environmental feature parameters include the scene type corresponding to the speech to be verified; enhancing the registered speech based on the environmental feature parameters includes: determining the template noise corresponding to the scene type based on the scene type corresponding to the speech to be verified, and superimposing the template noise on the registered speech.

[0031] According to the embodiments of this application, the registered speech is enhanced by superimposing template noise on the registered speech, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0032] In some implementations, the scene type corresponding to the speech to be verified is determined by recognizing the speech using a scene recognition algorithm. In some implementations, the scene recognition algorithm is any one of the following: GMM algorithm; DNN algorithm.

[0033] In some implementations, the scenario type for the voice to be verified is any of the following: home scenario; in-vehicle scenario; noisy outdoor scenario; meeting room scenario; cinema scenario. The scenario types in the implementations of this application cover places where users engage in daily activities, which is beneficial to improving user experience.

[0034] In some implementations, the environmental parameter features of the voice to be verified include the distance between the user generating the voice and the terminal device; enhancing the registered voice based on the environmental feature parameters includes: performing far-field simulation on the registered voice according to the distance between the user generating the voice and the terminal device. Specifically, performing far-field simulation on the registered voice is used to simulate the acquisition distance of the registered voice (the distance between the voice acquisition device for the registered voice and the user generating the registered voice) to the acquisition distance of the voice to be verified (the distance between the voice acquisition device for the voice to be verified and the user generating the voice).

[0035] According to the embodiments of this application, by performing far-field simulation on the registered speech, the attenuation component of the speech to be verified during the propagation process can be taken into account, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0036] In some implementations, far-field simulation of the registered speech is performed based on the distance between the user who generated the speech to be verified and the terminal device, including: establishing an impulse response function of the acquisition location of the speech to be verified based on the mirror source model method according to the distance between the user who generated the speech to be verified and the terminal device; and convolving the impulse response function with the audio signal of the registered speech to perform far-field simulation of the registered speech.

[0037] In some implementations, the speech to be verified and the enhanced registered speech are both processed by the same front-end processing algorithm. Front-end processing can remove interference factors in the speech, which helps to improve the accuracy of voiceprint recognition.

[0038] In some implementations, the front-end processing algorithm includes at least one of the following processing algorithms: echo cancellation; dereverberation; active noise reduction; dynamic gain; directional sound pickup.

[0039] In some implementations, there are multiple registered voice messages; and the server enhances each of the multiple registered voice messages based on environmental noise and / or environmental characteristic parameters to obtain multiple enhanced registered voice messages.

[0040] According to the implementation method of this application, multiple enhanced registered voices are obtained. The voice to be verified can be matched with the multiple enhanced registered voices separately to obtain multiple similarity matching results. Then, the similarity between the voice of the speaker to be verified and the voice of the registered speaker can be judged comprehensively based on the multiple similarity matching results. This allows the error of a single matching result to be averaged, which is beneficial to improving the accuracy of voiceprint recognition and the robustness of the voiceprint recognition algorithm.

[0041] In some implementations, comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user includes: extracting feature parameters of the voice to be verified and feature parameters of the enhanced registered voice using a feature parameter extraction algorithm; performing parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice using a parameter recognition model to obtain the voice template of the speaker to be verified and the voice template of the registered speaker, respectively; and matching the voice template of the speaker to be verified and the voice template of the registered speaker using a template matching algorithm, and determining that the voice to be verified and the registered voice come from the same user based on the matching result.

[0042] In some implementations, the feature parameter extraction algorithm is the MFCC algorithm, the log mel algorithm, or the LPCC algorithm; and / or, the parameter recognition model is the identity vector model, the time-delay neural network model, or the ResNet model; and / or, the template matching algorithm is the cosine distance method, the linear discriminant method, or the probabilistic linear discriminant analysis method.

[0043] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing instructions executable by one or more processors of the electronic device; and a processor, which, when executing the instructions in the memory, causes the electronic device to perform the speaker recognition method provided in any embodiment of the first aspect of this application. The beneficial effects achievable in this third aspect can be referred to in conjunction with the beneficial effects of the method provided in any embodiment of the first aspect, and will not be repeated here.

[0044] Fourthly, embodiments of this application provide a voice enhancement system, including a terminal device and a server communicatively connected to the terminal device, wherein...

[0045] The terminal device collects the voice to be verified and sends it to the server. The server is used to determine the environmental noise and / or environmental feature parameters contained in the voice to be verified, and enhance the registered voice based on the environmental noise and / or environmental feature parameters. The server also compares the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user. The server is also used to send the determination result that the voice to be verified and the registered voice come from the same user to the terminal device.

[0046] According to the embodiments of this application, the registered speech is enhanced based on the noise components in the speech to be verified, so that the enhanced registered speech and the speech to be verified have similar noise components. Thus, the main difference between the speech to be verified and the enhanced registered speech lies in the difference between their effective speech components. By comparing the two using a speaker recognition algorithm, a more accurate recognition result can be obtained. Furthermore, in this embodiment, the user only needs to record the registered speech in a quiet environment, eliminating the need to record the registered speech in multiple scenarios, resulting in a better user experience. In this embodiment, the speaker recognition algorithm is implemented on a server, saving local computing resources on the terminal device.

[0047] In some implementations, the registered speech is the speech from the registered speaker, collected in a quiet environment. This way, the registered speech contains no significant noise components, which can improve recognition accuracy.

[0048] In some implementations, enhancing the registered speech based on environmental noise includes superimposing environmental noise onto the registered speech. The method implemented in this application obtains enhanced registered speech by superimposing environmental noise onto the registered speech; the algorithm is simple.

[0049] In some embodiments, the ambient noise is the sound picked up by the secondary microphone of the terminal device. Embodiments of this application can conveniently determine the noise contained in the speech to be verified.

[0050] In some implementations, the duration of the voice message to be verified is shorter than the duration of the voice message to be registered. This allows users to record shorter voice messages to be verified, which improves the user experience.

[0051] In some implementations, the environmental feature parameters include the scene type corresponding to the speech to be verified; enhancing the registered speech based on the environmental feature parameters includes: determining the template noise corresponding to the scene type based on the scene type corresponding to the speech to be verified, and superimposing the template noise on the registered speech.

[0052] According to the embodiments of this application, the registered speech is enhanced by superimposing template noise on the registered speech, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0053] In some implementations, the scene type corresponding to the speech to be verified is determined by recognizing the speech using a scene recognition algorithm. In some implementations, the scene recognition algorithm is any one of the following: GMM algorithm; DNN algorithm.

[0054] In some implementations, the scenario type for the voice to be verified is any of the following: home scenario; in-vehicle scenario; noisy outdoor scenario; meeting room scenario; cinema scenario. The scenario types in the implementations of this application cover places where users engage in daily activities, which is beneficial to improving user experience.

[0055] In some implementations, the environmental parameter features of the voice to be verified include the distance between the user generating the voice and the terminal device; enhancing the registered voice based on the environmental feature parameters includes: performing far-field simulation on the registered voice according to the distance between the user generating the voice and the terminal device. Specifically, performing far-field simulation on the registered voice is used to simulate the acquisition distance of the registered voice (the distance between the voice acquisition device for the registered voice and the user generating the registered voice) to the acquisition distance of the voice to be verified (the distance between the voice acquisition device for the voice to be verified and the user generating the voice).

[0056] According to the embodiments of this application, by performing far-field simulation on the registered speech, the attenuation component of the speech to be verified during the propagation process can be taken into account, so that the enhanced registered speech and the speech to be verified have noise components that are as close as possible, which is beneficial to improving the recognition accuracy.

[0057] In some implementations, far-field simulation of the registered speech is performed based on the distance between the user who generated the speech to be verified and the terminal device, including: establishing an impulse response function of the acquisition location of the speech to be verified based on the mirror source model method according to the distance between the user who generated the speech to be verified and the terminal device; and convolving the impulse response function with the audio signal of the registered speech to perform far-field simulation of the registered speech.

[0058] In some implementations, the speech to be verified and the enhanced registered speech are both processed by the same front-end processing algorithm. Front-end processing can remove interference factors in the speech, which helps to improve the accuracy of voiceprint recognition.

[0059] In some implementations, the front-end processing algorithm includes at least one of the following processing algorithms: echo cancellation; dereverberation; active noise reduction; dynamic gain; directional sound pickup.

[0060] In some implementations, there are multiple registered voice messages; and the server enhances each of the multiple registered voice messages based on environmental noise and / or environmental characteristic parameters to obtain multiple enhanced registered voice messages.

[0061] According to the implementation method of this application, multiple enhanced registered voices are obtained. The voice to be verified can be matched with the multiple enhanced registered voices separately to obtain multiple similarity matching results. Then, the similarity between the voice of the speaker to be verified and the voice of the registered speaker can be judged comprehensively based on the multiple similarity matching results. This allows the error of a single matching result to be averaged, which is beneficial to improving the accuracy of voiceprint recognition and the robustness of the voiceprint recognition algorithm.

[0062] In some implementations, comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user includes: extracting feature parameters of the voice to be verified and feature parameters of the enhanced registered voice using a feature parameter extraction algorithm; performing parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice using a parameter recognition model to obtain the voice template of the speaker to be verified and the voice template of the registered speaker, respectively; and matching the voice template of the speaker to be verified and the voice template of the registered speaker using a template matching algorithm, and determining that the voice to be verified and the registered voice come from the same user based on the matching result.

[0063] In some implementations, the feature parameter extraction algorithm is the MFCC algorithm, the log mel algorithm, or the LPCC algorithm; and / or, the parameter recognition model is the identity vector model, the time-delay neural network model, or the ResNet model; and / or, the template matching algorithm is the cosine distance method, the linear discriminant method, or the probabilistic linear discriminant analysis method.

[0064] Fifthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method provided in any embodiment of the first aspect of this application, or cause the computer to perform the method provided in any embodiment of the second aspect of this application. The beneficial effects achievable through the fifth aspect can be found in the beneficial effects of the methods provided in any embodiment of the first or second aspect, and will not be repeated here. Attached Figure Description

[0065] Figure 1a Exemplary application scenarios of the speech enhancement method provided in the embodiments of this application are illustrated;

[0066] Figure 1b This paper illustrates another exemplary application scenario of the speech enhancement method provided in the embodiments of this application;

[0067] Figure 2 A schematic diagram of the structure of the speech enhancement device provided in the embodiments of this application is shown;

[0068] Figure 3 A flowchart of a speech enhancement method provided in one embodiment of this application is shown;

[0069] Figure 4 A flowchart of a speech enhancement method provided in another embodiment of this application is shown;

[0070] Figure 5 This illustrates an application scenario of the speech enhancement method provided in an embodiment of this application;

[0071] Figure 6 A structural diagram of the electronic device provided in an embodiment of this application is shown;

[0072] Figure 7 A block diagram of a system-on-a-chip (SoC) provided in an embodiment of this application is shown. Detailed Implementation

[0073] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0074] Speaker recognition technology (also known as voiceprint recognition technology) is a technique that uses the uniqueness of a speaker's voiceprint to identify the speaker's identity. Because each person's vocal organs (such as tongue, teeth, larynx, lungs, nasal cavity, vocal tract, etc.) have innate differences, and vocal habits also have acquired differences, each person's voiceprint characteristics are unique. By analyzing voiceprint characteristics, the speaker's identity can be identified.

[0075] The specific process of speaker identification involves collecting the voice of the speaker whose identity is to be verified and comparing it with the voice of a specific speaker to confirm whether the speaker whose identity is to be verified is indeed that specific speaker. In this paper, the voice of the speaker whose identity is to be verified is referred to as "voice to be verified," and the speaker whose identity is to be verified is referred to as "speaker to be verified"; the voice of the specific speaker is referred to as "voice to be registered," and the specific speaker is referred to as "registered speaker."

[0076] refer to Figure 1a Taking the voiceprint unlock function of a mobile phone (i.e., unlocking the phone screen through voiceprint recognition) as an example, the above process will be introduced. Before using the voiceprint unlock function, the phone owner records their own voice (the registration voice) through the phone's microphone.

[0077] When unlocking a phone screen via voiceprint recognition is required, the current user records real-time voice input through the phone's microphone (this voice is to be verified). The phone's built-in voiceprint recognition program compares this voice input with the registered voice input to determine if the current user is the phone's owner. If the voice input matches the registered voice input, the current user is identified as the owner, authentication is successful, and the phone unlocks the screen. If the voice input does not match the registered voice input, the current user is not identified as the owner, authentication is unsuccessful, and the phone may refuse to unlock the screen.

[0078] The above example of using a mobile phone's voiceprint unlock function illustrates the application of voiceprint recognition technology. However, this application is not limited to this; voiceprint recognition technology can be applied to other scenarios requiring speaker identification. For example, voiceprint recognition technology can be applied in the home environment for voice control of smartphones, smart cars, and smart homes (e.g., smart audio-visual devices, smart lighting systems, smart door locks); it can also be applied in the payment field, combining voiceprint authentication with other authentication methods (e.g., passwords, dynamic verification codes) to perform dual or multiple authentications of the user's identity, thereby improving payment security; it can also be applied in the field of information security, using voiceprint authentication as a method of logging into accounts; and it can also be applied in the judicial field, using voiceprints as auxiliary evidence for identity verification.

[0079] Furthermore, the primary device for voiceprint recognition can be any electronic device other than a mobile phone, such as mobile devices including wearable devices (e.g., wristbands, headphones), in-vehicle terminals, etc.; or fixed devices, including smart home devices, network servers, etc. In addition, voiceprint recognition algorithms can be implemented not only on the terminal but also in the cloud. For example, after a mobile phone collects the voice to be verified, it can send the collected voice to the cloud, where a voiceprint recognition algorithm identifies the voice. After the cloud completes the recognition, the result is returned to the mobile phone. Through cloud-based recognition, users can share cloud computing resources, thus saving local computing resources on the mobile phone.

[0080] like Figure 1b In the scenario shown, when collecting the voice of the speaker to be verified, if there is ambient noise, this noise will be picked up by the microphone and become part of the voice to be verified. Thus, the voice to be verified includes not only the speaker's voice but also noise components, which reduces the voiceprint recognition rate.

[0081] This embodiment does not limit the scenario for voiceprint recognition. For example, it can also be a home scenario, a vehicle scenario, a meeting scenario, a cinema scenario, etc.

[0082] When a phone owner needs to unlock their phone via voiceprint recognition, if there is noise in the surrounding environment, the phone's microphone will pick up both the owner's voice and the ambient noise. This can cause the phone to compare the real-time voice with the owner's pre-installed registration voice, potentially resulting in a mismatch. Even if the current user is the owner, the phone may still show a "user authentication failed" message, negatively impacting the user experience.

[0083] In existing technologies, some solutions remove noise components from the speech to be verified by denoising, thereby improving the voiceprint recognition rate. However, the denoised speech still contains some noise components, and some effective speech components (the speaker's speech components) are also removed. As a result, the denoised speech may still not be correctly recognized, and the improvement in voiceprint recognition rate is not significant.

[0084] In existing technologies, another solution improves voiceprint recognition rates by recording registration voice messages in different scenarios. Specifically, users record registration voice messages in multiple different scenarios (e.g., home, movie theater, noisy outdoor environment). During voiceprint recognition, the voice message to be verified is compared with the registration voice messages recorded in the corresponding scenarios to improve the recognition rate. However, this existing technology requires users to record registration voice messages in multiple different scenarios, resulting in a lower user experience.

[0085] To address this, this application provides a speech enhancement method to improve the recognition rate and robustness of voiceprint recognition methods, thereby enhancing the user experience. In this application, after acquiring the speech to be verified, noise components corresponding to the noise components in the speech to be verified are superimposed onto the registered speech. Then, the registered speech with the added noise components is compared with the speech to be verified to obtain the recognition result. In other words, this application enhances the registered speech based on the noise components in the speech to be verified, so that the enhanced registered speech and the speech to be verified have similar noise components. Thus, the main difference between the speech to be verified and the enhanced registered speech lies in the difference between their effective speech components. By comparing the two using a voiceprint recognition algorithm, a more accurate recognition result can be obtained. Furthermore, in this application's implementation, users only need to record the registration speech in a quiet environment, eliminating the need to record registration speech in multiple scenarios, thus providing a better user experience.

[0086] Here, "effective speech components" refers to speech components from the speaker. For example, the effective speech components in the speech to be verified are the speech components of the speaker to be verified, and the effective speech components in the enhanced registered speech are the speech components of the registered speaker.

[0087] The following still combines Figure 1b The technical solution of this application is introduced by the voiceprint unlocking function of the mobile phone, but it is understood that this application is not limited thereto.

[0088] Figure 2 The structure of mobile phone 100 is shown. Mobile phone 100 may include processor 110, external memory interface 120, internal memory 121, antenna, communication module 150, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, camera 193, display screen 194, etc.

[0089] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0090] Processor 110 may include one or more processing units, such as application processor (AP), modem processor, controller, digital signal processor (DSP), baseband processor, etc. Different processing units may be independent devices or integrated into one or more processors.

[0091] The processor can generate operation control signals based on the instruction opcode and timing signals to control the instruction fetching and execution.

[0092] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0093] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, and / or a general-purpose input / output (GPIO) interface, etc.

[0094] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals.

[0095] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, an audio module 170, etc. The GPIO interface can also be configured as an I2S interface, etc.

[0096] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the mobile phone 100. In other embodiments of this application, the mobile phone 100 may also adopt different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0097] The wireless communication function of mobile phone 100 can be realized through antenna, communication module 150, modem processor and baseband processor.

[0098] Antennas are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antennas can be reused as diversity antennas for a wireless local area network. In some other embodiments, antennas can be used in conjunction with tuning switches.

[0099] The communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the mobile phone 100. The communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The communication module 150 can receive electromagnetic waves via an antenna, filter and amplify the received electromagnetic waves, and transmit them to a modem processor for demodulation. The communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0100] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and housed within the same device as the communication module 150 or other functional modules.

[0101] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the mobile phone 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0102] The internal memory 121 can be used to store computer executable program code, which includes instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), a voiceprint recognition program, a voice signal front-end processing program, etc. The data storage area may store data created during the use of the mobile phone 100 (such as audio data, phonebook, etc.), and data required for voiceprint recognition, such as audio data of registered voices, trained voice parameter recognition models, etc. Furthermore, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the mobile phone 100 by running instructions stored in the internal memory 121 and / or instructions stored in memory located in the processor.

[0103] The mobile phone 100 can achieve audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0104] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0105] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Mobile phone 100 can listen to music or make hands-free calls through the speaker 170A.

[0106] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the mobile phone 100 answers a call or voice message, the receiver 170B can be brought close to the user's ear to listen to the voice.

[0107] Microphone 170C, also known as a "mic," "microphone," or "voice transducer," is used to convert sound signals into electrical signals. When recording registration or verification voice, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Mobile phone 100 can be configured with at least one microphone 170C.

[0108] In other embodiments, the mobile phone 100 may be equipped with two microphones 170C, which, in addition to collecting sound signals, can also achieve noise reduction. Specifically, the mobile phone 100 has one microphone on the bottom and one on the top. One microphone 170C is located on the bottom side of the mobile phone 100, and the other microphone 170C is located on the top side of the mobile phone 100. When a user makes a call or sends a voice message, their mouth is usually close to the bottom side microphone 170C. Therefore, the user's voice will generate a larger audio signal Va in this microphone, which is referred to herein as the "main mic". At the same time, the user's voice will also generate a certain amount of audio signal Vb in the top side microphone 170C. However, since this microphone is farther from the user's mouth, the audio signal Vb in this microphone is significantly smaller than the audio signal Va in the main mic, which is referred to herein as the "secondary mic".

[0109] Regarding environmental noise, since the noise source is usually 100 meters away from the mobile phone, it can be assumed that the distance between the noise source and the main mic and the secondary mic is basically the same. That is, it can be assumed that the intensity of the noise collected by the main mic and the secondary mic is basically the same.

[0110] The difference in signal strength caused by the positional difference between two microphones can be used to separate noise signals from user speech signals. For example, by subtracting the signal from the secondary microphone from the signal picked up by the primary microphone, the user's speech signal can be obtained (this is the principle of dual-mic active noise cancellation). Furthermore, by removing the user's speech signal from the primary microphone signal, the noise signal can be separated. Alternatively, since the audio signal Vb on the secondary microphone is significantly smaller than the audio signal Va on the primary microphone, the signal picked up by the secondary microphone can be considered noise.

[0111] The above provides one way to set up the dual microphones on the phone 100. However, this is only an example. Other microphone setups are possible, such as placing the main microphone on the front of the phone 100 and the secondary microphone on the back.

[0112] In other embodiments, the mobile phone 100 may also be equipped with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, and achieve directional recording functions, etc.

[0113] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a Universal Serial Bus (USB) interface, or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, or a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0114] Example 1

[0115] The following combination Figure 1b This embodiment describes the technical solution in the context of mobile phone voiceprint unlocking. It is understood that this application is not limited to this, and the voice enhancement method of this application can also be applied to… Figure 1b Other scenarios besides the one shown.

[0116] refer to Figure 3 This embodiment provides a speech enhancement method. After acquiring the speech to be verified, noise contained in the speech to be verified is separated from it. Then, the separated noise is superimposed on the registered speech. In this way, the speech to be verified and the registered speech with added noise have similar noise components. The main difference between the two lies in the difference between their effective speech components, thereby improving the voiceprint recognition rate and the robustness of the voiceprint recognition method. Specifically, the speech enhancement method provided in this embodiment includes the following steps:

[0117] S110: Collect registration voice. To provide voiceprint unlocking functionality, mobile phone 100 has a voiceprint unlocking application (which can be a system application or a third-party application). To utilize the voiceprint unlocking function of mobile phone 100, when the owner of mobile phone 100 registers a user account for the voiceprint unlocking application, their voice is collected through mobile phone 100. The voiceprint unlocking application uses this voice as the reference voice for subsequent voiceprint recognition; this voice is the registration voice. However, this application is not limited to this. For example, in other embodiments, when mobile phone 100 is powered on for the first time, the owner enters the registration voice through the setup wizard of mobile phone 100, and the voiceprint unlocking application of mobile phone 100 uses this voice as the reference voice for voiceprint recognition.

[0118] Here, the registration voice is a voice recorded by the owner of the mobile phone number 100 in a quiet environment, so there is no obvious noise component in the registration voice.

[0119] The signal-to-noise ratio (SNR) in the registered voice recording environment (i.e., the ratio of the intensity of the host's voice signal to the intensity of the noise signal) is used to characterize the recording environment. When the SNR in the recording environment is higher than a set value (e.g., 30 dB), the recording environment is considered quiet. Alternatively, when the intensity of the noise signal in the registered voice recording environment is lower than a set value (e.g., 20 dB), the recording environment is considered quiet.

[0120] In this embodiment, the registration voice from the phone owner is captured through the microphone of the mobile phone 100. This registration voice is near-field voice. When recording the registration voice, the distance between the phone owner's mouth and the main microphone of the mobile phone 100 is kept within 30cm to 1m. For example, the phone owner holds the mobile phone 100 and speaks directly into the main microphone, keeping the distance between the phone owner's mouth and the main microphone of the mobile phone 100 within 30cm. This can avoid the attenuation of the phone owner's voice due to the long transmission distance.

[0121] The user records six voice messages during the registration process to form six registration voice messages. Recording multiple voice messages helps improve the flexibility of speech recognition and the richness of voiceprint information.

[0122] To ensure a good user experience and sufficient voiceprint information in each registration voice message, the length of each voice message is 10–30 seconds. Furthermore, each voice message corresponds to different text content to enrich the voiceprint information contained within. After acquiring the registration voice message, the mobile phone 100 stores the audio signal in its internal memory. However, this application is not limited to this; the mobile phone 100 can also upload the audio signal of the registration voice message to the cloud for voiceprint recognition via a cloud-based recognition mode.

[0123] The above-described recording methods, recording lengths, and quantities of registration voice messages are merely illustrative examples, and this application is not limited thereto. For instance, in other examples, registration voice messages can be recorded using other recording devices (e.g., voice recorders, dedicated microphones, etc.), the number of registration voice messages can be one, and the length of the registration voice message can be greater than 30 seconds.

[0124] For the sake of narrative coherence, step S110 is mentioned first. It is understood that step S110, as the data preparation process of the speech enhancement method, is relatively independent of a single speech enhancement process and does not need to occur together with other steps of the speech enhancement method each time.

[0125] S120: Collects the voice to be verified. The voice to be verified is the voice recorded by the current user of the phone in a noisy environment. In other words, the phone user can unlock the phone screen in this scenario using voiceprint recognition. Furthermore, the current user of the phone is the person currently operating the phone (100), which could be the phone owner or someone else.

[0126] In this embodiment, the voice to be verified is collected through the microphone of the mobile phone 100. When the screen of the mobile phone 100 is locked, the microphone of the mobile phone 100 is turned on. At this time, the current user of the mobile phone 100 can record the voice to be verified through the microphone of the mobile phone 100 to unlock the phone through voiceprint recognition. For example, when the user needs to operate the mobile phone 100 from a distance (e.g., to open an application on the phone (e.g., a music application, a phone application)), or when the user needs to operate the mobile phone 100 when their hands are occupied (e.g., while doing housework), they can input the voice to be verified through the microphone of the mobile phone 100 to unlock the phone through voiceprint recognition.

[0127] The speech to be verified is speech with specific content. In other implementations, the speech to be verified can also be speech with arbitrary text content.

[0128] In this embodiment, the length of the voice to be verified is 10-30 seconds. This allows the voice to contain relatively rich voiceprint information, which is beneficial for improving the voiceprint recognition rate. However, this application does not limit this. For example, in some embodiments, the length of the voice to be verified is less than 10 seconds. In this case, the length of the voice to be verified is less than the length of the registered voice. This allows users to record shorter voice recordings, which is beneficial for improving the user experience. When the length of the voice to be verified is less than the length of the registered voice, a portion of the voice can be extracted from the voice to be verified and concatenated with the original voice to be verified. This concatenated voice has a length that is basically the same as the registered voice. In this way, in the subsequent steps of this embodiment (described in detail below), the feature parameters extracted from the registered voice and the feature parameters extracted from the voice to be verified have the same dimension, which facilitates the comparison of their similarity. In the description herein, neither the original voice to be verified nor the concatenated voice to be verified is distinguished; both are referred to as the voice to be verified.

[0129] In this paper, concatenating speech A with speech B means joining speech A and speech B end-to-end so that the length of the concatenated speech is the sum of the lengths of speech A and speech B. Furthermore, this application does not limit the order in which speech A and speech B are joined; for example, speech A can be joined after speech B or before speech B.

[0130] S130: Determine the noise contained in the speech to be verified. In this embodiment, the noise contained in the speech to be verified is the sound produced by other sound sources besides the current user of mobile phone 100 in the identification scenario. For example, the sound of household appliances (e.g., vacuum cleaner) and the sound of water running while washing dishes in a home scenario; the sound of car radio and engine in a car scenario; the sound of the projection sound system and the voices of other audience members in a movie theater environment, etc.

[0131] In this embodiment, the sound picked up by the 100 microphones of the mobile phone is identified as noise contained in the speech to be verified, thus facilitating the identification of noise in the speech. However, this application is not limited to this. For example, in some embodiments, it is assumed that the initial segment of the speech to be verified contains only noise components. Therefore, the initial segment of the speech to be verified is copied multiple times and identified as noise contained in the speech. In other embodiments, the speech to be verified is divided into multiple speech frames, and the energy of each speech frame is calculated. Since the energy in noise is usually less than the energy in valid speech, when the energy in a speech frame is less than a predetermined value, the speech frame can be identified as a noise frame, thereby simplifying the noise extraction process. Furthermore, other methods in the prior art can also be used to determine the noise in the speech to be verified, which will not be elaborated upon here.

[0132] The energy of a speech frame is represented by the sum of the squares of the signal values ​​of each speech signal included in the frame. For example, let the signal value of the i-th speech signal in the speech frame be x. i If the number of speech signals in this speech frame is N, then the energy in this speech frame is

[0133] S140: Noise contained in the speech to be verified is superimposed on the registered speech to obtain enhanced registered speech. In this embodiment, in the time domain, the signal value of the noise signal is added to the signal value of the registered speech signal to obtain enhanced registered speech. However, this application is not limited to this; in other embodiments, the superposition of the registered speech signal and the noise signal can also be completed in the frequency domain. This application embodiment achieves the enhancement of the registered speech signal by simply superimposing the signal values ​​of the sound signals, and the algorithm is simple.

[0134] In this embodiment, the length of the noise is equal to the length of the registered speech; in other embodiments, the length of the noise may be less than the length of the registered speech.

[0135] In this embodiment, there are 6 registered voice recordings. Therefore, the noise contained in the voice to be verified is superimposed on each of the 6 registered voice recordings to obtain 6 enhanced registered voice recordings.

[0136] S150: Extract feature parameters of the speech to be verified and the enhanced registered speech. Since the MFCC method can better conform to the auditory perception characteristics of the human ear, this embodiment uses the Mel-Frequency Cepstrum Coefficient (MFCC) method to extract feature parameters from the speech signal.

[0137] First, taking the speech to be verified as an example, the process of feature parameter extraction will be introduced. For ease of description, S will be used. TThis represents the audio signal of the speech to be verified. Before feature extraction, the audio signal S of the speech to be verified is first... T The audio signal is divided into a series of speech frames x(n), where n is the number of speech frames. Considering that the motion model of the vocal organs remains relatively stable within 10–30 ms, the length of each speech frame is 10–30 ms. Specifically, in this embodiment, an audio signal S of length 10 s is... T It is divided into 500 audio frames.

[0138] For audio signal S T After frame segmentation, feature parameters are extracted from each speech frame x(n) using the MFCC method. The MFCC feature extraction method includes Fourier transform, Mel filtering, and discrete cosine transform (DCT) steps on the speech frame x(n). The feature parameters of the speech frame x(n) are the coefficients of each order of cosine function after the DCT. In this embodiment, the order of the DCT is 20; therefore, the MFCC feature parameters of each speech frame x(n) are 20-dimensional.

[0139] After concatenating the feature parameters of each speech frame x(n), the audio signal S of the speech to be verified is obtained. T The MFCC feature parameters can be understood to have a dimension of 20 × 500 = 10000.

[0140] The process for extracting feature parameters from enhanced registered speech can be referred to the above process and will not be repeated here. It can be understood that for each enhanced registered speech, a set of MFCC feature parameters is obtained.

[0141] It should be noted that the above is a principle explanation of the MFCC method. In actual implementation, the extraction process can be adjusted as needed. For example, the extracted MFCC feature parameters can be differentially calculated. For instance, after taking the first and second order differences of the extracted MFCC feature parameters, a set of 60-dimensional MFCC feature parameters is obtained for each speech frame. In addition, other parameters in the extraction process, such as the length and number of speech frames, and the order of the discrete cosine transform, can also be adjusted according to the device's computing power and recognition accuracy requirements.

[0142] In addition to the MFCC method, other methods can be used to extract feature parameters from speech signals, such as the logmel method and the Linear Predictive Cepstrum Coefficient (LPCC) method.

[0143] S160: Perform parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice to obtain the voice template of the current user of mobile phone 100 and the voice template of the owner of mobile phone 100, respectively. This application does not limit the recognition model for parameter recognition; it can be a probabilistic model, such as an identity vector (I-vector) model; or a deep neural network model, such as a time-delay neural network (TDNN) model, a ResNet model, etc.

[0144] The 10,000-dimensional feature parameters of the speech to be verified are input into the recognition model. After dimensionality reduction and abstraction by the recognition model, the speech template of the current user of mobile phone 100 is obtained. In this embodiment, the speech template of the current user of mobile phone 100 is a 512-dimensional feature vector, denoted as A.

[0145] Accordingly, the feature parameters of the 6 enhanced registered voices are input into the recognition model to obtain the voice templates of 6 mobile phone owners. Each voice template is a 512-bit feature vector. The 6 owner voice templates are denoted as B1, B2, ..., B6.

[0146] It is understood that the dimension of the feature vectors mentioned above is only an example and can be adjusted according to the computing power and recognition accuracy requirements of the device.

[0147] S170: Match the voice template of the owner of mobile phone 100 with the voice template of the current user of mobile phone 100 to obtain the recognition result. In this application, the template matching method can be cosine distance method, linear discriminant method, or probabilistic linear discriminant analysis method, etc. The following explanation uses cosine distance method as an example.

[0148] The cosine distance method evaluates the similarity between two feature vectors by calculating the cosine of the angle between them. Taking feature vector A (the feature vector corresponding to the voice template of the current user on mobile phone 100) and feature vector B1 (the feature vector corresponding to the voice template of the owner on mobile phone 100) as an example, the cosine similarity can be expressed as:

[0149]

[0150] Among them, a i Let b be the i-th coordinate in the feature vector A. i Let cosθ be the i-th coordinate in feature vector B1, and let θ1 be the angle between feature vectors A and B1. A larger value of cosθ1 indicates that feature vectors A and B1 are more similar in direction, and the two feature vectors have a higher similarity. Conversely, a smaller value of cosθ1 indicates a lower similarity between the two feature vectors.

[0151] For the 6 enhanced registration voices, 6 owner voice templates B1, B2, ..., B6 are obtained. Their cosine similarities with the current user voice template of mobile phone 100 are cosθ1, cosθ2, ..., cosθ6, respectively. The average of the 6 cosine similarities is taken to obtain the similarity between the current user voice and the owner voice, P = (cosθ1 + cosθ2 + ... + cosθ6) / 6.

[0152] If the similarity P between the current user's voice and the owner's voice is greater than a set value (e.g., 0.8), then the current user of phone 100 is determined to be the owner, and phone 100 unlocks its screen; otherwise, the current user of phone 100 is determined not to be the owner, and phone 100 will not unlock its screen.

[0153] In this embodiment, the speech to be verified is compared with six enhanced registered speech samples to obtain six cosine similarity calculation results. These six cosine similarity results are then averaged to obtain the final similarity P between the current user's speech and the device owner's speech. This embodiment can average the matching error between the speech to be verified and a single enhanced registered speech sample, which helps improve the accuracy and robustness of voiceprint recognition.

[0154] It should be noted that in this embodiment, the voiceprint recognition algorithm (the algorithm corresponding to steps S130 to S170) can be implemented on the mobile phone 100 to achieve offline voiceprint recognition; it can also be implemented in the cloud to save local computing resources of the mobile phone 100. When the voiceprint recognition algorithm is implemented in the cloud, the mobile phone 100 uploads the voice to be verified collected in step S120 to the cloud server. After the cloud server uses the voiceprint recognition algorithm to authenticate the identity of the current user of the mobile phone 100, it returns the authentication result to the mobile phone. The mobile phone 100 decides whether to unlock the screen based on the authentication result.

[0155] The above describes the implementation process of the speech enhancement method in this embodiment. However, it is understood that the above is only an exemplary description. Under the premise of conforming to the inventive concept of this application, those skilled in the art can make other modifications based on the above embodiments.

[0156] For example, in some embodiments, in addition to enhancing the registered speech based on the noise in the speech to be verified, a reverberation component is added to the registered speech to obtain the enhanced registered speech.

[0157] When sound waves propagate indoors, they are reflected multiple times by the room's walls and obstacles. As a result, even after the sound source stops, several sound waves are superimposed and mixed together, making people feel that the sound continues for a period of time after the sound source stops. This phenomenon of sound continuing due to multiple reflections of sound waves is called reverberation.

[0158] When voiceprint recognition is performed in an indoor setting, the speaker's voice will produce reverberation within the room. Reverberation, as an interference factor, can negatively impact the voiceprint recognition rate. Therefore, in some embodiments, reverberation prediction is performed on the registered speech based on the recognition scenario. Specifically, the reverberation of the registered speech in the recognition scenario is simulated, and the reverberation component generated in the recognition scenario is added to the registered speech based on the reverberation simulation. This makes the non-speech components of the voiceprint to be verified as close as possible to the non-speech components in the enhanced registered speech, thereby improving the voiceprint recognition rate and the robustness of the voiceprint recognition method.

[0159] Optionally, the reverberation generated by the registered speech in the recognition scenario can be estimated based on the Image Source Model (ISM) method. The Image Source Model method can simulate the reflection path of sound waves in a room and calculate the room impulse response (RIR) function of the sound field based on the delay and attenuation parameters of the sound waves. After obtaining the room impulse response function, the reverberation generated by the registered speech in the room is obtained by convolving the audio signal of the registered speech with the impulse response function.

[0160] Furthermore, in some cases, such as when using voice control for intelligent robots or smart homes, the distance between the speaker to be verified and the microphone may be relatively far (e.g., exceeding 1 meter). This causes some attenuation of the speaker's voice as it reaches the microphone. Therefore, in some embodiments, to account for the distance between the voice to be verified and the microphone, far-field simulation is performed on the registered voice when estimating reverberation using the mirror source model method. That is, when calculating the room's impact response function using the mirror source model method, the distance between the registered voice and the voice receiving device in the simulated sound field is set based on the distance between the speaker to be verified and the microphone. This allows the acquisition distance of the registered voice to be simulated at the same distance as the voice to be verified, thereby further reducing the differences between the voice to be verified and the enhanced registered voice other than the effective speech components, improving the voiceprint recognition rate and the robustness of the voiceprint recognition method.

[0161] For example, in some embodiments, before comparing the speech to be verified and the enhanced speech (i.e., before step S50), the speech to be verified is further processed, such as echo cancellation, dereverberation, active noise reduction, dynamic gain, and directional pickup. To reduce differences between the speech to be verified and the enhanced registered speech other than effective speech components, the enhanced registered speech undergoes the same front-end processing as the speech to be verified (i.e., the speech to be verified and the enhanced registered speech pass through the same front-end processing algorithm module), thereby further improving the voiceprint recognition rate and the robustness of the voiceprint recognition method.

[0162] For example, in some embodiments, the feature parameter extraction step of the speech signal (i.e., step S150) can be omitted, and the speech signal can be directly recognized by a deep neural network model.

[0163]

Example 2

[0164] refer to Figure 4 This embodiment provides another voice enhancement method. Unlike Embodiment 1, in this embodiment, after acquiring the voice to be verified, the acquisition scenario of the voice to be verified is also identified to obtain the scenario type corresponding to the voice to be verified. Then, in addition to determining the enhanced registration voice based on the noise contained in the voice to be verified, the enhanced registration voice is also determined based on the aforementioned scenario type. Specifically, the voice enhancement method executed by mobile phone 100 according to this embodiment includes the following steps:

[0165] S210: Collect registration voice. Here, the registration voice is the voice recorded by the owner of the mobile phone in a quiet environment, so that there is no obvious noise component in the registration voice.

[0166] S220: Collect the voice to be verified. Here, the voice to be verified is the voice recorded by the current user of the phone in a noisy environment. In other words, the phone user can unlock the phone screen in this scenario using voiceprint recognition. Furthermore, the previous user of the phone is the person currently operating the phone (100), who could be the phone owner or someone else.

[0167] S230: Determine the noise contained in the speech to be verified. In this embodiment, the noise contained in the speech to be verified is the sound generated by other sound sources besides the current user of mobile phone 100 in the identification scene.

[0168] S240: The noise contained in the speech to be verified is superimposed on the registered speech to obtain the enhanced registered speech. In this embodiment, in the time domain, the signal value of the noise signal is added to the signal value of the registered speech signal to obtain the enhanced registered speech.

[0169] In this embodiment, steps S210 to S240 are essentially the same as steps S110 to S140 in Embodiment 1, and the details of the steps will not be repeated. In this embodiment, the number of registered voices is the same as in Embodiment 1, that is, the number of registered voices is 6. Therefore, in step S240, the noise contained in the voice to be verified is superimposed on the 6 registered voices to obtain 6 enhanced registered voices.

[0170] S250: Determine the scene type corresponding to the speech to be verified. Specifically, after acquiring the speech to be verified, the scene type corresponding to the speech is identified using a speech recognition algorithm, such as the GMM method or the DNN method. In the speech recognition algorithm, the label value of the scene type can be: home scene; in-vehicle scene; noisy outdoor scene; meeting room scene; cinema scene, etc.

[0171] S260: Superimpose template noise onto the registered speech. The template noise is noise corresponding to the scene type determined in step S250. For example, the template noise is noise recorded in the scene determined in step S250. Multiple sets of template noise may correspond to each scene type. In this embodiment, it is assumed that the scene type corresponding to the speech to be verified determined in step S250 is a home scene, and three sets of template noise are recorded in the home scene (e.g., sounds generated by home audio-visual equipment, background voices generated during family members' conversations, and / or noise generated by home appliances, etc.).

[0172] Then, the three sets of template noise are superimposed onto the six registered speech samples, forming 3 × 6 = 18 enhanced registered speech samples. Together with the six enhanced registered speech samples formed in step S240, a total of 24 enhanced registered speech samples are formed in this embodiment.

[0173] S270: Extract the feature parameters of the speech to be verified and the feature parameters of the enhanced registered speech, which can be referred to step S150 in Embodiment 1. However, it can be understood that in this embodiment, feature parameters are extracted from the 24 enhanced registered speech samples respectively.

[0174] S280: Perform parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registration voice to obtain the voice template of the current user of mobile phone 100 and the voice template of the owner of mobile phone 100, respectively. Refer to S160 in Embodiment 1. However, it can be understood that in this embodiment, the 24 owner voice templates are denoted as B1, B2, ..., B24.

[0175] S290: Match the voice template of the owner of mobile phone 100 with the voice template of the current user of mobile phone 100 to obtain the recognition result. Refer to step S170 in Embodiment 1. However, it can be understood that in this embodiment, the cosine similarity between the 24 owner voice templates and the current user voice template of mobile phone 100 are cosθ1, cosθ2, ..., cosθ... 24 The average of the 24 cosine similarities is used to obtain the similarity P between the current user's voice and the machine owner's voice: P = (cosθ1 + cosθ2 + ... + cosθ) / 2. 24 ) / twenty four.

[0176] If the similarity P between the current user's voice and the owner's voice is greater than a set value (e.g., 0.8), then the current user of phone 100 is determined to be the owner, and phone 100 unlocks its screen; otherwise, the current user of phone 100 is determined not to be the owner, and phone 100 will not unlock its screen.

[0177] It is understood that the above is merely an exemplary description of the technical solution of this application, and those skilled in the art can make other modifications based on the above. For example, steps S230 and S240 can be omitted, that is, the step of enhancing the registered speech based on the noise contained in the speech to be verified can be omitted, and the registered speech can only be enhanced based on the template noise corresponding to the recognition scenario. In this way, there are 18 enhanced registered speeches, and their corresponding owner speech templates are B7, B2, ..., B24. Accordingly, the similarity between the current user speech and the owner speech of mobile phone 100 is P = (cosθ7 + cosθ2 + ... + cosθ7) / ( ... 24 ) / 18.

[0178] In addition, for technical details not mentioned in this embodiment, such as the main body of the voiceprint recognition algorithm (whether it is implemented locally on the mobile phone 100 or in the cloud), and other voice processing (e.g., reverberation prediction, far-field simulation, front-end processing, etc.), please refer to the introduction in Embodiment 1, and will not be repeated here.

[0179] In this paper, the scene type corresponding to the speech to be verified and the distance between the speaker and the microphone are all environmental feature parameters in the speech to be verified.

[0180]

Example 3

[0181] This embodiment modifies the application scenario of the speech enhancement method based on Embodiment 1. Specifically, in this embodiment, the speech enhancement method is applied to... Figure 5 This illustrates a scenario where the smart speaker 200 is controlled. The smart speaker 200 has voice recognition capabilities, allowing users to interact with it via voice to perform functions such as song playback, weather inquiries, schedule management, and smart home control.

[0182] In this embodiment, when a user issues a voice command to the smart speaker 200 to perform a certain operation (e.g., play the day's schedule, play songs from a specific catalog, control smart home devices, etc.), the smart speaker authenticates the user's identity based on a voiceprint recognition method to determine whether the current user is the owner of the smart speaker 200, and then determines whether the current user has the authority to control the smart speaker 200 to perform the operation.

[0183] Specifically, the speech enhancement method in this embodiment includes:

[0184] S310: Collect registration voice. In this embodiment, the registration voice of the owner of the smart speaker 200 is collected through the microphone of the smart speaker 200. However, this application is not limited to this. In other embodiments, the registration voice can also be collected through a mobile phone, a dedicated microphone, etc. After the registration voice is collected, it can be saved locally on the smart speaker 200 for the smart speaker 200 to recognize the user's voiceprint, thereby achieving offline voiceprint recognition. Alternatively, the registration voice can be uploaded to the cloud to utilize the cloud's computing resources for voiceprint recognition, thus saving the local computing resources of the smart speaker 200.

[0185] S320: Acquire the voice to be verified. In this embodiment, the voice to be verified is acquired through the microphone of the smart speaker 200. The acquisition parameters of the voice to be verified (e.g., the duration and text content of the voice) can be referred to the description in Embodiment 1, and will not be repeated here.

[0186] S330: Determine the noise contained in the speech to be verified. In this embodiment, the speech to be verified is divided into multiple speech frames, and the energy of each speech frame is calculated. Since the energy in noise is usually less than the energy in valid speech, when the energy in a speech frame is less than a predetermined value, the speech frame can be identified as a noise frame, thereby simplifying the noise extraction process.

[0187] S340: The noise contained in the speech to be verified is superimposed on the registered speech to obtain the enhanced registered speech. In this embodiment, in the time domain, the signal value of the noise signal is added to the signal value of the registered speech signal to obtain the enhanced registered speech.

[0188] S350: Extract feature parameters of the speech to be verified and the enhanced registered speech. For example, extract feature parameters of the speech to be verified and the enhanced registered speech using the MFCC method.

[0189] S360: Perform parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice to obtain the voice template of the current user and the voice template of the owner of the smart speaker 200, respectively. This embodiment does not limit the recognition model for parameter recognition; it can be a probabilistic model, such as an identity vector (I-vector) model; or a deep neural network model, such as a time-delay neural network (TDNN) model, a ResNet model, etc.

[0190] S370: Match the voice template of the smart speaker 200 owner with the voice template of the current user of the smart speaker 200 to obtain the recognition result. In this embodiment, the template matching method can be cosine distance method, linear discriminant method, or probabilistic linear discriminant analysis method, etc. If the similarity between the current user's voice and the owner's voice is greater than a set value, then it is determined that the current user of the smart speaker 200 is the owner. At this time, the smart speaker 200 responds to the user's voice command and performs the corresponding operation; otherwise, it is determined that the current user of the smart speaker 200 is not the owner, and the smart speaker 200 ignores the user's voice command.

[0191] It should be noted that, apart from the application scenario, the speech enhancement method in this embodiment is essentially the same as the speech enhancement method in Embodiment 1. Therefore, technical details not described in this embodiment can be referred to the description in Embodiment 1.

[0192] Similar to Embodiment 1, the voiceprint recognition algorithm (the algorithm corresponding to steps S330 to S370) can be implemented on the smart speaker 200 to achieve offline voiceprint recognition; it can also be implemented in the cloud to save local computing resources of the smart speaker 200. When the voiceprint recognition algorithm is implemented in the cloud, the smart speaker 200 uploads the voice to be verified collected in step S120 to the cloud server. After the cloud server uses the voiceprint recognition algorithm to authenticate the identity of the current user of the smart speaker 200, it returns the authentication result to the smart speaker 200. The smart speaker 200 decides whether to execute the user's voice command based on the authentication result.

[0193] In addition, those skilled in the art can also apply the speech enhancement method in Embodiment 2 to... Figure 5 The scenarios for controlling the smart speaker shown will not be elaborated further.

[0194] Now for reference Figure 6The diagram shows a block diagram of an electronic device 400 according to one embodiment of the present application. The electronic device 400 may include one or more processors 401 coupled to a controller hub 403. In at least one embodiment, the controller hub 403 communicates with the processor 401 via a multi-branch bus such as a Front Side Bus (FSB), a point-to-point interface such as a QuickPath Interconnect (QPI), or a similar connection 406. The processor 401 executes instructions controlling general types of data processing operations. In one embodiment, the controller hub 403 includes, but is not limited to, a Graphics & Memory Controller Hub (GMCH) (not shown) and an Input / Output Hub (IOH) (which may be on a separate chip) (not shown), wherein the GMCH includes memory and a graphics controller and is coupled to the IOH.

[0195] Electronic device 400 may also include a coprocessor 402 and a memory 404 coupled to a controller hub 403. Alternatively, one or both of the memory and GMCH may be integrated within the processor (as described in this application), with memory 404 and coprocessor 402 directly coupled to processor 401 and controller hub 403, which is located on a single chip with IOH.

[0196] Memory 404 may be, for example, Dynamic Random Access Memory (DRAM), Phase Change Memory (PCM), or a combination of both. Memory 404 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. The computer-readable storage medium stores instructions, specifically, temporary and permanent copies of those instructions. The instructions may include, when executed by at least one of the processors, causing the electronic device 400 to perform, as... Figure 3 , Figure 4 The instructions for the speech enhancement method. When the instructions are executed on a computer, the computer performs the methods disclosed in Embodiment 1 and / or Embodiment 2 above.

[0197] In one embodiment, coprocessor 402 is a dedicated processor, such as, for example, a high-throughput MIC (Many Integrated Core) processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU (General-purpose computing on graphics processing units), or an embedded processor, etc. Optional properties of coprocessor 402 are indicated by dashed lines. Figure 6 middle.

[0198] In one embodiment, electronic device 400 may further include a network interface (NIC, Network Interface Controller) 406. Network interface 406 may include a transceiver for providing a radio interface for electronic device 400 to communicate with any other suitable device (such as a front-end module, antenna, etc.). In various embodiments, network interface 406 may be integrated with other components of electronic device 400. Network interface 406 can implement the functions of the communication unit in the above embodiments.

[0199] Electronic device 400 may further include input / output (I / O) devices 405. I / O 405 may include: a user interface designed to enable a user to interact with electronic device 400; a peripheral component interface designed to enable peripheral components to also interact with electronic device 400; and / or sensors designed to determine environmental conditions and / or location information related to electronic device 400.

[0200] It is worth noting that, Figure 6 This is merely an example. That is, although... Figure 6 The electronic device 400 shown includes multiple devices such as a processor 401, a controller hub 403, and a memory 404. However, in practical applications, devices using the methods of this application may include only a portion of the devices in the electronic device 400. For example, it may include only the processor 401 and the network interface 406. Figure 6 The properties of the optional devices are shown by dashed lines.

[0201] Now for reference Figure 7 The diagram shown is a block diagram of a SoC (System on Chip) 500 according to an embodiment of this application. Figure 7 In the diagram, similar components share the same reference numerals. Additionally, dashed boxes are an optional feature for more advanced SoCs. Figure 7In this SoC 500, the following components are included: an interconnect unit 550 coupled to the processor 510; a system proxy unit 580; a bus controller unit 590; an integrated memory controller unit 540; a group or one or more coprocessors 520, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random-access memory (SRAM) unit 530; and a direct memory access (DMA) unit 560. In one embodiment, the coprocessor 520 includes a dedicated processor, such as, for example, a network or communication processor, a compression engine, a GPGPU (General-purpose computing on graphics processing units), a high-throughput MIC processor, or an embedded processor.

[0202] Static Random Access Memory (SRAM) cell 530 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. The computer-readable storage medium stores instructions, specifically, temporary and permanent copies of those instructions. These instructions may include, when executed by at least one of the processors, causing the SoC implementation to... Figure 3 , Figure 4 The instructions for the speech enhancement method. When the instructions are executed on a computer, the computer performs the methods disclosed in Embodiment 1 and / or Embodiment 2 above.

[0203] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0204] All methods and implementations of this application can be implemented in the form of software, magnetic files, firmware, etc.

[0205] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application-specific integrated circuit (ASIC), or a microprocessor.

[0206] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this paper are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0207] One or more aspects of at least one embodiment can be implemented by representational instructions stored on a computer-readable storage medium, the instructions representing various logics in a processor, which, when read by a machine, cause the machine to create logic for performing the techniques described herein. These representations, referred to as “IP (Intellectual Property) cores,” can be stored on a tangible computer-readable storage medium and provided to multiple customers or production facilities for loading into manufacturing machines that actually manufacture the logic or processor.

[0208] In some cases, an instruction translator can be used to translate instructions from a source instruction set to a target instruction set. For example, an instruction translator can transform (e.g., using static binary transformation, including dynamically compiled dynamic binary transformation), morph, emulate, or otherwise translate instructions into one or more other instructions that will be processed by the core. Instruction translators can be implemented in software, hardware, firmware, or a combination thereof. Instruction translators can be on the processor, off the processor, or partially on and partially off the processor.

Claims

1. A speech enhancement method applied to electronic devices, characterized in that, The method comprises: collecting a voice to be verified; determining an environmental feature parameter contained in the voice to be verified, the environmental feature parameter comprising a scene type corresponding to the voice to be verified, the scene type corresponding to the voice to be verified being determined according to a scene recognition algorithm for recognizing the voice to be verified; enhancing a registered voice based on the environmental feature parameter; comparing the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user; wherein the step of enhancing the registered voice based on the environmental feature parameter comprises: determining a plurality of groups of template noises corresponding to the scene type based on the scene type corresponding to the voice to be verified, and superimposing the plurality of groups of template noises on a plurality of the registered voices respectively to obtain a plurality of enhanced registered voices.

2. The method of claim 1, wherein, The length of the voice to be verified is less than the length of the registered voice.

3. The method of claim 1, wherein, The scene recognition algorithm is any one of the following: GMM algorithm; DNN algorithm.

4. The method of claim 3, wherein, The scene type of the voice to be verified is any one of the following: home scene; vehicle-mounted scene; outdoor noisy scene; meeting scene; cinema scene.

5. The method of claim 1, wherein, The voice to be verified and the enhanced registered voice are voices processed by the same front-end processing algorithm.

6. The method of claim 5, wherein, The front-end processing algorithm comprises at least one of the following processing algorithms: echo cancellation; dereverberation; active noise reduction; dynamic gain; directional sound pickup.

7. The method of claim 1, wherein, The number of the registered voices is a plurality of; and the plurality of registered voices are enhanced based on the environmental feature parameter to obtain a plurality of enhanced registered voices.

8. The method of claim 1, wherein, The comparison of the voice to be verified with the enhanced registered voice to determine that the voice to be verified and the registered voice come from the same user comprises: extracting feature parameters of the voice to be verified and feature parameters of the enhanced registered voice through a feature parameter extraction algorithm; performing parameter recognition on the feature parameters of the voice to be verified and the feature parameters of the enhanced registered voice through a parameter recognition model to respectively obtain a voice template of a voice to be verified speaker and a voice template of a registered speaker; matching the voice template of the voice to be verified speaker and the voice template of the registered speaker through a template matching algorithm, and determining that the voice to be verified and the registered voice come from the same user according to a matching result.

9. The method of claim 8, wherein: the feature parameter extraction algorithm is an MFCC algorithm, a log mel algorithm or an LPCC algorithm; and / or the parameter recognition model is an identity vector model, a time delay neural network model or a ResNet model; and / or the template matching algorithm is a cosine distance method, a linear discriminant method or a probabilistic linear discriminant analysis method.

10. A speech enhancement system characterized by The method comprises: a terminal device and a server in communication connection with the terminal device, wherein: the terminal device is configured to collect a voice to be verified and send the voice to be verified to the server; The server is configured to determine that the to-be-verified voice contains an environmental feature parameter, enhance the registered voice based on the environmental feature parameter, and compare the to-be-verified voice with the enhanced registered voice to determine that the to-be-verified voice and the registered voice come from the same user; the environmental feature parameter includes a scene type corresponding to the to-be-verified voice, and the scene type corresponding to the to-be-verified voice is determined according to a scene recognition algorithm. The server is further configured to send a determination result that the to-be-verified voice and the registered voice come from the same user to the terminal device. The enhancement of the registered voice based on the environmental feature parameter includes: determining a plurality of groups of template noises corresponding to the scene type based on the scene type corresponding to the to-be-verified voice, and superimposing the plurality of groups of template noises on a plurality of the registered voices respectively to obtain a plurality of enhanced registered voices.

11. The system of claim 10, wherein, The duration of the to-be-verified voice is less than the duration of the registered voice.

12. The system of claim 10, wherein, The scene type of the to-be-verified voice is any one of the following: a home scene, a vehicle-mounted scene, a noisy outdoor scene, a conference scene, and a cinema scene.

13. The system of claim 10, wherein, The to-be-verified voice and the enhanced registered voice are voices processed by the same front-end processing algorithm.

14. The system of claim 13, wherein, The front-end processing algorithm includes at least one of the following processing algorithms: echo cancellation, de-reverberation, active noise reduction, dynamic gain, and directional sound pickup.

15. The system of claim 10, wherein, The number of the registered voices is a plurality of; and the server enhances the plurality of registered voices based on the environmental feature parameter to obtain a plurality of enhanced registered voices.

16. The system of claim 10, wherein, The comparison of the to-be-verified voice with the enhanced registered voice to determine that the to-be-verified voice and the registered voice come from the same user includes: extracting feature parameters of the to-be-verified voice and feature parameters of the enhanced registered voice through a feature parameter extraction algorithm; performing parameter identification on the feature parameters of the to-be-verified voice and the feature parameters of the enhanced registered voice through a parameter identification model to obtain a voice template of a to-be-verified speaker and a voice template of a registered speaker respectively; matching the voice template of the to-be-verified speaker with the voice template of the registered speaker through a template matching algorithm, and determining that the to-be-verified voice and the registered voice come from the same user according to a matching result.

17. An electronic device, comprising: The electronic device includes: a memory configured to store instructions executed by one or more processors of the electronic device; the processor, when executing the instructions in the memory, can cause the electronic device to perform the voice enhancement method of any one of claims 1-9.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to perform the method of any one of claims 1-9.

Citation Information

Patent Citations

  • Speech recognition device

    JP1994138895A

  • Speaker verification apparatus and method utilizing voice information of a registered speaker with extracted feature parameter and calculated verification distance to determine a match of an input voice with that of a registered speaker

    US6879968B1