Upgrade method, upgrade apparatus, and electronic device

The upgrade method enhances voiceprint recognition performance and user experience by using verification speech for silent enrollment during system updates, addressing the disruption caused by traditional re-enrollment processes.

JP7744438B2Active Publication Date: 2025-09-25HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023568018
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-07
Filing Date
2022-04-21
Publication Date
2025-09-25
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

The frequent need for user re-enrollment during voiceprint recognition system upgrades significantly impacts user experience, as it disrupts the seamless operation of voice recognition systems in electronic devices.

Method used

An upgrade method that utilizes verification speech to perform enrollment silently, allowing the system to update voiceprint recognition models without user intervention, thereby maintaining performance and experience.

Benefits of technology

Enables seamless system upgrades by using verification speech for enrollment, improving voiceprint recognition performance and user experience by avoiding the need for repeated user interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007744438000001
    Figure 0007744438000001
  • Figure 0007744438000002
    Figure 0007744438000002
  • Figure 0007744438000003
    Figure 0007744438000003
Patent Text Reader

Abstract

The embodiment of the present application provides an upgrade method, an upgrade apparatus, and an electronic device. The method includes the steps of: an electronic device collects a first verification voice input by a user; processing the first verification voice to obtain a first voiceprint feature by using a first model stored in the electronic device; verifying the identity of the user based on the first voiceprint feature and a first user feature template stored in the electronic device; after the identity of the user is verified, if the electronic device has received a second model, processing the first verification voice to obtain a second voiceprint feature by using the second model; updating the first user feature template based on the second voiceprint feature and updating the first model by using the second model. In the embodiment of the present application, the verification voice acquired in the verification process is used as a new enrollment voice to complete the upgrade and enrollment of the voiceprint recognition system, so that the upgrade of the voiceprint recognition system can be implemented without the user's perception, and both the performance of voiceprint recognition and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of computer technology, and in particular to an upgrade method, an upgrade apparatus, and an electronic device. [Background technology]

[0002] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to Chinese Patent Application No. 202110493970.X, entitled "Upgrade Method, Upgrade Apparatus, and Electronic Device," filed with the State Intellectual Property Administration of China on May 7, 2021, the contents of which are incorporated herein by reference.

[0003] [background] Voiceprint recognition is a technology for automatically identifying and verifying the identity of a speaker based on a voice signal. The basic scheme of voiceprint recognition includes an enrollment procedure and a verification procedure. In the enrollment procedure, a voiceprint recognition system installed in an electronic device extracts voiceprint features from an enrollment speech input by a user by using a pre-trained depth model (referred to herein as a "voiceprint feature extraction model" or "model") and stores the voiceprint features as a user feature template in the electronic device. In the verification procedure, the voiceprint recognition system installed in the electronic device extracts voiceprint features as verification target features from a verification speech by using the same voiceprint feature extraction model as in the enrollment procedure, and verifies the identity of the user based on the verification target features and the user feature template obtained in the enrollment procedure.

[0004] Currently, when upgrading the voiceprint recognition system installed in an electronic device (for example, updating the voiceprint feature extraction model), the enrollment procedure needs to be performed again (specifically, the user inputs the enrollment voice again, and the electronic device uses the new voiceprint feature extraction model to extract voiceprint features from the new enrollment voice as a new user feature template). If re-enrollment is not performed, the verification target features extracted by the electronic device using the new voiceprint feature extraction model in the subsequent verification procedure may differ from those of the old user. Feature Template However, if the registration procedure is repeated every time an upgrade is performed, it will have a significant impact on the user experience.

[0005] Therefore, how to find a balance between voiceprint recognition performance and user experience has become an urgent problem to be solved. Summary of the Invention

[0006] The embodiments of the present application provide an upgrade method, an upgrade apparatus, and an electronic device for upgrading a voiceprint recognition system without user perception, taking into consideration both voiceprint recognition performance and user experience.

[0007] According to a first aspect, there is provided an upgrade method applicable to an electronic device, the method comprising the steps of: collecting, by the electronic device, a first verification speech input by a user; processing the first verification speech by using a first model stored in the electronic device to obtain a first voiceprint feature; and verifying the identity of the user based on the first voiceprint feature and a first user feature template stored in the electronic device, the first user feature template being a voiceprint feature obtained by the electronic device by processing the user's historical verification speech or enrollment speech by using the first model; and, after the user's identity has been verified, if the electronic device has received a second model, processing the first verification speech by using the second model to obtain a second voiceprint feature, updating the first user feature template stored in the electronic device based on the second voiceprint feature, and updating the first model stored in the electronic device by using the second model.

[0008] In an embodiment of the present application, during the upgrade of the voiceprint recognition system, the verification voice obtained in the verification process is used as a new enrollment voice to complete the upgrade and enrollment, so that the upgrade of the voiceprint recognition system can be performed without the user's knowledge, and both the voiceprint recognition performance and the user experience can be improved.

[0009] In a possible implementation, the electronic device may calculate a similarity between the first voiceprint feature and the first user feature template and determine whether the similarity is greater than a first verification threshold corresponding to the first model to verify the user's identity, where if the similarity is greater than the first verification threshold, the verification is successful, and if the similarity does not exceed the first verification threshold, the verification is unsuccessful. After processing the first verification voice by using the second model, if the electronic device has received a second verification threshold corresponding to the second model, the electronic device further updates the first verification threshold based on the second verification threshold.

[0010] In this way, different models correspond to different verification thresholds, and the electronic device can update the verification thresholds by using the fields during system upgrades, thereby further improving the performance of the voiceprint recognition system.

[0011] In a possible implementation, the electronic device may process the first verification voice by using the second model only if the quality of the first verification voice satisfies a first preset condition, including, but not limited to, that the similarity between the first voiceprint feature and the first user feature template is equal to or greater than a first no-enrollment threshold and / or that the signal-to-noise ratio of the first verification voice is equal to or greater than a first signal-to-noise ratio threshold.

[0012] In this way, the quality of the second voiceprint feature can be ensured, and the performance of the voiceprint recognition system undergoing the upgrade can be further ensured.

[0013] In a possible implementation, the first no-enrollment threshold is greater than or equal to a first validation threshold corresponding to the first model.

[0014] In this way, the quality of the second voiceprint feature can be further improved, and the performance of the upgraded voiceprint recognition system can be improved.

[0015] In a possible implementation, if the electronic device receives a second enrollment-free threshold after processing the first verification audio by using the second model, the electronic device may further update the first enrollment-free threshold based on the second enrollment-free threshold, and / or if the electronic device receives a second signal-to-noise ratio threshold after processing the first verification audio by using the second model, the electronic device may further update the first signal-to-noise ratio threshold by using the second signal-to-noise ratio threshold.

[0016] In this way, the enrollment-free threshold, the signal-to-noise ratio threshold, etc. can also be automatically updated to further improve the quality of the second voiceprint feature and improve the performance of the voiceprint recognition system undergoing the upgrade.

[0017] In a possible implementation, the electronic device may update the first user feature template stored in the electronic device based on a preset amount of the second voiceprint features only after the amount of the second voiceprint features acquired by the electronic device cumulatively reaches a preset amount, and update the first model stored in the electronic device by using the second model.

[0018] In this way, it is possible to ensure that the voiceprint recognition system undergoing the upgrade has multiple user feature templates (i.e., second voiceprint features), and the performance of the voiceprint recognition system undergoing the upgrade can be further improved.

[0019] In a possible implementation, after updating a first user feature template stored in the electronic device based on the second voiceprint feature and updating the first model stored in the electronic device by using the second model, the electronic device further collects a second verification voice input by the user, processes the second verification voice using the second model to obtain a third voiceprint feature, and verifies the identity of the user based on the third voiceprint feature and the second voiceprint feature.

[0020] In this way, after the upgrade is completed, the electronic device performs the verification procedure by using the new model and the new user feature template, so that the voiceprint recognition performance of the electronic device can be further improved.

[0021] In a possible implementation, before collecting the first verification voice input by the user, the electronic device may further prompt the user to input a verification voice, for example, by displaying prompt information on a display or outputting a prompt voice through a speaker.

[0022] In this way, the user experience can be improved.

[0023] According to a second aspect, there is provided an upgrade apparatus, which may be an electronic device or a chip in an electronic device, the apparatus including a unit / module configured to perform the method according to the first aspect or any one of the possible implementations of the first aspect.

[0024] For example, the device may include a data collection unit configured to collect a first verification speech input by a user, and a calculation unit configured to: process the first verification speech to obtain a first voiceprint feature by using a first model stored in the device, and verify the identity of the user based on the first voiceprint feature and a first user feature template stored in the device, where the first user feature template is a voiceprint feature obtained by the device by processing the user's historical verification speech or enrollment speech by using the first model; and, after the user's identity is verified, if the device has received a second model, process the first verification speech to obtain a second voiceprint feature by using the second model, and update the first user feature template stored in the device based on the second voiceprint feature, and update the first model stored in the device by using the second model.

[0025] According to a third aspect, an electronic device is provided, including a microphone and a processor. The microphone is configured to collect a first verification speech input by a user. The processor is configured to process the first verification speech to obtain a first voiceprint feature by using a first model stored in the electronic device, and verify the identity of the user based on the first voiceprint feature and a first user feature template stored in the electronic device, the first user feature template being a voiceprint feature obtained by the electronic device by processing the user's historical verification speech or enrollment speech by using the first model. After the user's identity is verified, if the electronic device has received a second model, process the first verification speech to obtain a second voiceprint feature by using the second model, and update the first user feature template stored in the electronic device based on the second voiceprint feature, and update the first model stored in the electronic device by using the second model.

[0026] According to a fourth aspect, there is provided a chip coupled to a memory in an electronic device and configured to perform the method according to the first aspect or any one of the possible implementations of the first aspect.

[0027] According to a fifth aspect, there is provided a computer storage medium storing computer instructions which, when executed by one or more processing modules, perform a method according to the first aspect or any one of the possible implementations of the first aspect.

[0028] According to a sixth aspect, there is provided a computer program product comprising instructions stored therein that, when executed on a computer, enable the computer to carry out a method according to the first aspect or any one of the possible implementations of the first aspect. [Brief explanation of the drawings]

[0029] [Figure 1] 1 is a schematic diagram illustrating a structure of an electronic device according to an embodiment of the present application. [Figure 2] 1 is a schematic diagram illustrating a structure of an electronic device according to an embodiment of the present application. [Figure 3] 1 is a flowchart illustrating an upgrade method according to an embodiment of the present application. [Figure 4] FIG. 10 is a schematic diagram illustrating the electronic device prompting the user to input an enrollment voice. [Figure 5] FIG. 1 is a schematic diagram showing a user activating an electronic device to collect a verification voice. [Figure 6A] FIG. 2 is a schematic diagram illustrating a specific registration-free upgrade process manner according to an embodiment of the present application; [Figure 6B] FIG. 10 is a schematic diagram illustrating another specific registration-free upgrade process manner according to an embodiment of the present application. [Figure 6C] FIG. 10 is a schematic diagram illustrating another specific registration-free upgrade processing manner according to an embodiment of the present application; [Figure 6D] FIG. 10 is a schematic diagram illustrating another specific registration-free upgrade processing manner according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0030] A voiceprint is a sound wave spectrum displayed by an electroacoustic instrument and carrying audio information. Voiceprints are characterized by stability, measurability, and uniqueness. A person's voice can remain stable for a long period of time, even after they reach adulthood. The size and shape of the vocal tract used by people when speaking vary greatly from person to person. Therefore, the voiceprint graphs of any two people are different, and the distribution of resonance peaks in the spectrograms of different people's voices is also different. Voiceprint recognition is the process of comparing the voices of two speakers at the same phonemes to determine whether the two speakers are the same person, thereby performing the function of "recognizing people by listening to their voices."

[0031] From an algorithm perspective, voiceprint recognition can further include text-dependent voiceprint recognition and text-independent voiceprint recognition. In a text-dependent voiceprint recognition system, a user's voice must be spoken based on the specified content, and voiceprint models of people are accurately defined one by one. During recognition, a voice must also be spoken based on the specified content. This allows for better recognition. However, this system requires the user's cooperation. If the voice spoken by the user does not match the specified content, the user cannot be correctly recognized. In a text-independent recognition system, the content spoken by the speaker is not specified, making it more difficult to define a model. However, this system is easy to use and can be widely applied. Considering practicality, text-dependent voiceprint recognition algorithms are currently commonly used in terminal devices.

[0032] Voiceprint recognition can include speaker identification (SI) and speaker verification (SV). Speaker identification is used to determine which one of several people is speaking a voice, which is a "multiple choice" question. Speaker verification is used to verify whether a voice is spoken by a particular person, which is a "true or false" question.

[0033] This specification mainly relates to the function of speaker verification. Hereinafter, unless otherwise specified, the function of voiceprint recognition refers to the function of speaker verification, i.e., the terms "function of speaker verification" and "function of voiceprint recognition" are interchangeable.

[0034] The speaker verification function includes an enrollment procedure and a verification procedure. In the enrollment procedure, before a user officially uses the voiceprint recognition function, the voiceprint recognition system collects enrollment speech input by the user, extracts voiceprint features from the enrollment speech according to a pre-trained depth model (referred to herein as a "voiceprint feature extraction model" or "model"), and saves the voiceprint features as a user feature template in the electronic device. In the verification procedure, when a user uses the voiceprint recognition function, the voiceprint recognition system collects verification speech input by the user, extracts voiceprint features from the verification speech as verification target features by using the same voiceprint feature extraction model as in the enrollment procedure, then evaluates the similarity between the verification target features and the user feature template obtained in the enrollment procedure, and verifies the user's identity based on the evaluation result.

[0035] Since user voice is a sensitive personal information and cannot be stored or uploaded to the cloud, in consideration of privacy safety, voiceprint recognition systems generally operate offline on electronic devices, and a trained voiceprint feature extraction model must be stored in the electronic device in advance.

[0036] However, with the emergence of electronic devices equipped with voiceprint recognition functions, models are rapidly iterated and updated, typically once a year. As electronic devices are updated, the demand for speaker verification technology also continues to increase. When speaker verification technology needs to be upgraded and the upgraded new algorithm needs to be compatible with older devices, the voiceprint recognition system installed in the older devices needs to be upgraded, i.e., a new voiceprint feature extraction model needs to be remotely pushed to the electronic device. After receiving the new voiceprint feature extraction model, the electronic device needs to perform the enrollment procedure again based on the new voiceprint feature extraction model (i.e., the user needs to re-input the enrollment voice, and the voiceprint recognition system uses the new voiceprint feature extraction model to extract voiceprint features from the newly input enrollment voice by the user as a new user feature template). Without re-enrollment, in subsequent verification procedures, by using the new voiceprint feature extraction model, the verification target features extracted by the voiceprint recognition system will be the same as those of the older user. Feature Template However, if the registration procedure is performed again after each upgrade, the user experience will be significantly affected.

[0037] In view of this, the embodiments of the present application provide an upgrade solution. After detecting the new feature extraction model, the electronic device directly performs user enrollment based on the verification voice obtained in the verification procedure, and the user does not need to provide the enrollment voice again for user enrollment. In this way, the voiceprint recognition system is upgraded without the user's perception, and both the voiceprint recognition performance and the user experience are improved.

[0038] It should be understood that the technical solutions in the embodiments of the present application can be applied to any electronic device with voiceprint recognition function. See Fig. 1. The electronic device in the embodiments of the present application has at least a data collection unit 01, a storage unit 02, a communication unit 03, and a calculation unit 04. These units can be connected and communicate with each other via input / output (I / O) interfaces.

[0039] The data collection unit 01 is configured to collect voice input by a user (such as an enrollment voice or a verification voice). A specific embodiment of the data collection unit 01 may be a microphone or an acoustic sensor, etc.

[0040] The storage unit 02 stores the voiceprint feature extraction model, the threshold value used by the voiceprint recognition function, and the user information obtained by the user registration module in the calculation unit 04. Feature Template and configured to store the

[0041] The communication unit 03 is configured to receive a new voiceprint feature extraction model, and may be further configured to receive a new threshold value and provide the new threshold value to the calculation unit 04 .

[0042] The calculation unit 04 includes: - Based on the enrollment voice obtained by the data collection unit 01, the user Feature Template Extract and user Feature Template a user registration module 401 configured to provide a verification module 402 with - The verification voice acquired by the data collection unit 01 and the user stored in the storage unit 02 Feature Template a verification module 402 configured to verify the identity of the speaker based on the model and the threshold to obtain a verification result; and - based on the new model received by the communication unit 03 and the verification speech and evaluation results (optional) obtained by the verification module, Feature Template Determine the new user Feature Template And based on the new model, the old user in the storage unit 02 Feature Template and a registration-free upgrade module 403 configured to update the old model. Optionally, the registration-free upgrade module further updates the threshold value stored in the storage unit 02.

[0043] In the embodiments of the present application, there may be multiple specific product forms related to electronic devices, including, but not limited to, mobile phones, tablet computers, artificial intelligence (AI) intelligent voice terminals, wearable devices, augmented reality (AR) / virtual reality (VR) devices, in-vehicle terminals, laptop computers, desktop computers, and smart home devices (such as smart televisions or smart speakers).

[0044] For example, the electronic device is a mobile phone. Figure 2 is a schematic diagram showing the hardware structure of a mobile phone 100 according to an embodiment of the present application.

[0045] Mobile phone 100 includes processor 110, internal memory 121, external memory interface 122, camera 131, display 132, sensor module 140, subscriber identity module (SIM) card interface 151, keys 152, audio module 160, speaker 161, receiver 162, microphone 163, headset jack 164, universal serial bus (USB) interface 170, charge management module 180, power management module 181, battery 182, mobile communication module 191, and wireless communication module 192. In some other embodiments, mobile phone 100 may further include motors, indicators, keys, and the like.

[0046] It should be understood that the hardware configuration shown in Figure 2 is merely exemplary. Mobile phone 100 in embodiments of the present application may have more or fewer components than those of mobile phone 100 shown, may combine two or more components, or may have a different configuration of components. The components shown may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0047] Processor 110 may include one or more processing units. For example, processor 110 may include an application processor (AP), a modem, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU), etc. The different processing units may be separate components or may be integrated into one or more processors.

[0048] In some embodiments, a buffer may further be arranged in processor 110 to store instructions and / or data. For example, the buffer in processor 110 may be a cache. The buffer may be configured to store instructions and / or data that have just been used, generated, or reused by processor 110. When the instructions or data need to be used again, processor 110 may retrieve the instructions or data directly from the buffer. This contributes to reducing the time it takes for processor 110 to retrieve the instructions or data, contributing to improving system efficiency.

[0049] The internal memory 121 may be configured to store programs and / or data, and in some embodiments, the internal memory 121 includes a program storage area and a data storage area.

[0050] The program storage area may be configured to store an operating system (e.g., an operating system such as Android or iOS) and computer programs required by at least one function. For example, the program storage area may store a computer program required by a voiceprint recognition function (e.g., a voiceprint recognition system). The data storage area may be configured to store data (e.g., voice data) created and / or collected during the use of the mobile phone 100. For example, when the processor 110 calls up programs and / or data stored in the internal memory 121, the mobile phone 100 is enabled to execute corresponding methods for implementing one or more functions. For example, when the processor 110 calls up some programs and / or data in the internal memory, the mobile phone 100 executes the upgrade method provided in the embodiments of the present application.

[0051] The internal memory 121 may be a high-speed random access memory and / or a non-volatile memory, etc. For example, the non-volatile memory may include one or more disk memory devices, flash memory devices, and / or universal flash storage (UFS).

[0052] The external memory interface 122 may be configured to connect to an external memory card (such as a microSD card) to expand the storage capabilities of the mobile phone 100. The external memory card communicates with the processor 110 via the external memory interface 122 to implement data storage functions. For example, the mobile phone 100 may save files, such as images, music, and videos, to the external memory card via the external memory interface 122.

[0053] The camera 131 may be configured to capture dynamic images, static images, and the like. Typically, the camera 131 includes a lens and an image sensor. Using an optical image generated by the lens, an object is projected onto the image sensor, which is then converted into an electrical signal for subsequent processing. For example, the image sensor may be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) photoelectric transistor. The image sensor converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP. It should be noted that the mobile phone 100 may include one or N cameras 131, where N is a positive integer greater than 1.

[0054] The display 132 may include a display panel configured to display a user interface. The display panel may use a liquid crystal display (LCD), an organic light emitting diode (OLED), an active matrix organic light emitting diode (AMOLED), a flexible light emitting diode (FLED), a mini LED, a micro LED, a micro OLED, or a quantum dot light emitting diode (QLED), etc. It should be noted that the mobile phone 100 may include one or M displays 132, where M is a positive integer greater than 1. For example, the mobile phone 100 may implement display functionality by using a GPU, the display 132, an application processor, etc.

[0055] The sensor module 140 may include one or more sensors, such as a touch sensor 140A, a gyroscope 140B, an acceleration sensor 140C, a fingerprint sensor 140D, and a pressure sensor 140E. In some embodiments, the sensor module 140 may further include an ambient light sensor, a distance sensor, a proximity sensor, a bone conduction sensor, a temperature sensor, and the like.

[0056] The SIM card interface 151 is configured to connect to a SIM card. The SIM card can be inserted into or removed from the SIM card interface 151 to connect to or disconnect from the mobile phone 100. The mobile phone 100 can support one or K SIM card interfaces 151, where K is a positive integer greater than 1. The SIM card interface 151 can support a nano-SIM card, a micro-SIM card, and / or other SIM cards. Multiple cards can be inserted into the same SIM card interface 151 at the same time. The multiple cards can be of the same type or different types. The SIM card interface 151 can be compatible with SIM cards of different types. The SIM card interface 151 can be compatible with an external memory card. The mobile phone 100 interacts with a network through the SIM card to implement functions such as telephone calls and data communications. In some embodiments, the mobile phone 100 may use an eSIM, i.e., an embedded SIM card, which may be embedded in the mobile phone 100 and cannot be separated from the mobile phone 100.

[0057] The keys 152 may include a power key, a volume key, etc. The keys 152 may be mechanical keys or touch keys. The mobile phone 100 may receive key inputs and generate key signal inputs related to user settings and function control of the mobile phone 100.

[0058] The mobile phone 100 may implement voice functions such as, for example, voice playback, recording, voiceprint registration, voiceprint verification, and voiceprint recognition functions via the voice module 160, speaker 161, receiver 162, microphone 163, headset jack 164, and application processor.

[0059] Audio module 160 may be configured to perform digital-to-analog and / or analog-to-digital conversion on audio data and may further be configured to encode and / or decode audio data. For example, audio module 160 may be located independently of the processor, or may be located within processor 110, or some functional modules of audio module 160 may be located within processor 110.

[0060] The speaker 161 is also called a "loudspeaker" and is configured to convert audio data into sound and play the sound. 100 may be configured to play music, answer calls in hands-free mode, or provide voice prompts via speaker 161.

[0061] The receiver 162, also called an "earphone," is configured to convert audio data into sound and play the sound. 100 When answering a call by using the receiver 162, the receiver 162 can be placed in close proximity to the human ear to answer the call.

[0062] The microphone 163, also referred to as a "mike" or "mic," is configured to collect sounds (e.g., ambient sounds, including sounds made by a person or a device). When making a call or transmitting voice, a user may speak from their mouth near the microphone 163, and the microphone 163 collects the voice emitted by the user. When the voiceprint recognition function of the mobile phone 100 is enabled, the microphone 163 may collect ambient sounds in real time and obtain voice data.

[0063] It should be noted that at least one microphone 163 may be arranged on the mobile phone 100. For example, two microphones 163 may be arranged on the mobile phone 100 to implement a noise reduction function in addition to sound collection. In another example, three, four, or more microphones 163 may be arranged on the mobile phone 100 to implement functions such as sound source recognition or directional recording when implementing sound collection and noise reduction.

[0064] The headset jack 164 is configured to connect to a wired headset. The headset jack 164 may be a USB interface 170 or a 3.5 mm open mobile terminal. End of the year The interface may be a standard interface of the open mobile terminal platform (OMTP) or a standard interface of the cellular telecommunications industry association of the USA (CTIA).

[0065] The USB interface 170 is an interface that complies with the USB standard, and may specifically be a mini-USB interface, a micro-USB interface, a USB Type-C interface, or the like. The USB interface 170 may be configured to connect to a charger to charge the mobile phone 100, or to perform data transmission between the mobile phone 100 and a peripheral device, or to connect to a headset to play audio through the headset. For example, the USB interface 170, in addition to functioning as the headset jack 164, may be further configured to connect to another mobile phone 100, such as an AR device or a computer.

[0066] The charging management module 180 is configured to receive charging input from a charger, which may be a wireless charger or a wired charger. In some embodiments involving wired charging, the charging management module 180 may receive charging input from the wired charger via the USB interface 170. In some embodiments involving wireless charging, the charging management module 180 may receive wireless charging input through a wireless charging coil of the mobile phone 100. The charging management module 180 may charge the battery 182 while the power management module 180 1 The mobile phone 100 is supplied with power through the power supply.

[0067] Power management module 181 is configured to connect battery 182, charging management module 180, and processor 110. Power management module 181 receives input from battery 182 and / or charging management module 180 and provides power to processor 110, internal memory 121, display 132, camera 131, etc. Power management module 181 may be further configured to monitor parameters such as battery capacity, battery cycle count, and battery health (leakage and impedance). In some other embodiments, power management module 181 may be alternatively located in processor 110. In some other embodiments, power management module 181 and charging management module 180 may alternatively be located in the same device.

[0068] The mobile communication module 191 may be applied to the mobile phone 100 and provide solutions including wireless communication such as 2G, 3G, 4G, and 5G. The mobile communication module 191 may include a filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like.

[0069] The wireless communication module 192 may be applied to the mobile phone 100 to provide solutions including wireless communication such as WLAN (e.g., Wi-Fi network), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR), etc. The wireless communication module 192 may be one or more components that integrate at least one communication processing module.

[0070] In some embodiments, Antenna 1 of mobile phone 100 is connected to mobile communication module 191, and Antenna 2 is connected to wireless communication module 192, so that mobile phone 100 can communicate with another device. Specifically, mobile communication module 191 can communicate with another device via Antenna 1, and wireless communication module 192 can communicate with another device via Antenna 1. 2 may communicate with another device via antenna 2.

[0071] For example, the mobile phone 100 may receive upgrade information (including a new voiceprint feature extraction model and a new threshold value, etc.) from another device via the wireless communication module 192, and then update the voiceprint recognition system installed in the electronic device based on the upgrade information (e.g., update the voiceprint feature extraction model, update the threshold value, etc.). Optionally, the other device may be a server of a cloud service vendor, for example, a platform established and maintained by the vendor or operator of the mobile phone 100 and providing required services in an on-demand and easily scalable manner by utilizing a network, and specifically, may be a server of the mobile phone vendor Huawei, for example. Of course, the other device may alternatively be another electronic device. The present application is not limited thereto.

[0072] The voiceprint recognition method provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings and application scenarios. The following embodiments can all be implemented in the mobile phone 100 having the above-mentioned hardware configuration.

[0073] FIG. 3 illustrates a schematic diagram of the steps of an upgrade method according to one embodiment of the present invention.

[0074] S301: The electronic device collects a first enrollment voice input by a user.

[0075] Specifically, the electronic device may collect ambient sounds through the microphone 163 to obtain a first enrollment voice recorded by the user.

[0076] In a specific implementation, the user may be prompted by the electronic device to speak a first enrollment voice. For example, as shown in FIG. 4 , the electronic device may display text on the display 132 to prompt the user to speak the enrollment voice "1234567." In another example, the electronic device may also perform a voice prompt through the speaker 161. There may be multiple scenarios in which the electronic device prompts the user to input the enrollment voice. For example, when the user enables the voiceprint recognition function of the electronic device for the first time, the electronic device may automatically prompt the user to speak the enrollment voice, or when the user enables the voiceprint recognition function of the electronic device for the first time, the user may operate the electronic device to prompt the user to speak the enrollment voice, or when the user subsequently enables the voiceprint recognition function, the user may activate the electronic device to prompt the user to speak the enrollment voice, as needed.

[0077] Optionally, when performing voiceprint enrollment, the user can input the enrollment voice multiple times to improve the accuracy of voiceprint recognition.

[0078] S302: The electronic device processes a first enrollment voice by using a pre-stored first model to obtain a first user feature template, and saves the first user feature template.

[0079] The first model is a voiceprint feature extraction model that is pre-trained by using a neural network. The input of this first model is speech, and the output is voiceprint features corresponding to the input speech. The algorithm on which the first model is based may be, but is not limited to, a filter bank (FBank) algorithm, a mel-frequency cepstral coefficients (MFCC) algorithm, a D-vector algorithm, among other algorithms.

[0080] Optionally, because the quality of the enrollment voice significantly affects the recognition accuracy, the electronic device may first perform quality detection on the first enrollment voice by using the pre-stored first model before processing the first enrollment voice. The first enrollment voice is used for enrollment only if the quality of the first enrollment voice meets a first preset requirement (in other words, the first enrollment voice is processed by using the pre-stored first model to obtain a first user feature template, and the first user feature template is saved). If the quality is poor, the first enrollment voice may be rejected from being used for enrollment, or the user may be prompted to re-enter the enrollment voice and try enrolling again, etc.

[0081] For example, the electronic device prompts the user to speak a keyword for the voice assistant, such as "Xiaoyi Xiaoyi," three times on the display 132. Each time the user speaks "Xiaoyi Xiaoyi," the microphone 163 of the electronic device sends the collected voice to the processor 110 of the mobile phone. The processor 110 segments the voice corresponding to the keyword and uses the segmented voice as the enrollment voice. The processor 110 then determines the signal-to-noise ratio of the enrollment voice and determines whether the signal-to-noise ratio meets the requirement. If the signal-to-noise ratio is smaller than a set threshold (i.e., the noise is too loud), the enrollment is rejected. For voices that pass the signal-to-noise ratio detection, the processor 110 performs calculations on the voice by using a first model to determine the user's voice. Feature Template Get the user Feature Template is stored in the internal memory 121.

[0082] Optionally, the electronic device may improve the accuracy of voiceprint recognition by storing multiple first user feature templates.

[0083] Steps S301 and S302 are performed before the electronic device registers the user's voice for the first time, specifically before the user uses the voiceprint recognition function for the first time. Once registration is complete for the first time, the user may begin using the voiceprint recognition function, as shown in steps S303 and S304.

[0084] S303: The electronic device collects a first verification voice input by the user.

[0085] In a specific implementation, the user may be prompted by the electronic device to speak a verification voice. The method used by the electronic device to prompt the user to speak the verification voice is similar to the method used by the electronic device to prompt the user to speak the enrollment voice, and the repeated parts will not be described separately.

[0086] There may be multiple scenarios in which the electronic device prompts the user to input a verification voice. For example, after the electronic device is powered on by the user, the electronic device may automatically prompt the user to speak a first verification voice. This first verification voice is used to confirm the user's identity and unlock the electronic device. Alternatively, the electronic device may automatically prompt the user to speak a first verification voice when the user opens an encrypted application (e.g., a diary). The first verification voice is used to confirm the user's identity and unlock the application. Alternatively, the electronic device may automatically prompt the user to speak a first verification voice when the user opens the application and attempts to log in to an account. The first verification voice is used to confirm the user's identity and automatically enter the user's account and password.

[0087] The electronic device may be activated by a user's operation and collect a first verification voice input by the user. For example, after receiving a verification instruction by the user activating the verification instruction by opening the electronic device, the electronic device prompts the user to input the first verification voice and collects the first verification voice input by the user. For example, the user may activate the verification instruction by touching a position on an icon corresponding to a voiceprint recognition function installed on the touchscreen of the electronic device, causing the electronic device to prompt the user to speak the first verification voice. In another example, the user may activate the verification instruction by manipulating a physical entity (e.g., a physical key, a mouse, or a joystick). In another example, the user may activate the verification instruction by using a specific gesture (e.g., double-clicking on the touchscreen of the electronic device), causing the electronic device to prompt the user to speak the first verification voice. In another example, the user may speak the keyword "voiceprint recognition" into the electronic device (e.g., a smartphone or an in-car device). After collecting the keyword "voiceprint recognition" uttered by the user through the microphone 163, the electronic device activates a verification prompt and prompts the user to speak a first verification voice.

[0088] Alternatively, when a user utters a control command to control the electronic device, the electronic device collects the control command and uses the control command as a first verification voice to perform voiceprint recognition. Specifically, when the electronic device receives the control command, it activates a verification instruction and performs voiceprint recognition using the control command as a first verification voice. For example, as shown in FIG. 5 , a user may utter a control command "play music" to an electronic device (e.g., a smartphone or an in-car device). After collecting the user's voice "play music" through the microphone 163, the electronic device performs voiceprint recognition using the voice as a first verification voice. In another example, a user may utter a control command "set to 27 degrees Celsius" to an electronic device (e.g., a smart air conditioner). After collecting the user's voice "set to 27 degrees Celsius" through the microphone 163, the electronic device performs voiceprint recognition using the voice as a first verification voice.

[0089] Optionally, when performing voiceprint verification, the user may input the verification voice multiple times to improve the accuracy of voiceprint recognition.

[0090] S304: The electronic device processes the first verification voice by using the first model to obtain a first voiceprint feature, and verifies the identity of the user based on the first voiceprint feature and a first user feature template stored in the electronic device.

[0091] First, in the enrollment procedure in step S302, the electronic device inputs a first verification voice into the same model (in other words, the first model), and the first model outputs voiceprint features.

[0092] The electronic device then calculates a similarity between the first voiceprint feature and the first user feature template. Methods for calculating the similarity may include, but are not limited to, a cosine distance (CDS) algorithm, a linear discriminant analysis (LDA) algorithm, and a probabilistic linear discriminant analysis (PLDA) algorithm, among other algorithms. For example, in evaluating a cosine distance model, a feature vector of the first voiceprint feature to be verified and a feature vector of the user feature template may be calculated. Feature Template The cosine value between the feature vector of the first voiceprint feature to be verified and the user's voiceprint feature is calculated, and the cosine value is used as the similarity score (in other words, the evaluation result). For example, in the evaluation of a probabilistic linear discriminant analysis model, a pre-trained probabilistic linear discriminant analysis model is used to calculate the similarity score between the first voiceprint feature to be verified and the user's voiceprint feature. Feature Template The similarity score (in other words, the evaluation result) between multiple users is calculated. Feature Template If the first voiceprint feature to be verified and multiple users have been registered, Feature Template It should be understood that a fused matching evaluation can be performed based on:

[0093] The electronic device then selects whether to accept or reject the control command corresponding to the verification voice based on the evaluation result. For example, the electronic device determines whether the similarity is greater than a first verification threshold corresponding to the first model. If the similarity is greater than the first verification threshold, the verification is successful. Specifically, the speaker of the verification voice matches the speaker of the enrollment voice, and then a corresponding control operation (e.g., unlocking the electronic device, opening an application, or logging in to an account with a password) is performed. If the similarity does not exceed the first verification threshold, the verification is unsuccessful. Specifically, the speaker of the verification voice does not match the speaker of the enrollment voice, and the corresponding control operation is not performed. Optionally, if the verification fails, the electronic device may display the verification result on the display 132 and prompt the user to re-enter the verification voice, or the electronic device may prompt the user to re-enter the verification voice and attempt re-verification.

[0094] S305: After the user's identity is verified, if the electronic device has received the second model, the electronic device processes the first verification voice by using the second model to obtain second voiceprint features, updates the first user feature template stored in the electronic device based on the second voiceprint features, and updates the first model stored in the electronic device by using the second model.

[0095] It should be understood that the time at which the electronic device receives the second model is later than the time at which the electronic device receives the first model, i.e., the second model is an updated model relative to the first model.

[0096] The second model is a pre-trained voiceprint feature extraction model using a neural network. The input of the second model is speech, and the output is voiceprint features corresponding to the input speech. The algorithm on which the second model is based may be, but is not limited to, the FBank algorithm, the MFCC algorithm, and the D-vector algorithm, among other algorithms.

[0097] The source of the second model may be actively pushed by a cloud server. For example, if an upgrade of the voiceprint recognition model installed on the electronic device is required, the cloud server may push a new model (such as a second model) to the electronic device.

[0098] After receiving the second model, the electronic device uses the first verification voice obtained in the previous verification procedure (assuming that the verification result of the verification procedure is successful, ensures that the verification voice (e.g., the first verification voice) obtained in the verification procedure is spoken by the enrollee), uses the first verification voice as a new enrollment voice, and processes the first verification voice by using the second model to obtain second voiceprint features. Then, the electronic device updates the first user feature template stored in the electronic device based on the second voiceprint features, and updates the first model stored in the electronic device by using the second model, thereby performing upgrade and enrollment without the user's knowledge (i.e., without the user having to perform the operation of recording the enrollment voice).

[0099] Specific implementations of the electronic device updating the first user feature template stored in the electronic device based on the second voiceprint feature include, but are not limited to, the following two ways.

[0100] Mode 1: The electronic device directly uses the second voiceprint feature as a new user feature template (to distinguish it from the first user feature template, the second voiceprint feature is referred to as the second user feature template in this specification), and replaces the first user feature template stored in the electronic device with the second user feature template.

[0101] Mode 2: The electronic device receives the second voice print characteristic and the first user Features A weighting / combination is performed on the templates to obtain a third user feature template, and the first user feature template stored on the electronic device is replaced with the third user feature template.

[0102] It should be understood that the two modes mentioned above are merely examples and are not limiting in nature.

[0103] Similarly, in the case of model updating, the electronic device may directly replace the first model with the second model, or the electronic device may perform weighting / combination on the first model and the second model. This application is not limited thereto. Furthermore, instead of directly updating the entire model, the electronic device may receive only some updated parameters in the model and then update the relevant parameters of the first model based on the updated parameters.

[0104] Optionally, different models may correspond to different verification thresholds. After processing the first verification audio using the second model, if the electronic device receives a second verification threshold corresponding to the second model, the electronic device may further update the first verification threshold based on the second verification threshold to perform the verification threshold update. In this case, the electronic device may process the first verification audio using the second model only after determining that both the second model and the second verification threshold have been received. The verification threshold may be updated by replacing the first verification threshold with the second verification threshold, or by performing weighting / combining on the second verification threshold and the first verification threshold and replacing the first verification threshold with the verification threshold obtained by performing the weighting / combining. The present application is not limited thereto.

[0105] Optionally, to ensure the performance of the upgraded voiceprint recognition system, the electronic device may process the first verification voice by using the second model (in other words, use the first verification voice as a new enrollment voice) only after determining that the quality of the first verification voice meets second preset requirements.

[0106] The first preset condition includes, but is not limited to, the following two types:

[0107] (1) The similarity between the first voiceprint feature and the first user feature template is equal to or greater than a first enrollment-free threshold.

[0108] The first no-registration threshold may be calculated based on the first verification threshold used in the verification procedure (e.g., the first no-registration threshold is several decibels higher than the first verification threshold), or may be preset by the electronic device (e.g., received from a cloud server and stored in advance). The present application is not limited thereto.

[0109] (2) The signal-to-noise ratio of the first verification audio is greater than or equal to a first signal-to-noise ratio threshold.

[0110] The first signal-to-noise ratio threshold may be obtained based on a specified threshold used in the registration procedure (S301) (e.g., the first signal-to-noise ratio threshold is equal to the specified threshold, or the first signal-to-noise ratio threshold is several decibels higher than the specified threshold), or may be preset by the electronic device (e.g., received from a cloud server and stored in advance). The present application is not limited thereto.

[0111] Typically, the first signal-to-noise ratio threshold is 20 dB or more. Optionally, in specific implementation, the value of the first signal-to-noise ratio threshold can be further fine-tuned according to the specific form of the electronic device. For example, for a mobile phone, the first signal-to-noise ratio threshold can be set to 22 dB, and for a smart speaker, the first signal-to-noise ratio threshold can be set to 20 dB.

[0112] Further optionally, the first enrollment-free threshold is equal to or greater than the first verification threshold corresponding to the first model, thereby ensuring that the quality of the verification voice used as the new enrollment voice is high, and further improving the performance of the upgraded voiceprint recognition system.

[0113] Furthermore, optionally, the cloud server may further push a new registration-free threshold to the electronic device, and the electronic device updates the registration-free threshold. For example, if the electronic device determines that the quality of the first verification voice meets the requirements based on the first registration-free threshold and receives a second registration-free threshold after processing the first verification voice by using the second model, the electronic device updates the first registration-free threshold based on the second registration-free threshold. The manner of updating the registration-free threshold may include replacing the first registration-free threshold with the second registration-free threshold, or performing weighting / combination on the second registration-free threshold and the first registration-free threshold, and replacing the first registration-free threshold with the registration-free threshold obtained by performing weighting / combination. The present application is not limited to this. In this way, the performance of the upgraded voiceprint recognition system can be further improved.

[0114] Furthermore, optionally, the cloud server may further push a new signal-to-noise ratio threshold to the electronic device, and the electronic device updates the signal-to-noise ratio threshold. For example, if the electronic device determines that the quality of the first verification voice meets the requirements based on the first signal-to-noise ratio threshold and receives the second signal-to-noise ratio threshold after processing the first verification voice using the second model, the electronic device updates the first signal-to-noise ratio threshold by using the second signal-to-noise ratio threshold. The signal-to-noise ratio threshold may be updated by replacing the first signal-to-noise ratio threshold with the second signal-to-noise ratio threshold, or by performing weighting / combining on the second signal-to-noise ratio threshold and replacing the first signal-to-noise ratio threshold with the signal-to-noise ratio threshold obtained by performing weighting / combining. The present application is not limited thereto. In this way, the performance of the upgraded voiceprint recognition system can be further improved.

[0115] It should be understood that the two conditions mentioned above (i.e., the no-registration threshold and the signal-to-noise ratio threshold) may be implemented separately or simultaneously. This application is not limited thereto. Furthermore, the two conditions mentioned above are merely examples and are not limiting. In specific implementation, the first preset condition may alternatively be implemented in another manner.

[0116] Optionally, when multiple user feature templates are stored in the electronic device, the electronic device may update the first user feature template stored in the electronic device by using a preset amount of the second user feature template only after the cumulative acquisition amount of the second user feature template reaches a preset amount, and update the first model stored in the electronic device by using the second model.

[0117] Specifically, after the verification procedure is completed and the user's identity is verified, the electronic device processes the currently acquired verification voice by using the second model to obtain at least one second voiceprint feature, and stores each second voiceprint feature in the internal memory 121 as a second user feature template; then, determines whether the amount of second user feature templates stored in the internal memory 121 has reached a preset amount (e.g., 3); if the amount of second user feature templates has not reached the preset amount, waits for the next verification procedure, and in the next verification procedure, collects a second user feature template based on the second model, and in the next verification procedure, collects the verification voice; if the amount of second user features has reached the preset amount, updates all first user feature templates by using all second user feature templates, and updates the first model stored in the electronic device by using the second model.

[0118] In this way, it can be ensured that the upgraded voiceprint recognition system has multiple available second user feature templates, and the performance of the upgraded voiceprint recognition system can be further improved.

[0119] It should be understood that the above-mentioned prerequisites (e.g., determining whether the electronic device has received the second model and / or the second verification threshold, and determining whether the quality of the first verification voice meets the second preset requirement) used to activate the electronic device to use the confirmation voice as a new enrollment voice (in other words, the electronic device processes the first verification voice by using the second model to obtain second voiceprint features (in other words, the second user feature template)) may be implemented in combination, and the order in which the electronic device determines the prerequisites may be changed.

[0120] For example, some possible specific implementations for step S305 are provided below.

[0121] In a first implementation, as shown in FIG. 6A, a verification procedure is executed, and after the verification is successful, the electronic device determines whether the evaluation result obtained in the verification procedure is greater than or equal to a first enrollment-free threshold, and if the evaluation result is less than the first enrollment-free threshold, proceeds to the next verification procedure; if the evaluation result is greater than or equal to the first enrollment-free threshold, continues to determine whether the electronic device has received a second model and a second verification threshold, and if the electronic device has not received the second model and the second verification threshold, proceeds to the next verification procedure; and if the electronic device has received the second model and the second verification threshold, processes the first verification voice by using the second model to obtain a second user feature template. The electronic device determines whether the accumulated amount of the second user feature templates has reached a preset amount, and if the accumulated amount of the second user feature templates has not reached the preset amount, proceeds to the next verification procedure, and if the accumulated amount of the second user feature templates has reached the preset amount, updates all first user feature templates stored in the electronic device by using all second user feature templates, updates the first model stored in the electronic device by using the second model, and updates the first verification threshold stored in the electronic device based on the second verification threshold.

[0122] In a second implementation, as shown in FIG. 6B, a verification procedure is executed, and after the verification is successful, the electronic device determines whether the electronic device has received a second model and a second verification threshold; if the electronic device has not received the second model and the second verification threshold, it proceeds to the next verification procedure; if the electronic device has received the second model and the second verification threshold, it continues to determine whether the evaluation result obtained in the verification procedure is greater than or equal to the first enrollment-free threshold; if the evaluation result is less than the first enrollment-free threshold, it proceeds to the next verification procedure; if the evaluation result is greater than or equal to the first enrollment-free threshold, it processes the first verification voice by using the second model to obtain a second user feature template. The electronic device determines whether the accumulated amount of the second user feature templates has reached a preset amount, and if the accumulated amount of the second user feature templates has not reached the preset amount, proceeds to the next verification procedure, and if the accumulated amount of the second user feature templates has reached the preset amount, updates all of the first user feature templates stored in the electronic device by using all of the second user feature templates, updates the first model stored in the electronic device by using the second model, and updates the first verification threshold stored in the electronic device based on the second verification threshold.

[0123] In a third implementation, as shown in FIG. 6C , a verification procedure is performed, and after the verification is successful, the electronic device determines whether the signal-to-noise ratio of the first verification voice is greater than or equal to a first signal-to-noise ratio threshold, and if the signal-to-noise ratio of the first verification voice is less than the first signal-to-noise ratio threshold, proceeds to the next verification procedure, and if the signal-to-noise ratio of the first verification voice is greater than or equal to the first signal-to-noise ratio threshold, continues to determine whether the electronic device has received a second model and a second verification threshold, and if the electronic device has not received the second model and the second verification threshold, proceeds to the next verification procedure, and if the electronic device has received the second model and the second verification threshold, processes the first verification voice by using the second model to obtain a second user feature template. The electronic device determines whether the accumulated amount of the second user feature templates has reached a preset amount, and if the accumulated amount of the second user feature templates has not reached the preset amount, proceeds to the next verification procedure, and if the accumulated amount of the second user feature templates has reached the preset amount, updates all first user feature templates stored in the electronic device by using all second user feature templates, updates the first model stored in the electronic device by using the second model, and updates the first verification threshold stored in the electronic device based on the second verification threshold.

[0124] In a fourth implementation, as shown in FIG. 6D , a verification procedure is executed, and after the verification is successful, the electronic device determines whether the signal-to-noise ratio of the first verification voice is greater than or equal to a first signal-to-noise ratio threshold, and if the signal-to-noise ratio of the first verification voice is less than the first signal-to-noise ratio threshold, proceeds to the next verification procedure; if the signal-to-noise ratio of the first verification voice is greater than or equal to the first signal-to-noise ratio threshold, continues to determine whether the evaluation result obtained in the verification procedure is greater than a first enrollment-free threshold, and if the evaluation result is less than the first enrollment-free threshold, proceeds to the next verification procedure; if the evaluation result is greater than or equal to the first enrollment-free threshold, continues to determine whether the electronic device has received a second model and a second verification threshold, and if the electronic device has not received the second model and the second verification threshold, proceeds to the next verification procedure; and if the electronic device has received the second model and the second verification threshold, processes the first verification voice by using the second model to obtain a second user feature template. The electronic device determines whether the accumulated amount of the second user feature templates has reached a preset amount, and if the accumulated amount of the second user feature templates has not reached the preset amount, proceeds to the next verification procedure, and if the accumulated amount of the second user feature templates has reached the preset amount, updates all first user feature templates stored in the electronic device by using all second user feature templates, updates the first model stored in the electronic device by using the second model, and updates the first verification threshold stored in the electronic device based on the second verification threshold.

[0125] It should be understood that the above lists only four possible combination modes and is not limiting in nature.

[0126] After the electronic device updates the first user feature template stored in the electronic device by using the second user feature template and updates the first model stored in the electronic device by using the second model (in other words, after the first registration-free upgrade is completed), when the user uses the voiceprint verification function again, the electronic device may perform a verification procedure by using the updated model, such as collecting a second verification voice input by the user, processing the second verification voice by using the second model to obtain a third voiceprint feature, and verifying the user's identity based on the third voiceprint feature and the second user feature template. For specific implementation methods, please refer to steps S303 and S304. Details will not be described again in this specification.

[0127] Indeed, after receiving the model that is upgraded to the second model, the electronic device performs a new upgrade of the registration-free upgrade. For example, if the electronic device receives a third model after the verification procedure based on the second model is completed and the user's identity is verified, the electronic device processes the second verification voice by using the third model to obtain a third user feature template, then uses the third user feature template to update the second user feature template stored in the electronic device, and uses the third model to update the second model stored in the electronic device. For specific implementation methods, please refer to step S305. Details will not be described again in this specification.

[0128] A user scenario is used as the above example to describe in detail the registration, verification, and registration-free upgrade of a single user. In specific implementation, the embodiments of the present application can also be applied to a multi-user scenario. The main differences in the multi-user scenario are as follows: the first registration procedure requires simultaneous registration of first user feature templates for multiple users; the verification procedure requires determining the user feature template for the current user from the user feature templates for multiple users to perform identity verification for the current user; and the registration-free upgrade procedure requires simultaneous update of the user feature templates for multiple users.

[0129] As can be seen from the above, in an embodiment of the present application, during the upgrade of the voiceprint recognition system, the verification voice obtained in the verification process is used as the new enrollment voice to complete the upgrade and enrollment, so that the upgrade of the voiceprint recognition system can be performed without the user's perception, and both the voiceprint recognition performance and the user experience can be improved.

[0130] Based on the same technical concept, an embodiment of the present application further provides a chip, which can be coupled to a memory in an electronic device and can perform the methods shown in Figure 3 and Figures 6A to 6D.

[0131] Based on the same technical concept, an embodiment of the present application further provides a computer storage medium, which stores computer instructions, which, when executed by one or more processing modules, implement the methods shown in Figure 3 and Figures 6A to 6D.

[0132] Based on the same technical concept, an embodiment of the present application further provides a computer program product including instructions, which store the instructions, and when the instructions are executed on a computer, the computer can perform the methods shown in Figure 3 and Figures 6A to 6D.

[0133] In this application, unless otherwise specified, " / " should be understood to mean "or." For example, A / B may represent A or B. In this application, "and / or" describes only a related relationship between related objects and represents that three relationships may exist. For example, A and / or B may represent the following three cases: only A is present, both A and B are present, or only B is present. "At least one" means one or more, and "multiple" means two or more.

[0134] In this application, terms such as "example," "in some embodiments," or "in some other embodiments" are used to indicate providing an example, illustration, or explanation. An embodiment or design scheme described in this application as an "example" should not be described as preferred or having more advantages over another embodiment or design scheme. Specifically, the term "example" is used to present concepts in a specific manner.

[0135] Furthermore, terms such as "first" and "second" in this application are merely used for distinction and description purposes and cannot be understood as an indication or implication of relative importance, or an implicit indication of indicated quantities of technical features, or an indication or implication of a sequence.

[0136] Those skilled in the art will understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Thus, the present application may use hardware-only embodiments, software-only embodiments, or embodiments using a combination of software and hardware. Furthermore, the present application may use the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0137] The present application will be described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present application. It should be understood that computer program instructions can be used to implement each process and / or each block in the flowcharts and / or block diagrams, and combinations of processes and / or blocks in the flowcharts and / or block diagrams. These computer program instructions can be implemented in a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or any other programmable data processing device to create a machine, and thus, executed by the computer or processor in any other programmable data processing device, these instructions create an apparatus for implementing the specific functions in one or more processes in the flowcharts and / or one or more blocks in the block diagrams.

[0138] These computer program instructions may be stored in a computer-readable memory that can direct a computer, or any other programmable data processing device, to operate in a specific manner, such that the instructions stored in the computer-readable memory generate artifacts including instruction devices that implement specific functions in one or more processes in the flowcharts and / or one or more blocks in the block diagrams.

[0139] The computer program instructions may be loaded into a computer or other programmable data processing device such that a sequence of operations and steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specific functions in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.

[0140] It is apparent that those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. The present application intends to cover these modifications and variations in the present application as long as they fall within the scope of protection defined by the claims of the present application and their equivalent technologies.

Claims

1. 1. An upgrade method applied to an electronic device, comprising: collecting a first verification voice input by a user; processing the first verification speech to obtain a first voiceprint feature by using a first model stored on the electronic device, and verifying the identity of the user based on the first voiceprint feature and a first user feature template stored on the electronic device, the first user feature template being a voiceprint feature obtained by the electronic device by processing a historical verification speech or an enrollment speech of the user by using the first model; After the identity of the user is confirmed, if the electronic device has received a second model, processing the first verification voice by using the second model to obtain a second voiceprint feature, updating the first user feature template stored on the electronic device based on the second voiceprint feature, and updating the first model stored on the electronic device by using the second model. A method comprising:

2. The step of verifying the identity of the user based on the first voiceprint characteristic and the first user characteristic template stored on the electronic device includes: calculating a similarity between the first voiceprint feature and the first user feature template and determining whether the similarity is greater than a first verification threshold corresponding to the first model, where if the similarity is greater than the first verification threshold, verification is successful, and if the similarity does not exceed the first verification threshold, verification is unsuccessful. Including, after processing the first verification speech using the second model; If the electronic device has received a second validation threshold corresponding to the second model, updating the first validation threshold based on the second validation threshold. The method of claim 1 further comprising:

3. said step of processing said first verification speech by using said second model comprising: processing the first verification speech by using the second model if the quality of the first verification speech satisfies a first preset condition; Including, the first preset condition includes a similarity between the first voiceprint feature and the first user feature template being equal to or greater than a first enrollment-free threshold, and / or a signal-to-noise ratio of the first verification voice being equal to or greater than a first signal-to-noise ratio threshold; The method of claim 1.

4. The method of claim 3 , wherein the first no-enrollment threshold is greater than or equal to a first validation threshold corresponding to the first model.

5. After processing the first verification speech using the second model, If the electronic device has received a second registration-free threshold, updating the first registration-free threshold based on the second registration-free threshold; and / or If the electronic device receives a second signal-to-noise ratio threshold, updating the first signal-to-noise ratio threshold based on the second signal-to-noise ratio threshold. The method of claim 3 further comprising:

6. updating the first user feature template stored on the electronic device based on the second voiceprint feature; and updating the first model stored on the electronic device by using the second model; After the amount of second voiceprint features acquired by the electronic device has cumulatively reached a preset amount, updating the first user feature template stored in the electronic device based on the preset amount of second voiceprint features, and updating the first model stored in the electronic device by using the second model. The method of claim 1 , comprising:

7. After the step of updating the first user feature template stored in the electronic device based on the second voiceprint feature and the step of updating the first model stored in the electronic device by using the second model, collecting a second verification voice input by the user; processing the second verification speech to obtain a third voiceprint feature by using the second model, and verifying the identity of the user based on the third voiceprint feature and the second voiceprint feature. The method of claim 1 further comprising:

8. before the step of collecting the first verification voice input by the user, prompting the user to input a verification voice; The method of claim 1 further comprising:

9. An upgrade device, comprising: a data collection unit configured to collect a first verification voice input by a user; a computing unit configured to process the first verification speech to obtain first voiceprint features by using a first model stored in the device, and verify the identity of the user based on the first voiceprint features and a first user feature template stored in the device, the first user feature template being voiceprint features obtained by the device by processing a historical verification speech or an enrollment speech of the user by using the first model; after the identity of the user is confirmed, if the device has received a second model, the computing unit processes the first verification speech to obtain second voiceprint features by using the second model, and updates the first user feature template stored in the device based on the second voiceprint features; and An apparatus comprising:

10. When verifying the identity of the user based on the first voiceprint feature and the first user feature template stored on the device, the computing unit: calculating a similarity between the first voiceprint feature and the first user feature template, and determining whether the similarity is greater than a first verification threshold corresponding to the first model, where if the similarity is greater than the first verification threshold, verification is successful, and if the similarity does not exceed the first verification threshold, verification is unsuccessful; specifically configured to the computing unit is further configured, after processing the first verification speech by using the second model, to update the first verification threshold based on a second verification threshold corresponding to the second model, if the device receives the second verification threshold; 10. The apparatus of claim 9.

11. When processing the first verification speech by using the second model, the computing unit: If the quality of the first verification speech satisfies a first preset condition, the first verification speech is processed by using the second model. It is specifically structured as follows: the first preset condition includes a similarity between the first voiceprint feature and the first user feature template being equal to or greater than a first enrollment-free threshold, and / or a signal-to-noise ratio of the first verification voice being equal to or greater than a first signal-to-noise ratio threshold; 11. Apparatus according to claim 9 or 10.

12. The apparatus of claim 11 , wherein the first no-enrollment threshold is greater than or equal to a first validation threshold corresponding to the first model.

13. The computing unit After processing the first verification audio by using the second model, if the device has received a second enrollment-free threshold, updating the first enrollment-free threshold based on the second enrollment-free threshold and / or if the device has received a second signal-to-noise ratio threshold, updating the first signal-to-noise ratio threshold by using the second signal-to-noise ratio threshold. The apparatus of claim 11 , further configured to:

14. When updating the first user feature template stored in the device based on the second voiceprint feature and updating the first model stored in the device by using the second model, the calculation unit: After the amount of second voiceprint features cumulatively acquired reaches a preset amount, the first user feature template stored in the device is updated based on the preset amount of second voiceprint features, and the first model stored in the device is updated by using the second model.

10. The apparatus of claim 9, specifically configured to:

15. The computing unit updating the first user feature template stored in the device based on the second voiceprint feature, and collecting a second verification voice input by the user after updating the first model stored in the device by using the second model; processing the second verification speech to obtain a third voiceprint feature by using the second model, and verifying the identity of the user based on the third voiceprint feature and the second voiceprint feature; The apparatus of claim 9 , further configured to:

16. The computing unit prompting the user to input the first verification voice before the data collection unit collects the first verification voice input by the user; The apparatus of claim 9 further configured to:

17. 1. An electronic device comprising a microphone and a processor, the microphone is configured to collect a first verification voice input by a user; The processor is configured to process the first verification speech to obtain first voiceprint features by using a first model stored on the electronic device, verify the identity of the user based on the first voiceprint features and a first user feature template stored on the electronic device, the first user feature template being voiceprint features obtained by the electronic device by processing a historical verification speech or an enrollment speech of the user by using the first model, and after the identity of the user is confirmed, if the electronic device has received a second model, process the first verification speech to obtain second voiceprint features by using the second model, update the first user feature template stored on the electronic device based on the second voiceprint features, and update the first model stored on the electronic device by using the second model. Electronic devices.

18. In verifying the identity of the user based on the first voiceprint characteristic and the first user characteristic template stored on the electronic device, the processor: calculating a similarity between the first voiceprint feature and the first user feature template, and determining whether the similarity is greater than a first verification threshold corresponding to the first model, where if the similarity is greater than the first verification threshold, verification is successful, and if the similarity does not exceed the first verification threshold, verification is unsuccessful; specifically configured to the processor is further configured, after processing the first verification audio by using the second model, to update the first verification threshold based on a second verification threshold corresponding to the second model, if the electronic device has received the second verification threshold.

18. The electronic device of claim 17.

19. When processing the first verification speech using the second model, the processor: If the quality of the first verification speech satisfies a first preset condition, the first verification speech is processed by using the second model. It is specifically structured as follows: the first preset condition includes a similarity between the first voiceprint feature and the first user feature template being equal to or greater than a first enrollment-free threshold, and / or a signal-to-noise ratio of the first verification voice being equal to or greater than a first signal-to-noise ratio threshold; 19. An electronic device according to claim 17 or 18.

20. 20. The electronic device of claim 19, wherein the first no-enrollment threshold is greater than or equal to a first validation threshold corresponding to the first model.

21. The processor: After processing the first verification audio by using the second model, if the electronic device has received a second enrollment-free threshold, updating the first enrollment-free threshold based on the second enrollment-free threshold and / or if the electronic device has received a second signal-to-noise ratio threshold, updating the first signal-to-noise ratio threshold by using the second signal-to-noise ratio threshold.

20. The electronic device of claim 19, further configured to:

22. When updating the first user feature template stored on the electronic device based on the second voiceprint feature and updating the first model stored on the electronic device by using the second model, the processor: After the amount of second voiceprint features acquired by the electronic device cumulatively reaches a preset amount, the first user feature template stored in the electronic device is updated based on the preset amount of second voiceprint features, and the first model stored in the electronic device is updated by using the second model.

20. The electronic device of claim 17, specifically configured to:

23. The processor: updating the first user feature template stored in the electronic device based on the second voiceprint feature and updating the first model stored in the electronic device by using the second model, and then collecting a second verification voice input by the user; processing the second verification speech to obtain a third voiceprint feature by using the second model, and verifying the identity of the user based on the third voiceprint feature and the second voiceprint feature; The electronic device of claim 17 , further configured to:

24. The processor: Prompting the user to input a verification voice before the microphone collects the first verification voice input by the user.

20. The electronic device of claim 17, further configured to:

25. A memory and a chip coupled to the memory, disposed within an electronic device, wherein the memory stores instructions that, when executed by the chip, enable the chip to perform a method according to any one of claims 1 to 8.

26. 9. A computer storage medium having stored thereon computer instructions that, when executed by one or more processing modules, perform the method of any one of claims 1 to 8.

27. A computer program comprising instructions which, when executed on a computer, enable the computer to carry out the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sound user authentication system

    JP2002269047A

  • System, method and computer program product for updating biometric model based on change in biometric feature

    JP2007249179A

  • Method and device for processing voice data

    JP2018081297A

  • Authentication system, authentication method, and program

    JP2019185605A

  • Text independent speaker recognition

    WO2020117639A2